cd/entity/RunInfra· home› entities› RunInfra
grep -l @runinfra /news/*.json | wc -l → 4

RunInfra

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

20:30
2026-08-26
runinfra.ai
large-language-models

The fastest and cheapest GLM 5.3 Flash endpoint

RunInfra is offering GLM 5.3 Flash, an LLM from zai-org, at $0.10 per 1M input tokens and $0.40 per 1M output tokens, with a 1,048,576-token context window and OpenAI-compatible chat completions. The …

13:26
2026-08-15
runinfra.ai
large-language-models

DeepSeek V4 Flash at 278 tok/s, full precision, no quantization

RunInfra lists DeepSeek V4 Flash, an LLM served as deepseek-ai/DeepSeek-V4-Flash-0731, at $0.13 per 1M input tokens and $0.27 per 1M output tokens, with a 1,048,576-token context window and OpenAI-com…

13:05
2026-07-30
github.com
artificial-intelligence

Kimi k3 run on RTX 5090

RunInfra enables running Kimi-Linear-48B, a distilled version of the full 2.78-trillion-parameter Kimi K3 model, on a single consumer GPU such as the RTX 5090 with 32 GB VRAM, achieving 113.83 tokens …

// co-occurs with top 8 entities
// topics top 6 topics