How Much Memory Does Your Agent Actually Need?
IBM Research's ALT K-Evolve framework shows that the optimal amount of agentic memory varies by model capability, with strong models like DeepSeek-V3.2 (671B MoE) gaining +9.5 percentage points in tas…
IBM Research's ALT K-Evolve framework shows that the optimal amount of agentic memory varies by model capability, with strong models like DeepSeek-V3.2 (671B MoE) gaining +9.5 percentage points in tas…
A June 2026 study by LightSpeed and Tencent found that four frontier models—Claude Sonnet 4.5, DeepSeek-V4-Pro, GLM-5, and Gemini-3-Flash—misjudged their total context size with median relative error …
Vernor Vinge, a computer scientist and science fiction author, predicted in his 1993 essay that the Singularity—a point where machines surpass human intelligence and accelerate progress beyond human c…
Code Arena's WebDev leaderboard, updated August 6, 2026, shows the Elo gap between proprietary and open-weight AI coding models has narrowed from roughly 150 points to just 11, with Anthropic's Claude…
TileRT's persistent engine on NVIDIA GPUs achieves up to 500 tokens/s/user on the InferenceX GLM5 FP8 744B benchmark on a single B200 decode server, approximately 3× faster than GB300 NVL72 running tr…
Self-improving AI harnesses are real and operational, with systems like Shanghai AI Lab's Self-Harness (arXiv 2606.09498) boosting held-out pass rates on Terminal-Bench-2 from 40.5% to 61.9% on MiniMa…
A new arXiv preprint (2607.28636v1) introduces Chain-of-Models (CoM), an automated audit pipeline in which a second model inspects a first model's reasoning trace to reduce cognitive biases in LLM jud…
A bootcamp graduate discovered they were overpaying for AI by up to 40x after analyzing API costs. By switching from GPT-4o to models like DeepSeek V4 Flash via Global API, the developer reduced a $50…
A developer benchmarked 10 large language models on five coding tasks, finding that DeepSeek V4 Flash offers the best value-to-quality ratio at $0.25 per million output tokens, while Qwen3-Coder-30B e…
A backend engineer migrated three production services from OpenAI to DeepSeek V4 Flash via a Global API endpoint, reducing monthly costs from $500 to approximately $12.50—a 40× price difference—while …
OpenClaw's GLM-5 inference optimization study found that adjusting parameters like chunked prefill size and request concurrency improved throughput and reduced latency, cutting serving costs by 10.4% …
A developer cut their monthly AI bill from $487.92 to $12.50 by switching from OpenAI's GPT-4o to DeepSeek V4 Flash via the Global API, achieving a 97.5% cost reduction. The migration required changin…
A bootcamp graduate tested ten AI coding models on five tasks, scoring them on code quality, readability, and explanation clarity. The experiment found that DeepSeek V4 Flash and DeepSeek Coder offer …
A developer migrated from OpenAI's GPT-4o to DeepSeek V4 Flash via a global API provider, reducing monthly costs from $487 to $12.50 with only two lines of code changed. The switch required only alter…
A B2B SaaS startup cut its LLM inference costs by 97% by switching from GPT-4o to cheaper alternatives like DeepSeek V4 Flash, reducing a $14,200 monthly OpenAI bill to an estimated $355. The develope…
An engineer at a company using OpenAI's GPT-4o for LLM inference discovered they were overpaying by up to 40x compared to alternatives like DeepSeek V4 Flash served through Global API. After benchmark…
Ollama's expansion to support Chinese AI models like Kimi-K2.5, GLM-5, MiniMax, and DeepSeek offers local deployment benefits but comes with hidden costs. Developers face documentation gaps, quantizat…
Prime Intellect released prime-rl 0.6.0, an open framework for asynchronous reinforcement learning on trillion-parameter Mixture-of-Experts models, enabling training on agentic RL workloads with optim…
Researchers developed VibeThinker-3B, a 3-billion-parameter language model that achieves reasoning performance matching or exceeding models orders of magnitude larger, scoring 94.3 on AIME26 and 80.2 …
Researchers introduced Self-Harness, a new paradigm enabling LLM-based agents to iteratively improve their own operating harnesses without human intervention. In tests on Terminal-Bench-2.0, Self-Harn…