The boring layer around your LLM call
A developer building LLM applications found that the model itself is only a third of the work, with the surrounding infrastructure—timeouts, retries, and token management—proving critical. They implem…
DeepSeek is a Chinese AI research laboratory that has developed highly capable open-source language models including DeepSeek-V3 and DeepSeek-R1, notable for their efficiency and performance.
A developer building LLM applications found that the model itself is only a third of the work, with the surrounding infrastructure—timeouts, retries, and token management—proving critical. They implem…
National Cyber Director Sean Cairncross said Tuesday at the Black Hat cybersecurity conference that the Trump administration wants U.S.-built open-source AI to become the preferred technology worldwid…
ByteDance's AI research division Seed has operated without using knowledge distillation since its inception, a divergence from other Chinese AI labs that learn from US-built models. The company aims f…
Goldman Sachs raised its run-rate revenue forecast for China's artificial intelligence model market by 30% to US$13 billion, citing aggressive price cuts, technical breakthroughs, and accelerating cor…
ModelPlane, a new OpenAI-compatible gateway, lets developers route AI requests across OpenAI, Anthropic, DeepSeek, OpenRouter, and custom services through a single API, with features like model groups…
DeepSeek released DeepSeek-V4-Flash-0731 on July 31, an updated version of its existing deepseek-v4-flash API model that now outperforms DeepSeek's V4-Pro-Preview on all nine agent and coding benchmar…
The Trump administration met with tech leaders on Tuesday to finalize a security review process for advanced AI models before release, but it will reportedly apply only to closed models from developer…
DeepSeek has reportedly reopened talks for a second funding round targeting RMB50 billion, potentially valuing the Chinese AI company at about RMB500 billion before the financing, with an agreement po…
DeepSeek V4 Flash ranked first in OpenRouter's weekly model-usage ranking for July 27 to Aug. 2, processing 7.22 trillion tokens on the multi-model aggregation platform. On Aug. 1, the model processed…
Chinese labs released 10 open-weight AI models in the past 30 days, including DeepSeek-V2.5 (236B parameters, Apache 2.0) and Qwen-1.5-110B-Chat, challenging US counterparts like Llama 3 and Gemma 2 w…
Data gravity is compounding integration debt in API-first AI stacks, making vendor lock-in a gradual process that most teams drift into rather than choose, according to a technical analysis by Glukhov…
DeepSeek released V4 Flash, an open-source AI model that outperforms most top Western models, costs 60% less per task than OpenAI's GPT-5.6 Luna after an 80% price cut, and runs on cheaper hardware, a…
A developer's measurement harness testing six Chinese LLM APIs found that GLM hallucinated a detailed, non-existent Chinese website for Airtable, complete with a .cn domain, pricing in RMB, and locali…
Two independent benchmark sites, BenchLM.ai and llm-stats.com, compared DeepSeek's V4 Flash 0731 update with Alibaba's Qwen3.6-27B, finding that DeepSeek V4 Flash is roughly 12× cheaper per token and …
DeepSeek's V4 Flash 0731, a 284B-parameter Mixture-of-Experts model with 13B active parameters and a native 1M-token context under an MIT license, is cheaper to use via the API than to run locally, ac…
DeepSeek's V4 Flash and Qwen-3.8-Max have rewritten the cost equation for large language models, with Nous Research offering DSV4 Flash at a 90% discount and OpenCode reporting 8 trillion tokens proce…
DeepSeek released DeepSeek-V4-Flash-0731 on July 31, a $0.14 per million input tokens model that outperforms its own flagship, a 1.6 trillion parameter model, on all nine agent benchmarks published by…
A developer cut a nightly batch classification job from 40 minutes to 6 minutes by switching to DeepSeek-V4-Flash and adopting an async request pattern. The model tier swap and concurrency changes wer…
DeepSeek's V4 Flash 0731 build corrupts integer fields in strict JSON schema outputs when thinking mode is enabled, failing 8 of 13 runs across two request paths, while disabling thinking fixes all ru…
OpenAI's GPT-5.6 Sol rewrote production GPU kernels, cutting serving costs by 20%, and an internal model Astra produced ten formally proved mathematical results, while OpenAI cut GPT-5.6 Luna API pric…