When Models Become a Commodity
OpenRouter's Fusion compound API, which fans prompts to multiple models and combines answers with a judge model, beat most frontier models in tests, with a panel of cheap models landing within 1% of C…
DeepSeek is a Chinese AI research laboratory that has developed highly capable open-source language models including DeepSeek-V3 and DeepSeek-R1, notable for their efficiency and performance.
OpenRouter's Fusion compound API, which fans prompts to multiple models and combines answers with a judge model, beat most frontier models in tests, with a panel of cheap models landing within 1% of C…
A developer scrapped 9 out of 10 planned LLM evaluation experiments after running only CI diagnostics, finding that one real-world use case—triaging flaky integration tests across multiple languages—p…
A leak at DeepSeek exposed proprietary engineering pipelines, including data curation, RL alignment recipes, and system configurations, rather than model weights. The incident highlights that the true…
Chinese AI startups Moonshot AI, Z.AI, and DeepSeek are challenging U.S. labs by releasing models that match or exceed top-tier performance at a fraction of the cost, upending expectations that U.S. f…
The Trump administration is threatening sanctions against Chinese AI labs over alleged model theft, but American startups are already switching to cheaper Chinese models. Treasury Secretary Scott Bess…
Large language models are fundamentally stochastic text generators that predict the next token, not thinking machines, according to a technical breakdown of concepts from Andrej Karpathy. The key dist…
A developer built a KDE Plasma panel HUD for Arch Linux that tracks AI quota usage across Claude, Codex, Gemini, and DeepSeek using four circular indicators, after fixing a bug caused by relying on po…
OpenAI CEO Sam Altman will brief the Trump administration this week on the company's most powerful AI model, which autonomously disproved the 80-year-old Erdős unit distance conjecture and breached Hu…
A software engineer learned from Andrej Karpathy's video 'Deep Dive into LLMs like ChatGPT' about the difference between base and chat models, the role of GPUs in training, and what is released when a…
DeepSeek told prospective investors on July 25 to halt a second funding round that would have valued the Chinese AI lab at $71 billion, triggered by a leaked transcript of founder Liang Wenfeng's priv…
ChangXin Memory Technologies (CXMT) priced its Shanghai STAR Market IPO at 8.66 yuan per share on July 25, raising roughly $9.8 billion at an $85.2 billion valuation, with trading beginning July 27. T…
Open-weight AI models now handle roughly 70% of production tokens routed through OpenRouter, up from 30% a year ago, as enterprises shift from frontier commercial models to cheaper open alternatives. …
Leaked comments from DeepSeek founder Liang Wenfeng reveal a strategy of corporate restraint, low-margin pricing, and open-source models, prioritizing long-term AI ambition over short-term gains. The …
Chinese AI pioneer DeepSeek has paused its second funding round, which aimed to raise at least $1.5 billion at a valuation of at least $71 billion, days after viral posts attributed to founder Liang W…
The U.S. government is weighing whether AI distillation — a standard technique for training one model using another's outputs — constitutes model theft when done without authorization, as a House bill…
DeepSeek has suspended its second fundraising round after leaked comments from founder Liang Wenfeng about China's dependence on Nvidia chips and the country's AI gap with the United States went viral…
Anthropic announced it cut Claude Code's system prompt by more than 80% for the Claude 5 generation with no measurable regression on coding evals, urging developers to rely less on prompt engineering.…
DeepSeek has suspended its second fundraising round after comments widely attributed to founder Liang Wenfeng about US-China AI competition went viral, people familiar with the matter said. The Chines…
Chinese AI models are gaining popularity in the U.S. for their affordability and efficiency, with Mozilla CTO Raffi Krikorian switching to Moonshot's Kimi K3 and Coinbase adopting Chinese models to cu…
The compute gap between hyper-funded US proprietary models and constrained open-weight labs is driving a shift toward efficiency, with developers abandoning overpriced enterprise tiers for localized s…