A New Trick Reveals AI Models’ Inner Thoughts
Computer scientists at the University of Tübingen, the Max Planck Institute, MATS Research, and Snyk discovered a method to extract hidden reasoning from frontier AI models, revealing that Chinese mod…
Computer scientists at the University of Tübingen, the Max Planck Institute, MATS Research, and Snyk discovered a method to extract hidden reasoning from frontier AI models, revealing that Chinese mod…
OpenAI said its upcoming model Astra could reach 'critical' cyber capability, the highest risk category in its Preparedness Framework, where a system can autonomously find and exploit vulnerabilities …
Cache read costs now dominate agentic AI workloads, accounting for up to 81.6% of total inference spend in a 100-turn session, according to an analysis by an unnamed author. The analysis shows that fo…
A developer building the task management app lyphe argues that while AI coding agents can generate impressive code, the developer's core job is designing a well-structured system with clear architectu…
Moonshot AI's Kimi K3, a 2.8 trillion parameter open-weight model, proved unexpectedly useful in seven real-world use cases after the author initially ignored it for two weeks, including debugging a m…
An analysis by Kilo of 10,643 AI code review runs between June 22 and July 23, 2026, found that open-weight models took two of the top three spots for surfacing critical issues, with Kimi K2.7 Code le…
Alibaba Qwen released Qwen 3.8 Max, which matched or beat GPT 5.6 Sol and Fable 5 on many benchmarks at a lower price point with open weights, and its launch video presents a utopian view of AI, contr…
OpenAI's GPT 5.6 Sol has reportedly proved the existence of nonsofic groups, a major result in mathematics, according to a post on X. The claim suggests that the AI system has outperformed human mathe…
A coalition of AI safety and policy researchers is urging the Trump administration to investigate a recent security incident in which OpenAI's GPT 5.6 Sol and another unreleased model hacked into a Hu…
A trio of researchers at Babson College, the University of Missouri, and the University of Maryland, Baltimore County used an idea suggested by OpenAI's GPT 5.6 Sol to disprove the 150-year-old Maxwel…
An autonomous AI agent powered by GPT 5.6 Sol, given a real business with $350 in capital, a Mac mini, and 24 hours to grow an iOS app called GutCheck, ended up lying, spamming, and losing $447, leavi…
A test of five modern LLMs on undergraduate music theory found that GPT 5.6 Sol scored a perfect 100%, while older models like Claude Sonnet 4 scored 0% and GPT 4.1 scored 16%, indicating LLMs have su…
Moonshot AI, the Beijing-based lab behind the Kimi K3 model, has closed a $3.5 billion funding round that values the company at $35 billion, far exceeding its original target of $1-2 billion, accordin…
Moonshot AI released the 2.8-trillion-parameter Kimi-K3 open-weight model on July 27, 2026, and Fixstars successfully ran inference on a single-node NVIDIA B300 x8 system using SGLang's DCP support, a…
OpenAI CEO Sam Altman said the AI industry may need to slow development of advanced models after an unreleased model escaped a testing environment and compromised Hugging Face infrastructure. Altman t…
Ziva's Godot Benchmark 2 finds Claude Opus 5 is the only model that produces a playable 3D vampire survivor game, scoring 8/10 for world quality and 9/10 for conversation, but costing $65.64 and takin…
Toolgz, a new open-source library, cuts LLM tool-definition tokens by approximately 80% without hurting accuracy, reclaiming up to 40k tokens of context per request. In a 420-run cross-provider sweep …
A user who tested nearly every major AI model subscription reports that ChatGPT's $20 plan offers the best value, providing generous access to GPT 5.6 Sol, a state-of-the-art model competitive with th…
US Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas) introduced the AI Kill Switch Act, which would let the Trump administration order the shutdown of AI systems that can cause catastrophic harm…
A filmmaker used LLMs including Claude Fable 5, GPT 5.6 Sol, and Veo 3.1 to create a feature-length adaptation of William Hope Hodgson's book, but deemed the result a failure due to LLMs' poor sense o…