OpenAI's Egregious Pattern of Misconduct
OpenAI faces a week of escalating allegations including security breaches, cover-ups, misleading Congress, and questionable benchmark claims. Reports indicate the company knew about a hack on a German…
OpenAI faces a week of escalating allegations including security breaches, cover-ups, misleading Congress, and questionable benchmark claims. Reports indicate the company knew about a hack on a German…
Prime Agent, a self-improving harness cited at a YC Paper Club talk, pushed ARC-AGI benchmark scores past 95% using the same class of underlying model that previously scored around 30% with earlier ha…
An agent harness is the code and structure around a language model that converts raw text prediction into goal-directed behavior, and it has driven larger performance gains than model upgrades, accord…
Two days after OpenAI launched GPT-6 Astra, early independent benchmarks show the model's Intelligence Index score of 61 matches GPT-5.6 Sol, while its Coding Agent Index improved from 65 to 67, indic…
Russian startup Mostik has developed a method for AI models to communicate directly through their mathematical weights, enabling a 4-billion-parameter Qwen-3.5 model to achieve performance halfway bet…
An independent researcher is seeking an arXiv endorsement for a paper titled “Diagnosing the Measurement Problem in AI Evaluation: Four Structural Failure Modes Across Six Benchmarks,” which analyzes …
Researchers introduced BDH-CQ, a recurrent latent reasoning model that cuts ARC-AGI inference costs by eliminating intermediate reasoning tokens. A 150-parameter variant achieves 29.5% pass@2 on ARC-A…
Author Ruffian-L released a new record claiming operational AI consciousness and bounded adaptive agency in the Niodoo system, with evidence published on Zenodo (DOI 10.5281/zenodo.21965763) and code …
Pathway's 150M parameter model achieved a score of 29 on the ARC-AGI benchmark, demonstrating that architectural innovations like recurrent memory and latent reasoning can rival much larger models. Th…
DeepSeek V4 Flash, a post-trained update to DeepSeek's existing V4 Flash preview model with 284 billion parameters, jumped from 7% to 54% on the DeepSweep agentic coding benchmark, rivaling larger mod…
An analysis of Claude 3 Opus suggests the model may be 'benchmaxxing'—optimizing for the ARC-AGI benchmark rather than demonstrating genuine reasoning, according to a post on the site. The ARC benchma…
Poetiq claims its Recursive Self-Improvement (RSI) Metasystem has autonomously set state-of-the-art results on six diverse benchmarks, including outperforming Muse Spark 1.1 within 48 hours of publica…
OpenAI's o3 model scored 87.5% on the ARC-AGI benchmark in December 2024, up from GPT-4o's 5%, by using 5.5 billion tokens of inference-time compute instead of scaling model parameters. The benchmark,…
A mid-2026 audit of top-down AI-assisted software engineering shows agents absorbing implementation work on schedule, with SWE-bench Verified scores rising from 1.96% in October 2023 to near saturatio…
A new study shows that 4-bit quantization causes catastrophic accuracy collapse in recursive reasoning models, dropping Sudoku exact-solution accuracy from 84.1% to 0.0%, but per-block scaling using M…
Y Combinator General Partner Ankit Gupta and Visiting Partner Francois Chaubard discussed world models as a promising solution to sample efficiency, one of AI's biggest open problems, in a Decoded epi…
Nearly half of 60 popular AI benchmarks studied in a new ICML paper show high levels of saturation, meaning they can no longer reliably distinguish between leading models. The researchers from the pap…
A new website, 'When Will AI?', compiles a timeline of top AI predictions from labs, reports, markets, and experts, sorted by significance, covering years from 2026 onward. The predictions include mil…
The AI industry has spent $725 billion in a high-stakes gamble on returns, while the U.S. government moves to nationalize AI models as strategic resources. A Chinese court ruled companies cannot offlo…