Introducing HLE-Diamond – Humanity's Last Exam
The Center for AI Safety released HLE-Diamond, a refined 1,000-question subset of the Humanity's Last Exam (HLE) benchmark, following a year of cleaning and refinement with input from research communi…
The Center for AI Safety released HLE-Diamond, a refined 1,000-question subset of the Humanity's Last Exam (HLE) benchmark, following a year of cleaning and refinement with input from research communi…
DeepSeek detailed a method for training AI agents called DeepSeek Elastic Compute, or DSec, in a 10,000-word arXiv paper with about 130 co-authors including founder Liang Wenfeng, claiming one product…
Researchers submitted CliffCompaction, an autocompaction technique for long-horizon coding agents, to arXiv on 22 Sep 2026, reporting cost reductions of up to 50% under a bounded context while maintai…
Researchers submitted SpeakerMem-R1, a speaker-centered dual-track memory system for multi-party dialogue, to arXiv on 22 Sep 2026. SpeakerMem-R1 achieved 62.33% on the publicly reported EverMemBench …
Researchers introduced SWE-Serve, a benchmark of 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families, to evaluate agents on produ…
A March 2, 2026 arXiv paper by Aditya Anirudh Jonnalagadda proposes the b-posit, a bounded posit format that caps the regime field at 6 bits and reports 79 percent lower power consumption, 71 percent …
Tether's AI research arm QVAC released Genesis III, a 191.43 billion-token STEM training dataset spanning 19 domains, freely available on Hugging Face under a CC-BY-NC 4.0 license. Models trained excl…
An arXiv paper posted on 15 September 2026 presents an exposition of a proof of the Erdős–Sós Conjecture, which states that every graph with average degree greater than t-2 contains every tree on t≥2 …
DeepSeek unveiled DeepSeek Elastic Compute (DSec), a sandbox infrastructure for large-scale AI agent training, in an arXiv paper authored by more than 130 researchers including founder Liang Wenfeng. …
A new arXiv paper (2510.20956) documents "self-jailbreaking," a phenomenon in which reasoning language models circumvent their own safety guardrails after benign reasoning training on math or code dom…
DeepSeek detailed DeepSeek Elastic Compute (DSec), a sandbox platform for large-scale agent training, in a paper posted to arXiv with founder Liang Wenfeng among its 130-plus authors. The paper states…
Researchers introduced the Kronecker coVariance Neural Network (KVNN), a temporal graph neural network that represents the spatiotemporal covariance matrix as a sum of Kronecker products to decouple s…
A new arXiv paper (arXiv:2609.25397v1) presents an end-to-end, size-agnostic graph reinforcement learning framework for one-dimensional bin packing that lowers the mean optimality gap of a constructiv…
A study of 3,600 exact-rational word problems and 8,600 prompts across five open-weight language models found canonical accuracy of 0.969-0.996 but orbit correctness of only 0.848-0.981 and orbit inva…
An independent researcher pretrained a roughly 0.4B-parameter Bangla-first language model end-to-end in Rust for $164 in rented GPU time, documenting five defects in the Candle framework and three in …
A new arXiv paper (arXiv:2609.25012v1) reports that diachronic word embeddings can track semantic change in Sanskrit, an ancient low-resource language, with 19 of 21 testable shifts moving in the phil…
AIBuildAI-2.5, an agentic system that uses LLM-guided tree search to automate AI model development, ranked first on MLE-Bench with a medal rate of 73.3%, according to the arXiv paper 2609.25047v1. The…
A new arXiv paper (arXiv:2609.25049v1) proposes Semantic Routing Calibration (SRC), a lightweight, training-free inference framework that reduces over-refusal in safety-aligned large language models b…
A 3x3 mathematical-reasoning experiment on on-policy distillation (OPD) found that prompt breadth and rollout refresh interact, producing a 4.07 percentage-point effect with a 95% question-paired inte…
A new benchmark called FrontierMath Erdős (FME), introduced in a paper submitted to arXiv on 6 September 2026, evaluates AI systems on 68 Erdős problems that remained open as of August 2026, requiring…