We Tried ISO-AdamW. AdamW Kept Its Job.
A controlled pilot on a dedicated NVIDIA H200 GPU found ISO-AdamW scored 758 correct answers versus 754 for baseline AdamW on a 1,000-question held-out test, a +0.4% delta with a 95% confidence interv…
A controlled pilot on a dedicated NVIDIA H200 GPU found ISO-AdamW scored 758 correct answers versus 754 for baseline AdamW on a 1,000-question held-out test, a +0.4% delta with a 95% confidence interv…
A reasoning-aware compression framework that selectively restores the most quantization-sensitive circuits to FP16 achieves Pareto-optimal energy-accuracy points unreachable by uniform quantization, a…
Researchers introduced AdaptiveSpec, a training-free per-step speculative decoding method that adapts both draft verification and tree shape from internal decoding signals, improving throughput over E…
A new study from arXiv (2609.02899v1) finds that benchmark contamination inflates absolute scores but rarely reorders large language model (LLM) leaderboards, with a rank correlation of 0.997 between …
LLM benchmark scores are becoming unreliable because models are trained on the same tests they are evaluated on, leading to inflated results that fail to reflect real-world performance, according to a…
Researchers introduced Proof-Verified Benchmark Rewriting (RePro), the first framework integrating Lean-oriented neural automated theorem provers into benchmark rewriting, ensuring rewritten math prob…
A new dependency-free audit script from developer Ashish Sinha found that six human-authored Hugging Face benchmarks—BIRD-CRITIC 1.0, Spider, GSM8K, and HumanEval—are clean, with only 18 genuinely dup…
XPress, a lightweight causal refiner for diffusion drafters in speculative decoding, raises acceptance length by ~30% on average (up to +56%) and decoding throughput by ~1.3× on average (up to 1.7×) c…
A new study from arXiv (2608.23807v1) characterizes serving of masked diffusion language models (dLLMs) using LLaDA-8B-Instruct with a D2F LoRA adapter on a single NVIDIA H200 GPU, finding that reques…
CodeSOTA, an open data terminal for reinforcement learning environments and state-of-the-art models, has launched a registry tracking 9,102 results, 163 models, 371 datasets, and 9 capability areas, w…
Google DeepMind, with researchers from the University of Texas at Austin, published a paper (arXiv:2608.17981) introducing 'recirculation,' an inference-time technique that feeds deep-layer activation…
Researchers introduced Thermo-FL, a thermal-aware federated fine-tuning framework for large language models on edge devices, which uses device temperature to control local adapter training and sparse …
Researchers introduced Skill Entropy RL, a reward that quantifies the difficulty of switching between reasoning skills, and showed it more than doubles the accuracy of Qwen3 models on the Skill²-Bench…
Fortitude Omnis Group's OmnisBench benchmark initially showed small models performing surprisingly well, but community feedback revealed potential contamination from old benchmarks like HumanEval and …
A developer testing Qwen3.5 4B via Ollama found that enabling chain-of-thought can return an empty response while the thinking field holds content, and that self-critique can degrade accuracy, citing …
Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…
A new empirical analysis from arXiv (2608.19072v1) finds that large language model (LLM) agents post-training an LLM lock in their training strategy at the very beginning and spend the remaining budge…
Inco AI released DFlash 2, an updated version of its parallel speculative-decoding method, reporting a 21% increase in average accepted token sequence length over the original DFlash, with only 1.3% a…
IBM Research has introduced BenchDrift, an open-source tool that quantifies how much large language model benchmark scores shift when test prompts are rephrased without changing meaning, revealing tha…
The developers behind the open-source OmnisBench project have released a benchmark designed to measure the true cost savings of LLM routing. Running 364 coding and math tasks across four models, they …