cd/entity/MATH-500· home entities MATH-500
grep -l @math-500 /news/*.json | wc -l → 20

MATH-500

mentions 20 type Organization feed RSS

// recent coverage 20 mentions

04:00
2026-08-17
machinebrief.com
artificial-intelligence

KV Cache Compression Through the Lens of Transform Coding

Researchers introduced Attention-Aware Transform Coding (AATC), a KV cache compression method that reduces memory use by approximately 5.8x while maintaining near-lossless accuracy on models like Llam…

14:34
2026-08-06
ycrootaccess.com
machine-learning

Eight ML Papers, Explained by the Researchers Behind Them

At YCML, YC's first machine learning research showcase at Startup School 2026, eight researchers presented papers on topics including model reasoning, formal mathematics, video agents, and robotics. N…

01:05
2026-07-26
promptcube3.com
large-language-models

Kimi K3 vs GPT-5 vs Claude 4 Opus: 2026 Comparison

Kimi K3 is 30x cheaper than GPT-5 for output tokens while leading the LMArena leaderboard with 1,289 ELO, according to a 2026 comparison. For a customer support bot handling 10M input and 5M output to…

19:55
2026-07-14
machinebrief.com
artificial-intelligence

Cracking the Code: How SPARK Enhances AI Reasoning

SPARK, a new approach to diagnosing hidden-state failures in AI reasoning, boosted accuracy on the MATH-500 benchmark for Qwen3-4B from 82.0% to 84.6% and for Qwen3-8B from 82.4% to 85.6%, according t…

16:01
2026-06-30
pub.towardsai.net
ai-agents

AI Agent Evaluation: How to Know If Your Agent Actually Works

A developer recounts pushing an agent into production that failed after a CRM dropdown change, highlighting the inadequacy of model-level evaluation for agent systems. The article argues that agents m…

20:00
2026-06-22
haoailab.com
large-language-models

JetSpec

JetSpec, a new speculative decoding method, trains a causal parallel draft head over fused hidden states from a frozen target model, enabling lossless verification of candidate trees in one forward pa…

04:00
2026-06-16
arxiv.org
large-language-models

Evaluating the Robustness of Proof Autoformalization in Lean 4

Researchers at UC Riverside introduced the first robustness study for proof autoformalization in Lean 4, testing seven LLM-based models on perturbed informal proofs. All models showed sensitivity to g…

// co-occurs with top 8 entities
// topics top 6 topics