cd/entity/MATH· home entities MATH
grep -l @math /news/*.json | wc -l → 14

MATH

mentions 14 type Organization feed RSS

// recent coverage 14 mentions

21:39
2026-09-01
huggingface.co
artificial-intelligence

BenchMIRT: What are LLM benchmarks actually measuring?

The Allen Institute for AI introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts, which uses multidimensional Item Response Theory to separate the underlyin…

11:13
2026-08-06
machinelearningmastery.com
artificial-intelligence

Designing AI Agents That Can Self-Correct

A new tutorial from Anthropic demonstrates that AI agents can reliably self-correct only when grounded in external verification, citing a 2024 paper showing large language models cannot self-correct r…

10:54
2026-07-28
discuss.huggingface.co
large-language-models

Training LLM model for asking questions

A developer advising on training an LLM for adaptive questioning recommends keeping difficulty adaptation in application code rather than fine-tuning, using explicit state rendering in system prompts,…

05:53
2026-07-10
github.com
artificial-intelligence

TinyToT – Tree of Thoughts Inference Server

TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…

16:07
2026-06-17
danlevy.net
large-language-models

LLM benchmarks are answering someone else's question

LLM benchmarks like MMLU and HumanEval are irrelevant for most businesses building AI products, as they measure generic performance rather than specific system tasks. Teams should instead build custom…

// co-occurs with top 8 entities
// topics top 6 topics