cd/entity/GSM8K· home entities GSM8K
grep -l @gsm8k /news/*.json | wc -l → 67

GSM8K

mentions 67 type Organization page 3/4 feed RSS

// recent coverage 67 mentions

01:42
2026-07-11
lesswrong.com
artificial-intelligence

The Termination Circuit (how reasoning models stop thinking).

Researchers discovered that reasoning models like o1 and R1 often overthink, computing answers at around 30% of their chain-of-thought but continuing for the remaining 70%. The termination decision is…

05:53
2026-07-10
github.com
artificial-intelligence

TinyToT – Tree of Thoughts Inference Server

TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…

01:11
2026-07-08
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now

NVIDIA open-sourced Nemotron-Labs-TwoTower, a diffusion language model that generates text 2.42x faster than its autoregressive counterpart without retraining original weights. The model achieves 98.7…

04:00
2026-06-29
arxiv.org
large-language-models

Masked Language Flow Models

Researchers introduced Masked Language Flow Models (MLFMs), combining masked diffusion and flow-based methods for efficient language generation. MLFMs enable conditional generation via continuous flow…

04:00
2026-06-24
arxiv.org
machine-learning

Weight-Space Geometry of Offline Reasoning Training

Researchers compared six offline reinforcement-learning methods for distilling reasoning from large language models into smaller ones, finding that SFT, RFT, and RIFT produce nearly identical weight u…

03:30
2026-06-21
FareedKhan-dev.github.io
large-language-models

Train LLM from Scratch

A developer trained a large language model from scratch using plain PyTorch, implementing the full post-training pipeline including SFT, reward modeling, DPO, PPO, and GRPO on public datasets, all run…

05:02
2026-06-19
discuss.huggingface.co
large-language-models

When Should LLMs Verify Instead of Think Longer?

Researchers introduced SEVRA, a serving-layer controller that decides when a frozen reasoning model should verify its answer instead of thinking longer, finding that selective verification improves ac…

23:39
2026-06-05
arxiv.org
machine-learning

Discrete Tilt Matching

Researchers have developed Discrete Tilt Matching (DTM), a likelihood-free method for fine-tuning masked diffusion large language models using reinforcement learning. The approach recasts fine-tuning …

← prev page 3 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics