cd/entity/GSM8K· home› entities› GSM8K
grep -l @gsm8k /news/*.json | wc -l → 98

GSM8K

mentions 98 type Organization page 4/5 feed RSS

// recent coverage 98 mentions

09:00
2026-07-24
promptcube3.com
large-language-models

How to Write Prompts So AI Reliably Outputs What You Want

The Role-Context-Task-Constraint (RCTC) model is the most effective framework for structured prompting, eliminating ambiguity by assigning a professional persona, providing context, specifying an impe…

06:46
2026-07-22
runtimewire.com
large-language-models

Head to head: gpt-oss-120b vs Kimi K3

In a head-to-head test of eight GSM8K math word problems, gpt-oss-120b and Kimi K3 tied on seven tasks, both returning the correct numeric answer every time. The only separation came on the eighth tas…

23:00
2026-07-17
latentheat.dev
artificial-intelligence

My model learned to think. It never learned to listen.

A developer known as flirp spent 21 GPU-hours training Qwen2.5-1.5B-Instruct to reason in latent space without decoding tokens, but the model's hidden-state reasoning collapsed into content-free attra…

01:42
2026-07-11
lesswrong.com
artificial-intelligence

The Termination Circuit (how reasoning models stop thinking).

Researchers discovered that reasoning models like o1 and R1 often overthink, computing answers at around 30% of their chain-of-thought but continuing for the remaining 70%. The termination decision is…

05:53
2026-07-10
github.com
artificial-intelligence

TinyToT – Tree of Thoughts Inference Server

TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…

01:11
2026-07-08
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now

NVIDIA open-sourced Nemotron-Labs-TwoTower, a diffusion language model that generates text 2.42x faster than its autoregressive counterpart without retraining original weights. The model achieves 98.7…

← prev page 4 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics