cd/entity/Qwen3-8B· home entities Qwen3-8B
grep -l @qwen3-8b /news/*.json | wc -l → 64

Qwen3-8B

mentions 64 type Organization page 3/4 feed RSS

// recent coverage 64 mentions

11:11
2026-07-11
machinebrief.com
machine-learning

Procedural Memory: A New Era in Reinforcement Learning

Procedural Memory Distillation (PMD) advances reinforcement learning by transforming experiences across episodes into actionable intelligence, outperforming previous models like SDPO by 3.8-5.5% on SC…

01:42
2026-07-11
lesswrong.com
artificial-intelligence

The Termination Circuit (how reasoning models stop thinking).

Researchers discovered that reasoning models like o1 and R1 often overthink, computing answers at around 30% of their chain-of-thought but continuing for the remaining 70%. The termination decision is…

15:21
2026-07-10
byteiota.com
artificial-intelligence

NVIDIA Nemotron-Labs-Diffusion Kills the Draft Model

NVIDIA released Nemotron-Labs-Diffusion, a single model that eliminates the need for a separate draft model in speculative decoding, achieving 6.82 accepted tokens per forward pass in self-speculation…

13:24
2026-07-10
arxiv.org
large-language-models

DominoTree

Researchers introduced DominoTree, a training-free best-first draft tree method for speculative decoding that uses Domino's conditional correction to achieve up to 6.6x speedup over autoregressive dec…

14:01
2026-06-27
dev.to
large-language-models

The Developer's Guide to Trimming AI API Costs Without Crying

A backend engineer at an unnamed company slashed their team's LLM API costs from $11,400 to $1,830 per month by switching to cheaper models for most tasks and implementing tiered routing. The team rep…

16:20
2026-06-26
dev.to
large-language-models

How I Cut Our AI API Bill by 95%: What Actually Worked

A developer cut their company's AI API bill by 95% from $11,000 to under $400 per month by implementing per-request model routing and tiered escalation. The team replaced expensive GPT-4o calls with c…

20:00
2026-06-22
haoailab.com
large-language-models

JetSpec

JetSpec, a new speculative decoding method, trains a causal parallel draft head over fused hidden states from a frozen target model, enabling lossless verification of candidate trees in one forward pa…

00:57
2026-06-22
lesswrong.com
large-language-models

NLA explanations can be shortened without harming reconstruction

Researchers trained Qwen3-8B natural language autoencoders with varying length penalties and found that explanation length can be significantly reduced without harming reconstruction fidelity, suggest…

← prev page 3 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics