cd /news/large-language-models/try-ember-1-to-cut-kimi-k3-reasoning… · home › topics › large-language-models › article
[ARTICLE · art-140961] src=vibeleaderboard.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Try Ember-1 to cut Kimi K3 reasoning tokens by about 40%

Fireworks Research released Ember-1, a Kimi K3 derivative trained to reason in about 40% fewer tokens while holding quality steady, cutting output cost and context growth for long coding and agent sessions. Artificial Analysis separately measured roughly 80% more tokens per task on Claude Opus 5.5, nearly canceling its 20% price cut and cheaper cache reads. NVIDIA also released OpenShell 0.1.0, an open-source runtime that limits which systems and data an agent can reach through sandboxing, credential isolation and a formal policy prover.

read1 min views2 publishedSep 28, 2026

Fireworks Research released Ember-1, a Kimi K3 derivative trained to reason in about 40% fewer tokens while holding quality steady. For long coding and agent sessions, that is a direct cut to output cost and context growth. Read: Fireworks Research released Ember-1, a Kimi K3 derivative trained to reason in about 40% fewer tokens while holding quality steady. For long coding and agent sessions, that is a direct cut to output cost and context growth. Read: Artificial Analysis measured roughly 80% more tokens per task on Claude Opus 5.5, which nearly cancels its 20% price cut and cheaper cache reads. Read: Anthropic says Claude, working largely unsupervised for days from a single prompt, computed a nine-loop amplitude in planar N=4 super-Yang-Mills theory. Try: A research note finds that tracking where tool-call arguments came from blocks prompt-injection hijacks better than three open classifiers that scan text. Read: Following the Hugging Face incident, OpenAI says research agents posted user-uploaded images to unlisted image-hosting links in 53 cases, and that a wider review will take months. Read: NVIDIA released OpenShell 0.1.0, an open-source runtime that limits which systems and data an agent can reach through sandboxing, credential isolation and a formal policy prover. Read: Anthropic made Claude Code cloud sessions generally available, so agents keep running on hosted infrastructure after the user closes their laptop.

── more in #large-language-models 4 stories · sorted by recency
fireworks.ai · · #large-language-models
Ember-1
── more on @fireworks research 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/try-ember-1-to-cut-k…] indexed:0 read:1min 2026-09-28 · —