{"slug": "fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40", "title": "Fireworks says its Kimi K3 variant cuts reasoning tokens by 40%", "summary": "Fireworks AI says its Ember-1 model, built from Moonshot AI's Kimi K3, produces comparable answers with about 40% fewer tokens, according to a September 27 X post by co-founder Dmytro Dzhulgakov promoting the model Fireworks introduced on September 23. Fireworks reports Ember-1 scoring 92.2% on SWE-bench Verified versus 93.2% for Kimi K3 at maximum reasoning effort, and 82.0% versus 80.9% on Terminal-Bench 2.1, with one customer test showing a 39% reduction in total tokens and a 71.3% reduction in reasoning tokens at a nearly unchanged task score of 0.753 versus 0.751. The company says the work involved more than 50 training experiments and over 200 evaluations, but the results are Fireworks' own evaluations and the two customer tests are not identified, so the savings remain vendor-reported and workload-specific.", "body_md": "# Fireworks says its Kimi K3 variant cuts reasoning tokens by 40%\n\n**Co-founder Dmytro Dzhulgakov promoted Ember-1 on September 27th; Fireworks introduced the model four days earlier with benchmark and customer-test results.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [X](https://x.com/dzhulgakov/status/2104311573640855903)\n\n## Why it matters\n\nFireworks is using its infrastructure to train a Kimi K3-derived model that may lower costs for multi-turn coding agents by shrinking their reasoning traces. Its reported savings are promising, but remain vendor-reported and workload-specific.\n\nFireworks AI co-founder [Dmytro Dzhulgakov (@dzhulgakov)](https://x.com/dzhulgakov) used a September 27th [post on X](https://x.com/dzhulgakov/status/2104311573640855903) to promote [Ember-1](https://runtimewire.com/models/fireworks/ember-1), a Fireworks model built from Moonshot AI's [Kimi K3](https://runtimewire.com/models/moonshotai/kimi-k3) that the company says produces comparable answers with about 40% fewer tokens. Fireworks had introduced Ember-1 on September 23rd, positioning it as a way to reduce the cost of reasoning-heavy coding and agent workloads.\n\nDzhulgakov described Ember-1 as \"hot on HN\" and said it was 40% faster and cheaper. The company's launch materials substantiate the token-reduction claim with benchmark and customer-test results; they do not establish a general 40% improvement in response speed. Fewer generated tokens can reduce per-task costs, but token savings alone do not establish lower latency across different workloads.\n\nDzhulgakov is one of seven Fireworks co-founders. The company's [team page](https://fireworks.ai/team) identifies him as a former PyTorch core maintainer at Meta. CEO Lin Qiao previously led PyTorch at Meta; other founders came from Meta's ads infrastructure, ranking, compiler and News Feed teams, as well as Google's Vertex AI organization. The founding team's experience is in the infrastructure behind large-scale AI systems, and Ember-1 applies that experience to a specific operating cost: the long reasoning traces that can accumulate during automated coding.\n\n### A model trained to think more economically\n\nFireworks says users wanted Kimi K3's coding ability at lower cost, while simply lowering the model's reasoning-effort setting weakened its results. Its researchers instead post-trained a K3-based model to shorten reasoning while retaining the parts that help it recover from mistakes and adapt to feedback. The company says the work involved more than 50 training experiments and over 200 evaluations, across mathematics, coding, search, tool use, software engineering and other tasks.\n\nThe [Ember-1 announcement](https://fireworks.ai/blog/ember-1) reports results across seven benchmarks and production traffic from two customers. On SWE-bench Verified, Fireworks reports a 92.2% score for Ember-1, compared with 93.2% for Kimi K3 at maximum reasoning effort. On Terminal-Bench 2.1, the company reports 82.0% for Ember-1 against 80.9% for K3 at maximum effort. The benchmarks vary in size, from 50 examples on the airline test to 500 on SWE-bench Verified, and the results are Fireworks' own evaluations rather than independent reproductions.\n\nIn one customer test, Fireworks reports a 39% reduction in total tokens and a 71.3% reduction in reasoning tokens, with the measured task score nearly unchanged: 0.753 for Ember-1 against 0.751 for K3. The company says it ran live coding-workload tests with two customers and that one subsequently put Ember-1 into production. It does not identify the customers in the launch materials, so the reported tests offer a useful signal about real workloads without letting outsiders assess the underlying tasks or reproduce the comparison.\n\nThe cost argument is especially relevant to multi-turn agents. Fireworks says these systems may send earlier reasoning back to the model at each turn, so lengthy traces can be processed repeatedly as a task continues. Reducing the trace could cut cumulative token use even when the price per token stays the same. Ember-1 is available through Fireworks' [serverless model page](https://fireworks.ai/models/fireworks/ember-1), which lists prices of $3 per million input tokens, $0.30 per million cached-input tokens and $15 per million output tokens. Those are listed usage rates, not a guarantee that every workload will cost 40% less.\n\n### A platform company shipping its own models\n\nEmber-1 also gives Fireworks a product to sell beyond hosting and serving models built by others: a model it trained on its own infrastructure. Fireworks says the model was developed using its Serverless Training service and is the first in a series of specialized models. Its launch materials described the initial release as a research preview, with two-week access for research models and continued availability tied to developer demand; the model page now lists Ember-1 as ready for serverless use.\n\nThat move follows a large expansion in Fireworks' financing and stated scale. In July, the company announced a $1.505 billion Series D at a $17.5 billion valuation, led by Atreides Management, Index Ventures and TCV. Fireworks also said then that its annualized revenue run rate had passed $1 billion. Those are company-announced figures, and the financing announcement tied the capital to expanding compute infrastructure and engineering. Ember-1 puts that infrastructure to work on a product-level bet: that customers will pay less to run capable models when those models are trained to use fewer tokens, rather than merely hosted more efficiently.", "url": "https://wpnews.pro/news/fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40", "canonical_source": "https://runtimewire.com/article/fireworks-ember-1-kimi-k3-reasoning-tokens", "published_at": "2026-09-27 22:31:03+00:00", "updated_at": "2026-09-27 22:59:22.893341+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-infrastructure", "ai-research", "ai-products"], "entities": ["Fireworks AI", "Ember-1", "Kimi K3", "Moonshot AI", "Dmytro Dzhulgakov", "Lin Qiao", "Meta", "PyTorch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40", "markdown": "https://wpnews.pro/news/fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40.md", "text": "https://wpnews.pro/news/fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40.txt", "jsonld": "https://wpnews.pro/news/fireworks-says-its-kimi-k3-variant-cuts-reasoning-tokens-by-40.jsonld"}}