{"slug": "focusing-on-post-training", "title": "Focusing on Post-Training", "summary": "Fireworks post-trained Kimi K3 to create Ember-1, a model that uses roughly 40% fewer tokens while maintaining comparable quality across Fireworks' evaluations, according to the company's announcement. The model was trained to produce shorter reasoning traces, an approach that makes token usage part of the training objective through token-efficient and budgeted reinforcement learning methods. The result illustrates that post-training an existing open-weight LLM can deliver substantial efficiency gains without duplicating pre-training efforts at frontier scale.", "body_md": "# Focusing on Post-Training\n\n[Ember-1](https://sebastianraschka.com/llm-architecture-gallery/#card-ember-1) is a nice example of what can be built on top of an existing open-weight LLM. Fireworks started with [Kimi K3](https://sebastianraschka.com/llm-architecture-gallery/#card-kimi-k3) and post-trained it to produce shorter reasoning traces. According to the [Fireworks announcement](https://fireworks.ai/blog/ember-1), Ember-1 uses roughly 40% fewer tokens while maintaining comparable quality across their evaluations.\n\nI was recently spontaneously asked on a podcast, given an imaginary multi-million dollar budget, how I would approach developing a frontier LLM with a limited budget. My recommendation was to start with an existing model and spend that budget on post-training. ([Here’s a link](https://www.youtube.com/watch?v=fpYG6OBKEZw), but sorry, it’s in German; my only German podcast.)\n\nThe point is, even millions of dollars is not a lot for training a competitive LLM at the frontier (500B parameters or larger). And given that there is such a strong and good selection of open-weight models out there, there is currently no reason to spend all that money on duplicating pre-training efforts.\n\nEmber-1 is a concrete example of the kind of improvement I had in mind. Sure, here they only focused on efficiency while maintaining existing capabilities, but a 40% reduction in token use is a lot. (Another angle could be maintaining the current efficiency level of an existing model but pushing its capabilities in newer harnesses, for example.)\n\nHow did they achieve the 40% reduction in token use? They trained the model to produce shorter reasoning traces. This connects to the token-efficient and budgeted reinforcement learning methods I covered in [Controlling Reasoning Effort in LLMs](https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms).\n\nThe general idea is to make token usage part of the training objective. For example, we can specify a reward that factors in both answer correctness and response length, or encourages correct solutions within a specified token budget. The challenge is to reduce unnecessary reasoning while retaining the ability to work through difficult problems.\n\nThat’s the connection I wanted to illustrate in the figure below. (But please note that Fireworks hasn’t disclosed enough of Ember-1’s training recipe to identify the exact algorithm, so the connection to those methods is conceptual.)\n\nPS: I am not affiliated with Fireworks in any way; I just thought that this was an interesting case study.", "url": "https://wpnews.pro/news/focusing-on-post-training", "canonical_source": "https://sebastianraschka.com/blog/2026/focusing-on-llm-post-training.html", "published_at": "2026-09-27 22:12:03+00:00", "updated_at": "2026-09-27 22:30:08.037024+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "ai-infrastructure"], "entities": ["Fireworks", "Ember-1", "Kimi K3", "Sebastian Raschka"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/focusing-on-post-training", "markdown": "https://wpnews.pro/news/focusing-on-post-training.md", "text": "https://wpnews.pro/news/focusing-on-post-training.txt", "jsonld": "https://wpnews.pro/news/focusing-on-post-training.jsonld"}}