Show HN: Distill and serve small models with frontier quality for half the cost Experiential Labs launched wmo serve, an open-source tool that routes repetitive agent tasks to distilled smaller models, achieving frontier quality at 40%+ lower cost. The tool uses an OpenAI-compatible endpoint and continuously improves models through distillation from open-source models and token compaction. Hi HN, we built world-model-optimizer, an open-source tool for continually improve models specialized to agents. Today we are launching wmo serve , a tool to route repetitive tasks to distilled smaller models. Agent traces you already capture are opportunities to get signal on how to make your model cheaper, faster, better. We continuously improve - your specialized model through distillation from open source models - model routing to frontier + custom models - token compaction to remove noise and save tokens Demo: https://www.youtube.com/watch?v=2 m4Ze6mdko https://www.youtube.com/watch?v=2 m4Ze6mdko Pass in traces and an OpenRouter key, and wmo starts a local OpenAI-compatible endpoint to run with your model at a lower cost with equivalent quality. Behind the scenes a router decides which tasks should go to the frontier versus your model. Tinker continually trains as new traces arrive. We also offer a hosted solution for anyone that just wants a frontier quality endpoint with self-improvement over time at a 40%+ lower cost. Sign up for the waitlist at https://experientiallabs.ai https://experientiallabs.ai and give the repo a star Comments URL: https://news.ycombinator.com/item?id=49063454 https://news.ycombinator.com/item?id=49063454 Points: 3 Comments: 0