{"slug": "a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review", "title": "A $500 RL fine-tune of a 9B open model beat frontier models on catalog review", "summary": "A $500 reinforcement-learning fine-tune of a 9B open model outperformed frontier models on catalog review, according to a report detailing three real-world deployments. Bridgewater Associates' trained model makes roughly 30% fewer mistakes than the best frontier model at a fraction of the inference cost, Harvey's legal agent outperforms GPT-5.5 and Claude Opus 4.8 on its rubrics, and Intercom's Fin Apex resolves more issues than frontier models while being cheaper to run.", "body_md": "Over the past two years this has hardened into a playbook: an open-source model,\nproprietary task data, and a reinforcement-learning stage against a scored version of\nthe workflow. Below, we discuss three scenarios where this approach has been applied to real-world tasks.\n\n**Bridgewater Associates** is one of the largest hedge funds in the world. Its\nanalysts sift a constant stream of articles, filings, and emails, judging which\ndocuments are relevant to the firm's investment thesis and where boilerplate\ncontent begins. The catch is that *relevant* means relevant by Bridgewater's internal\njudgment, and no amount of prompting got frontier models to absorb that judgment\nreliably. So, the company decided to train an open-source\nmodel on labels from its own expert investors. The trained model\n[makes\nroughly 30% fewer mistakes than the best frontier model, at a fraction of the\ninference cost](https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/).\n\n**Harvey** builds AI agents for law firms. Its hardest workloads are\nlong-horizon: transaction due diligence and legal memo drafting, where the agent\nnavigates large document sets, errors compound across steps, and even the best\nfrontier models at maximum reasoning effort kept falling short of the quality bar.\nHarvey ran reinforcement learning on an open-weight model and got\n[a\nlegal agent that outperforms both GPT-5.5 and Claude Opus 4.8 on its rubrics](https://www.harvey.ai/blog/training-a-legal-agent-with-applied-compute).\n\n**Intercom** is a customer-service platform whose AI agent, Fin, resolves\nalmost two million customer issues a week. At that volume the problem is unit\neconomics: frontier per-call pricing adds up fast, and every point of resolution rate\nmatters. So Intercom's AI group post-trained its own vertical support model, Fin Apex,\non billions of customer-service interactions. Intercom reports that it\n[resolves\nmore issues than the best frontier models while being cheaper to run](https://www.intercom.com/blog/announcing-fin-apex-the-age-of-vertical-models-is-here/).\n\nThe same shape repeats well beyond these three. The\n[appendix](#appendix) collects eight more deployments, with what each\nmodel was trained to do and what changed once it shipped.", "url": "https://wpnews.pro/news/a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review", "canonical_source": "https://fermisense.com/when-machines-take-the-wheel/", "published_at": "2026-07-28 02:18:53+00:00", "updated_at": "2026-07-28 02:52:33.285964+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-products"], "entities": ["Bridgewater Associates", "Harvey", "Intercom", "Fin Apex", "GPT-5.5", "Claude Opus 4.8"], "alternates": {"html": "https://wpnews.pro/news/a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review", "markdown": "https://wpnews.pro/news/a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review.md", "text": "https://wpnews.pro/news/a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review.txt", "jsonld": "https://wpnews.pro/news/a-500-rl-fine-tune-of-a-9b-open-model-beat-frontier-models-on-catalog-review.jsonld"}}