cd /news/artificial-intelligence/a-500-rl-fine-tune-of-a-9b-open-mode… · home topics artificial-intelligence article
[ARTICLE · art-76265] src=fermisense.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

A $500 reinforcement-learning fine-tune of a 9B open model outperformed frontier models on catalog review, according to a report detailing three real-world deployments. Bridgewater Associates' trained model makes roughly 30% fewer mistakes than the best frontier model at a fraction of the inference cost, Harvey's legal agent outperforms GPT-5.5 and Claude Opus 4.8 on its rubrics, and Intercom's Fin Apex resolves more issues than frontier models while being cheaper to run.

read2 min views1 publishedJul 28, 2026
A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
Image: source

Over the past two years this has hardened into a playbook: an open-source model, proprietary task data, and a reinforcement-learning stage against a scored version of the workflow. Below, we discuss three scenarios where this approach has been applied to real-world tasks.

Bridgewater Associates is one of the largest hedge funds in the world. Its analysts sift a constant stream of articles, filings, and emails, judging which documents are relevant to the firm's investment thesis and where boilerplate content begins. The catch is that relevant means relevant by Bridgewater's internal judgment, and no amount of prompting got frontier models to absorb that judgment reliably. So, the company decided to train an open-source model on labels from its own expert investors. The trained model [makes roughly 30% fewer mistakes than the best frontier model, at a fraction of the

inference cost](https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/). Harvey builds AI agents for law firms. Its hardest workloads are long-horizon: transaction due diligence and legal memo drafting, where the agent navigates large document sets, errors compound across steps, and even the best frontier models at maximum reasoning effort kept falling short of the quality bar. Harvey ran reinforcement learning on an open-weight model and got [a

legal agent that outperforms both GPT-5.5 and Claude Opus 4.8 on its rubrics](https://www.harvey.ai/blog/training-a-legal-agent-with-applied-compute). Intercom is a customer-service platform whose AI agent, Fin, resolves almost two million customer issues a week. At that volume the problem is unit economics: frontier per-call pricing adds up fast, and every point of resolution rate matters. So Intercom's AI group post-trained its own vertical support model, Fin Apex, on billions of customer-service interactions. Intercom reports that it [resolves

more issues than the best frontier models while being cheaper to run](https://www.intercom.com/blog/announcing-fin-apex-the-age-of-vertical-models-is-here/). The same shape repeats well beyond these three. The appendix collects eight more deployments, with what each model was trained to do and what changed once it shipped.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bridgewater associates 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-500-rl-fine-tune-o…] indexed:0 read:2min 2026-07-28 ·