cd /news/large-language-models/what-we-learned-fine-tuning-our-own-… · home › topics › large-language-models › article
[ARTICLE · art-144200] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

What we learned fine-tuning our own coding model on a $100 budget

ElderAI, a small team building the ATLAS Code coding model for agent tools like Cline, Aider, Continue and Cursor, reported that none of its fine-tuning runs on roughly $100 of prepaid GPU compute has yet cleared its pre-registered quality gate, so its invite-only preview still serves the original starting checkpoint. The team detailed a hashed gate file written before metrics are computed, a midcheck early stop that cuts a failed run's cost to about $1.40 from about $3, and a shift to reporting edit_applies alongside exact-match scoring after byte-for-byte comparisons failed for both the base model and every fine-tune. Total GPU spend across several pilots, two midcheck-stopped runs and the latest gate run came to about $15.

by read3 min views1 publishedOct 2, 2026

We're a small team at ElderAI building ATLAS Code, a coding model for agent tools (Cline, Aider, Continue, Cursor and anything else that takes an OpenAI-compatible base URL). We do our own training runs on rented GPUs with about $100 of prepaid compute. Here's an honest account of the last two days.

The short version: none of our fine-tunes has cleared its quality gate yet. That's why the invite-only preview still runs the starting checkpoint we're trying to improve. We'll swap in a fine-tune once one earns it, and we'll say so when that happens. Here's what we learned while getting there.

On a small budget it's tempting to look at a run, find the metric that went up, and call it a win. We stopped letting ourselves do that.

Every run now starts with a small gate file written before any metric is computed. It's hashed, and the launcher refuses to start if the file changes. For our latest run it said:

The starting checkpoint's numbers are measured in the same job with the same harness, so we never compare against a number from a different setup.

The gate has already stopped us from fooling ourselves more than once. In the latest run, the fine-tune was ahead on the edit metric at the 40% checkpoint but failed the tool-call parse bar by about one call in a hundred. The gate said stop, so we stopped. (Funny detail: the starting checkpoint also sat right at that bar in the same job. That tells us the bar is very tight, but we don't get to loosen it after seeing the result.)

Our first instinct was to pour agent-style data (read a file, call edit_file, finish) into training to make the model better at tool use. Format did improve. But the general coding check slipped a little at the same time, enough to fail the "lose at most one problem" rule on several runs.

What helped:

The tradeoff is real, and on a small model you feel it fast. We don't see it as a bug to patch. It's the main thing we have to manage.

Our first edit metric was strict: after the model's tool call, the file had to match the real commit byte for byte. Almost everything failed, the starting checkpoint and every fine-tune alike, so we went through every failure by hand. We found:

spacing=3 to 6. No model can recover the old_str values that weren't unique in the file. So we changed what we measure instead of tuning the scorer until the numbers looked better:

edit_applies: Exact match stays the official number. The others are reported next to it so we can see why something failed.

Every run has a watchdog on the box that controls it:

A watchdog can also be too strict. Two of our recent attempts were stopped by the projected-cost rule while the job itself was healthy. ETA jitter early in training pushed the projection a few cents over the cap. We paid about $1.18 to learn that, then fixed the projection headroom.

Our biggest single saving is the midcheck early stop. A run that fails at 40% costs about $1.40 instead of about $3.

For scale: our recent attempts (several pilots, two full runs that stopped at midcheck, and the latest gate run) came to about $15 of GPU time in total. ATLAS Code is in invite-only preview behind an OpenAI-compatible /v1 API, with a Playground in your account. New accounts get 200 free credits, and plans are capped: when credits run out, requests stop, with no overage. We don't train on your prompts or code.

If you run Cline, Aider, Continue or a similar agent and want to tell us where it breaks, request access at elderai.cloud. We read everything people send.

── more in #large-language-models 4 stories · sorted by recency
── more on @elderai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-we-learned-fine…] indexed:0 read:3min 2026-10-02 · —