{"slug": "2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt", "title": "2/184 to 170/184: one epoch on two R9700s took a 14B coder from useless to 5 points behind gpt-5.6-sol", "summary": "A fine-tuned Qwen2.5-Coder-14B-Instruct model compiled 170 of 184 MQL5 test prompts after one epoch of training on the author's dataset, up from 2 of 184 for the untouched base model, according to a published benchmark card. The same 184 prompts scored 179 of 184 for gpt-5.6-sol, leaving the local 14B about 5 points behind the frontier API model. Training ran under ROCm on three AMD Radeon AI PRO R9700 cards across two machines totaling 96GB of VRAM, and the release includes the 184 prompts, scoring rules, per-item results, and a verifier script, but not the training data or tuned weights.", "body_md": "Been a while. Last time I posted the dataset had passed 300k\n\nentries and I had one 14B fine tune with 81% structural\n\ncorrectness. Since then I did the thing I should have done from\n\nthe start, I built a proper benchmark and published it so people\n\ncan check my numbers instead of taking my word for it.\n\nWhat got published:\n\n184 test prompts for MQL5 (the language for algo trading bots),\n\nthe scoring rules, per item results for every model, and a\n\nverifier script that checks the release hashes and recomputes the\n\nheadline numbers from the rows. A pass is zero compiler errors\n\nand a built binary. Thats it. Compiling is a low bar and the card\n\nsays so, it does not mean the bot trades well.\n\nThe numbers:\n\nBase Qwen2.5-Coder-14B-Instruct, untouched: 2 of 184 compiled.\n\nSame model after one epoch on my data: 170 of 184.\n\ngpt-5.6-sol on the same 184 prompts: 179 of 184.\n\nSo the local 14B went from useless to about 5 points behind a\n\nfrontier API model on my niche. One epoch. I am not going to\n\npretend thats the finish line but it is the first number I have\n\nthat somebody else can reproduce, and that matters more to me\n\nthan the number itself.\n\nHardware, since this is L1T:\n\nEverything is still AMD and still local, but the layout changed\n\nsince April. The main rig is now a 7950X3D with 64GB and one\n\nRadeon AI PRO R9700 in it, the 9800X3D and the 9070XT are out.\n\nThat box is where both local arms of the benchmark were run, base\n\nmodel and tuned model, same card, same settings. The\n\nThreadripper 2970WX has two R9700s now and does the training\n\nunder ROCm. So three R9700s total across two machines, 96GB of\n\nVRAM, no CUDA anywhere. I fought ROCm plenty (some of you saw the\n\nWindows thread) but it does the job and the numbers came out of\n\nit.\n\nLinks:\n\nHugging Face:\n\nGitHub mirror:\n\nWhat is not in the release is the training data and the tuned\n\nweights. What is in the release is every prompt and every result\n\nrow, so you can run your own model against the same 184 and\n\ncompare.\n\nCaveat thats also on the card: the test prompts come from the\n\nsame generator family as the training data, so this measures how\n\ngood the model got at my kind of spec, not how it does on random\n\nhuman written requests. I have a bigger private holdout in the\n\nworks and a version 1.1 of the card coming with those results.\n\nStill grinding. If you are on AMD and want to compare notes on\n\nROCm training or eval, ask away.\n\nMQL5 is a trademark of its owner. Not affiliated, independent\n\nproject.", "url": "https://wpnews.pro/news/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt", "canonical_source": "https://forum.level1techs.com/t/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-points-behind-gpt-5-6-sol/255875#post_1", "published_at": "2026-09-12 12:43:09+00:00", "updated_at": "2026-09-12 13:10:27.290629+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools", "developer-tools"], "entities": ["Qwen2.5-Coder-14B-Instruct", "gpt-5.6-sol", "MQL5", "AMD Radeon AI PRO R9700", "ROCm", "Hugging Face", "GitHub", "Threadripper 2970WX"], "alternates": {"html": "https://wpnews.pro/news/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt", "markdown": "https://wpnews.pro/news/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt.md", "text": "https://wpnews.pro/news/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt.txt", "jsonld": "https://wpnews.pro/news/2-184-to-170-184-one-epoch-on-two-r9700s-took-a-14b-coder-from-useless-to-5-gpt.jsonld"}}