cd /news/large-language-models/i-fine-tuned-qwen3-14b-for-my-ai-com… · home › topics › large-language-models › article
[ARTICLE · art-144369] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

I fine-tuned Qwen3-14B for my AI companion app for $1.30. The hard part wasn't the model

Richael (dh8116), a Year 11 student in Auckland, built Soulor AI, an AI companion app with persistent memory and a five-perspective Simulation Mode, on a LoRA fine-tune of Qwen3-14B served with vLLM on Modal. The fine-tune cost about $1.30 and took half an hour, scoring 15/15 on companion chat and 15/15 on feed comments at production prompt size, after he rewrote the training corpus so every example was written inside the JSON envelope the app parses. Serving on an H100 cut reply latency from 7.9s to 2.0s, and a warm-container probe in the gateway lets the fine-tune serve conversations when a container is already up while falling back to hosted models during the 108s cold start.

by read3 min views1 publishedOct 3, 2026

Soulor AI is an AI companion app with persistent memory and five relationship stages, from Stranger to Soulmate, plus a Simulation Mode that gives a panel of five AI perspectives on any situation. I'm Richael (dh8116), a Year 11 student in Auckland, and I built it. This post is about the model underneath: a LoRA fine-tune of Qwen3-14B, served with vLLM on Modal, and the four things that mattered more than I expected.

The app doesn't want free text from the model. Every reply comes back in a JSON envelope the frontend parses: the message, plus fields the app acts on. My first adapter learned the persona perfectly and couldn't produce that envelope at all. It sounded right, and the app couldn't use any of it.

The fix was the training data, not the hyperparameters. I rewrote the corpus so every example was written inside the envelope, the same shape production asks for. The retrained adapter scores 15/15 on companion chat and 15/15 on feed comments at production prompt size.

Training cost about $1.30 and took half an hour.

Lesson: if your app parses the model's output, the output format is part of the behaviour you're fine-tuning, so train on it.

Replies took 7.9s. I spent a while on serving flags before measuring properly. Decoding a 14B model is bound by memory bandwidth, so each token means reading the weights again, and no flag changes how fast the card can read them. Moving to an H100 took it to 2.0s.

The hosted Qwen fallback had a different problem. With thinking enabled it took 14.6s, and the reasoning was eating the token budget and truncating replies. Turning enable_thinking off gave 5.7s and complete answers.

The Modal endpoint scales to zero, which is what keeps it affordable. A cold start takes 108s, though, and nobody waits 108s for a chat reply.

So the gateway asks one question before every request: is a container warm? If yes, the fine-tune goes first in the provider chain. If not, it skips straight to the hosted models. The probe doubles as the warm-up, so the first message of a conversation is answered by a hosted model and the rest can go to the fine-tune. The voice can shift once, early, but I don't pay for an idle GPU between conversations.

Signups leaving unreachable accounts, cleared chats coming back, companions forgetting the last few turns, and chat hanging for a minute all looked unrelated. They came from just two patterns:

await, with its error swallowed. Both are now fixed where they can't be forgotten: writes are awaited, and every provider call has a timeout.

The same engine runs the companion and the panel. In Simulation Mode, one situation goes to five analysts: an Optimist, a Cynic, a Mentor, a Status Observer and a Gossiper. Because it's the same app, the panel already knows the context you gave your companion. You can rehearse a hard conversation before having it for real, then switch back mid-conversation without starting over.

Soulor is in early access and free to start: https://soulor-ai.vercel.app/ I'm happy to go deeper on any of this in the comments: the envelope corpus, the warm gate, or the eval harness. More of what I build, including a weekly Triton kernel series benchmarked against PyTorch, is at https://dh8116.github.io/

── more in #large-language-models 4 stories · sorted by recency
── more on @soulor ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-fine-tuned-qwen3-1…] indexed:0 read:3min 2026-10-03 · —