Fine-tuning Qwen3.5-4B to replace Gemini Flash-Lite in a RAG pipeline Local Minutes fine-tuned Qwen3.5-4B with LoRA on 24,000 examples generated by Claude Opus 5.5, completing one training pass on a single rented H100 GPU in 1.6 hours for about $7, after Gemini 2.5 Flash-Lite's retirement and the nearly 4x higher cost of Gemini 3.5 Flash-Lite threatened the nonprofit L3C's RAG pipeline budget. On a 100-excerpt pilot, Opus 5.5 contexts scored 4.36 vs 2.69 out of 5 for usefulness, 2% vs 28% for unsupported claims, and 1% vs 44% for restated lines, while an earlier student trained on 23,800 Gemini-generated contexts inherited Gemini's flaws — roughly 20% unsupported claims and 41% restated lines. Pipeline costs Local Minutes finds meeting documents, converts each document to text, splits the text into excerpts, adds context to each excerpt, and embeds the result for search. Adding context is the most expensive of those steps, because it runs once for every excerpt. Gemini 2.5 Flash-Lite is great because it's very inexpensive, even compared to newer versions of Flash-Lite. Being forced to migrate to later versions of Flash-Lite would significantly impact the cost of the Local Minutes pipeline. Training Qwen3.5-4B With Gemini 2.5 Flash-Lite being retired, and its successor, Gemini 3.5 Flash-Lite being nearly 4x the cost, Local Minutes, being an L3C with limited resources, couldn't absorb that large of a cost increase, so we needed a cheaper model. Ideally an open source one that we can deploy as needed and don't need to worry about it being changed or deprecated in the future. Small open models are cheap to run, but out of the box they're bad at this job: an untrained 4B model made an unsupported claim in more than a third of its contexts see below . The standard way to make a small model good at one narrow job is to train it on examples: a capable teacher model does the job thousands of times, and a small student model is trained to produce the same answers from the same inputs. For the student I picked Qwen3.5-4B, a recent open model with 4 billion parameters: small enough to run on a single inexpensive GPU, and open, so no one can retire it. First try: Gemini as the teacher The easiest solution seemed to be to mimic the contexts generated by Gemini, which were already being saved with excerpts. Using the sample of 23,800 contexts generated by Gemini in production, I created a dataset for Qwen that reproduced Gemini's results, good and bad. When Claude Opus 5.5 was used as a blind evaluator, Gemini was given high marks for usefulness, but it had an issue making claims that weren't supported by the excerpt, happening in about 20% of contexts. A second check, with a classifier using TypeSafe's Jev model, found another habit: 41% of the lines in Gemini's contexts just restate the excerpt, which gives search nothing new, which was another bad pattern learned from Gemini 2.5 Flash-Lite. The same blind evaluator couldn't tell the student's contexts apart from Gemini's: a student copies its teacher's mistakes along with everything else. Second try: Opus 5.5 as the teacher The next test was to see if we could generate contexts using Opus 5.5 and train the 4B model on them. On a 100-excerpt pilot, Claude Opus 5.5 wrote far better contexts: usefulness 4.36 vs 2.69 out of 5, unsupported claims 2% vs 28%, restated lines 1% vs 44%. I sampled 25,206 excerpts from random documents across the past year, at random positions within each the first sample leaned too heavily on the openings of documents , and had Opus write context for each, and this became the new training data. Fine-tuning SFT A model's knowledge lives in its weights: large matrices of numbers that each layer multiplies its input by. Qwen3.5-4B has about 4 billion of them. Training means nudging those numbers so the model's output gets closer to the example answers. Nudging all 4 billion takes a lot of GPU memory. LoRA low-rank adaptation leaves the original weights frozen and learns a small correction for each matrix instead: two thin matrices whose product has the same shape as the original, added on top. Training went through the 24,000 examples once, on one rented H100 GPU, in 1.6 hours, for about $7. During training the model reads each request and predicts the teacher's generated context one token at a time. At each step it assigns a probability to every possible next token. The loss is small when it gave the token the teacher actually wrote a high probability, and large when the teacher's choice surprised it. After each batch of examples, the LoRA matrices are nudged in the direction that lowers the loss, so the model's predictions move closer to what the teacher would write. Only the context is scored, not the request, so the model learns to write context rather than to reproduce the documents it was given. Grading each generated excerpt context Verifying a given context is correct requires reading its source pages and its excerpt to see if the context covers everything it should in the excerpt, is accurate, and doesn't restate what is already evident in the excerpt, and there are thousands of them. So I used Claude Opus as a judge, set up to remove as much bias as possible: - Blind. Contexts from every system for the same excerpt are shown together, labeled only A, B, C. The judge doesn't know which model wrote which. - Rotated. Language-model judges tend to favor answers in certain positions, so every system appears in every position equally often. - Specific. The judge lists each unsupported claim and each missed reference it finds, then scores usefulness from 1 to 5. Having to list errors makes it look for them. Here's how the students did on 400 excerpts none of them had trained on: The student trained on Opus's contexts landed much closer to its teacher than to Gemini, with a third as many unsupported claims as production. The student trained on Gemini's contexts made fewer errors than the untrained model but was no more useful, and scored below Gemini: it learned Gemini's habits, including restating the excerpt, rather than better context. The student was still well short of its teacher. Bigger students trained on the same contexts did only a little better that side quest is at the end of Serving the model serving , so I kept working on the 4B. Testing search A judge's score is a proxy for retrieval quality or whether people can find what they're looking for. So I tested search directly. Most searches work fine with or without good context, so I focused on the cases context exists for: queries that can only find the right excerpt through its context. 1. Opus read 1,016 test excerpts and their production contexts, and flagged every context missing or misstating a fact a searcher would need. It found 235, about one in four. 2. For each, it wrote a search query that needs that fact, without seeing any context. 3. I indexed all 7,395 excerpts of those documents the way production does Cohere embeddings for semantic search, Pinecone's sparse model for keyword search , once with each system's contexts. 4. For each query, I checked whether the right excerpt came back first. With keyword search, the right excerpt came first 76% of the time with context from the RL-trained model that shipped, against 50% with production's. That comparison flatters the new contexts. The queries were chosen where production's contexts failed, so almost any other context starts ahead: the 4B's contexts before fine-tuning already reached 62%. The fair measure of what fine-tuning bought is the base model against the fine-tuned one: 62% to 76% for keyword search, and 60% to 69% for semantic search. Preference training DPO The above supervised fine-tuning amounts to imitation training and it taught the student to write correct context, but not to cover everything in the excerpt that needs explaining: it missed references nearly three times as often as its teacher. It was rewarded for correct context but not penalized for omitting context. Preference training teaches a model from pairs of answers to the same input: a chosen one and a rejected one. Imitation only ever shows the model good answers; preference training also shows it bad ones, so it learns what to avoid as well as what to copy. The method I used is DPO direct preference optimization . The rejected contexts came from the student itself, so the training targets the mistakes it actually makes i.e. on-policy . I had it write four contexts for each of 12,000 training excerpts at a temperature of 0.8. Temperature sets how much randomness goes into picking each next token. Production runs at 0: the model always takes its most likely token, so it writes the same context every time. At 0.8 it sometimes takes a less likely one, so the four versions differ and show the range of mistakes the student is capable of making, including ones it makes only occasionally that a single deterministic answer would hide. Much higher, and the samples would fill with mistakes it never makes in practice. I found the samples with mistakes and paired each with a corrected version. This is a real pair, with the names changed: In practice, training measures how likely the student is to write each context in a pair. If it already clearly favors the chosen one, the pair changes almost nothing; if it doesn't, the weights are adjusted so the chosen context becomes more likely and the rejected one less likely. What I learned, in order: - It over-corrects. The first round cut missed references from 21% to 10% but raised unsupported claims from 9.5% to 15%. The model learned to explain more of the excerpt but didn't prioritize accuracy. - It learns patterns you didn't intend. In one round the rejected context was always the longer one, so the model learned that shorter is safer and became terse. The next round balanced the two. - Pairs have to be the model's own mistakes. Reusing the 4B's pairs on a larger 9B model made the 9B worse usefulness 3.96 → 3.68 . Pairs built from the 9B's own contexts made it better 3.91 → 4.01 . My guess as to why is that the 9B rarely made the 4B's mistakes, so pushing them down taught it little about its own. The main thing left to learn between each pair was the 4B's phrasing, so the 9B's writing drifted, while the mistakes it did make never appeared in the pairs. - Would more imitation examples have done the same job? No. Training on 6,000, 12,000 and 24,000 of Opus's contexts kept lowering the loss, but judged quality barely moved 3.79 → 3.83 → 3.85 . The student sounds more like Opus without getting much better. It traded one mistake for another. Judged blind against the imitation model on 631 test excerpts, the best preference-trained version was exactly as useful, but made a different kind of error: Every pair rewarded explaining more, and none rewarded stopping where the source stops. So the model learned to fill gaps it couldn't actually fill. For this task that's the worst kind of error, because it's a confident statement about a real person. Reinforcement learning GRPO Preference training learns from pairs collected once. The model never finds out how the contexts it writes after an update would be graded. Reinforcement learning RL is done online so the model writes contexts, a grader scores them, and training shifts the model toward whatever scored well. Then it writes new contexts and the loop repeats, so the feedback always matches its current behavior. I used GRPO group relative policy optimization , the method behind DeepSeek-R1. Starting from the imitation model, each step went like this: 1. Take 32 training excerpts and have the model write 8 contexts for each, with some randomness again using temperature . 2. Have Claude Opus 5.5 grade every context. 3. Compare each context's score with the average of its group of 8. The LoRA weights are then nudged so that every token in an above-average context becomes a little more likely, and every token in a below-average one a little less likely, in proportion to how far from the average it scored. It's the same kind of update as fine-tuning, but on the model's own writing, with the grade setting the direction and size. 4. Add a penalty for drifting far from the imitation model, so it can't wander into nonsense. This method keeps the training trajectory on policy, which will result in more internally coherent, incremental modifications to the weights. A hundred steps meant 25,600 graded contexts. The grader In RL, the grader's score is the only thing that tells the model what a good context is, so it has to reward everything we want and penalize everything we don't. Preference training fell short here: every pair rewarded explaining more, and none penalized guessing. So this grader checks each claim a context makes, in two stages: - An answer key, once per excerpt. Opus lists what a reader of the excerpt alone would need explained, and for each item whether the source settles it, leaves it ambiguous, or doesn't say. - A grade, once per context. Opus sorts how the context handled each item on the key, and lists anything else it states that the source doesn't support. Opus only sorts; it never assigns points. The points live in code, so every context for the same excerpt is scored on the same terms: | What the context did | Points | |---|---| | Resolved a reference the source settles | +1 | | Got one wrong, or committed to an answer the source doesn't give | −3 | | Said the source doesn't say, when it doesn't | +0.5 | | Said it's unclear, when the source settles it | −0.5 | | Any other unsupported claim | −3 | | Wrong meeting or subject | −1.5 | | An item that only restates the excerpt | −0.25 | The −3 means a resolution is only worth stating if the model is more than 75% sure of it; below that, the expected cost of being wrong outweighs the point for being right. Before training on this grader, I checked it against the blind judge: on whether a context contains an unsupported claim, they agreed 95% of the time. Reinforcement learning moved the model toward its teacher on both kinds of mistake at once, which preference training never managed. Against the imitation model, unsupported claims fell from 11.9% to 7.9% and missed references from 15.2% to 9.5%, and usefulness rose from 4.08 to 4.19. Serving the model Serving a model means loading its weights onto a GPU and sending it requests. You can rent a GPU by the hour and keep it running, or use a serverless service that starts a GPU when requests arrive, bills by the second while it works, and shuts it down when idle. Local Minutes' excerpts come in nightly bursts, so serverless fits. The catch is a cold start: the first request after a quiet stretch waits a few minutes while the model loads. Two speeds matter. Latency is how long one excerpt takes: about 1.3 seconds for the 4B. Throughput is how many excerpts per second the GPU handles when it's busy, and that's what sets the price. Most of the GPU's work goes into reading the ~2,500-token request, not writing the ~100-token context, so a smaller model, which reads faster, costs less. The test excerpts understate the real workload: production requests are about twice as long, because long documents contribute many excerpts. So to measure cost and speed for real, I replayed a full production day of requests through the 4B on Runpod Serverless: I measured this with the first version that shipped. The RL-trained model is the same size on the same GPU. Its contexts are about 10% longer, which barely moves the cost, because reading the request is most of the work. Done The fine-tuned, RL-trained Qwen3.5-4B is now the Local Minutes contextualizer. It runs on Runpod Serverless, and each night the pipeline sends it the new excerpts together, so the GPU spins up, works through them, and shuts down. It gets the same request Gemini did and returns the same three fields, so nothing else in the pipeline had to change. The model is on Hugging Face as DuaneMB/local-minutes-contextualizer-4b https://huggingface.co/DuaneMB/local-minutes-contextualizer-4b , with a model card covering how to call it, how it was trained, and its limitations. What it cost | Item | Cost | |---|---| Against Gemini 3.5 Flash-Lite, the fine-tuned model saves about $920 per million excerpts, so the project pays for itself after about 1.2 million excerpts, and every million after that is 92% cheaper. A few other notes 1. Qwen3.5 knows 248,000 different tokens, and computing a score for every one of them at every position ran the GPU out of memory. Computing scores only where the context is fixed it, and made that part of training about 8× cheaper. 2. Small models sometimes repeat themselves forever. A mild penalty for repeating text 1.05 stopped it. 3. The first reinforcement-learning run barely moved: after 25 steps the model had drifted almost nowhere from where it started. Increasing the learning rate fixed it.