{"slug": "training-a-small-model-to-run-a-house", "title": "Training a small model to run a house", "summary": "A developer fine-tuned Qwen3-1.7B, an open model small enough to run on a consumer GPU, to control a home through Home Assistant, raising its score on the Home Assistant community's voice-command test from 75.0% to 89.8%. The training used GRPO reinforcement learning, in which the model generated eight answers per command and a real Home Assistant checked which attempts left the house in the correct end state, running for seven hours on a single RTX 3060. The resulting model, hua-1.7b, was trained on 403 generated tests across 85 synthetic homes totaling 4,318 sentences, with the community's assist-mini, assist and questions sets held out for scoring.", "body_md": "# Training a small model to run a house\n\nI trained Qwen3-1.7B, an open model small enough to run on a consumer GPU, to control a house through Home Assistant. On the Home Assistant community's voice-command test it went from 75.0% to 89.8%.\n\nThe method is reinforcement learning with GRPO. The model never sees a correct answer. It tries each command eight times, a real Home Assistant checks which tries left the house right, and the model is nudged toward those. Training took seven hours on one RTX 3060, and the result is published as [hua-1.7b](https://huggingface.co/neelabhbuilds/hua-1.7b).\n\nThe rest of this post explains how it works, step by step, down to the math of a single training step.\n\n## 01 The tasks\n\n403\n\ntests across 85 pretend homes, 4,318 sentences in all.\n\nThe Home Assistant community has [a public test](https://github.com/allenporter/home-assistant-datasets) that scores models on this, with a leaderboard. Each task is a sentence and a synthetic home. The model answers with a tool call, Home Assistant runs it, and the task passes only if the whole home ends up the way it should.\n\nThe test has three parts. The assist-mini set is 196 spoken commands in small homes with few devices, the one the leaderboard is sorted by. The assist set is 460 commands in a medium-sized home, written as corner cases to trip models up. The questions set is 370 questions about the house, where the right response is an answer, not an action.\n\nEvery answer can be checked by running it: the house ends up right or wrong, 1 or 0. That makes this a good fit for reinforcement learning.\n\nI trained on my own tests, built by a generator in the same format as the community's. The community's three sets were kept out of training and used only to score the model.\n\n## 02 How it works\n\nThe model reads two things: Home Assistant's own prompt for the home, then the command. The prompt lists the rooms, each device with its name and its kind (light, lock, sensor), and the tools the model can use. A tool is an action Home Assistant can carry out, such as TurnOn or TurnOff. The model never changes the house itself. It writes a tool call: which tool, and which device to use it on. For \"Unlock the August Lock\", the right tool call is:\n\n```\nTurnOff\n  name: August Lock\n  kind: lock\nTurnOff    name: August Lock    kind: lock\n```\n\nThe name picks the device. The kind is there because the lock's sensor is also named August Lock. On a lock, TurnOff means unlock. Home Assistant runs the call, and the scorer compares the whole home with the expected end state.\n\nGRPO (Group Relative Policy Optimization) was introduced by DeepSeek in [the DeepSeekMath paper](https://arxiv.org/abs/2402.03300) (2024). The model gets one prompt and writes a group of answers to it. Each answer is scored. Answers that score above the group's average are made more likely, and answers below it less likely. Each answer is judged only against its own group: that is the \"group relative\" in the name.\n\nHere, the prompt is one command and the group is eight answers. Each answer scores 1 if the house ends up right, 0 if not. Compare each answer with the group's average. Answers above it get their probability nudged up; answers below it get nudged down. That is one step. Move to the next command and repeat, 1,600 steps in all.\n\n## The numbers behind one stepFive short parts, worked on one real sentence.OPENCLOSE\n\n### 1 The model answers eight times\n\nThe model gets \"Unlock the August Lock\" (the test from section 01) and answers it eight times. Each answer scores 1 if the house ends up right, 0 if not.\n\nThe trainer also gave a point for any valid tool call. In 1,587 of the 1,600 steps all eight answers earned it, so it changed nothing.\n\n### 2 Better answers are pushed up, worse ones down\n\nThe average score is 0.75. An answer above the average is made more likely. An answer below it is made less likely. The size of the push is the gap from the average, divided by the spread of the scores (0.463):\n\npush (score − average) / spread\n\nright answer (1 − 0.75) / 0.463 = +0.54\n\nwrong answer (0 − 0.75) / 0.463 = −1.62\n\n### 3 Equal scores push nothing\n\nThe pushes always add up to zero: six of +0.54 and two of −1.62. If all eight answers score the same, every push is zero and the model learns nothing. That is why the first run failed.\n\n### 4 Only the unsure parts move\n\nThe model writes an answer one token at a time (a token is a word or part of a word), and it picks each token with some probability. A push up raises the probability of every token in that answer. A push down lowers it.\n\nBefore training, the model was already over 99.9% sure of 29 of the 32 tokens in the right answer, so those barely change. The learning happens at the three places where it was unsure, shown before training and after all 1,600 steps:\n\n### 5 How the weights change\n\nThe loss is minus the push times the log probability of each token, averaged over all the tokens of the eight answers. Backpropagation gives the gradient, and each weight takes a small step against it:\n\n```\np.data += -lr * p.grad\n```\n\nwith lr = 5e-6, on the LoRA weights only. The real optimizer is AdamW, which sizes each weight's step separately; the idea is the same.\n\nEach step took 16.4 seconds. Almost all of it is the model: writing the eight answers, then updating its weights. The scorer checks 62 answers a second.\n\n## 03 What broke\n\n90%\n\nof the first run's steps had nothing to learn from: all eight answers scored the same.\n\nThe first run chose its training data by test, and that was the mistake. Before training, I asked the untrained model every sentence eight times and kept every test it passed only some of the time: 353 tests, 3,793 sentences. But a test has about eleven sentences, and each step uses only one. Inside a mixed test, most sentences still came back right all eight times or wrong all eight times. When the eight scores are equal, every answer sits at the average, so every nudge is zero and the step teaches nothing. I stopped the run at step 197.\n\n## 04 The fix\n\n616\n\nof the 4,318 sentences, kept for training.\n\nThe fix is to choose by sentence, not by test: keep only the sentences the untrained model got right on some tries and wrong on others. [The DAPO paper](https://arxiv.org/abs/2503.14476) (2025) does the same thing during training and calls it dynamic sampling.", "url": "https://wpnews.pro/news/training-a-small-model-to-run-a-house", "canonical_source": "https://www.neelabhbuilds.com/writing/training-a-small-model-to-run-a-house", "published_at": "2026-10-07 15:38:59+00:00", "updated_at": "2026-10-07 15:50:01.504916+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-agents"], "entities": ["Qwen3-1.7B", "Home Assistant", "hua-1.7b", "GRPO", "DeepSeek", "DeepSeekMath", "RTX 3060", "Hugging Face"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/training-a-small-model-to-run-a-house", "markdown": "https://wpnews.pro/news/training-a-small-model-to-run-a-house.md", "text": "https://wpnews.pro/news/training-a-small-model-to-run-a-house.txt", "jsonld": "https://wpnews.pro/news/training-a-small-model-to-run-a-house.jsonld"}}