This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
The plan was simple: before heading out for a picnic, you tell a robot what to pack, step away from the screen, and go. The robot does the rest.
I built that loop in simulation with an open vision-language-action (VLA) model, SmolVLA. It reads two camera images, the arm's state, and a sentence like "pick up the butter and place it in the basket", and outputs continuous arm actions. The picnic items are the ten groceries of the LIBERO-Object benchmark: alphabet soup, cream cheese, salad dressing, bbq sauce, ketchup, tomato sauce, butter, milk, chocolate pudding, and orange juice.
Then I tested the part that decides whether you actually get to leave: can you ask in your own words?
Each row is one item. Left to right: Original (the exact training sentence), then A, B, C, and D, four ways a person might say it instead (listed in the table under Step 2). Each tile shows the first successful episode out of ten, or episode 0 when all ten failed. For a closer look at the videos, see the full-resolution versions on GitHub.
The training sentence worked every time. The paraphrases split in two: A and D, which keep "put ... in the basket", worked about half the time (22/40 and 26/40), while B and C almost never worked (2/40 each).
First, the baseline: all ten items with their original instructions, 10 episodes each from slightly different starting layouts.
89 of 100 episodes succeeded. BBQ sauce was the weakest (6/10); four items were perfect (10/10). That matches a public run of the same checkpoint in huggingface/lerobot#4614 (42/50, 84%), so the setup reproduces what others see.
I took the four items that scored 10/10, so any drop points at the wording and not the grasp, and gave each one four paraphrases:
| Instruction | Success | |
|---|---|---|
| Original | pick up the {item} and place it in the basket | 40/40 |
| A | grab the {item} and put it in the basket | 22/40 |
| B | {item} in the basket | 2/40 |
| C | pack the {item} for our picnic | 2/40 |
| D | could you put the {item} in the picnic basket? | 26/40 |
Across all four paraphrases: 52/160 (33%). Every failure ran to the 280-step limit without placing the item.
Three things stood out:
If only the exact template works, the person stays at the screen rephrasing instead of heading out. For a Touch Grass project, that is the result that matters.
These are hypotheses; the runs cannot separate them.
train_expert_only: False, so the SmolVLM2 language layers were trained on those same template sentences and may have drifted from general English.lerobot/pi05_libero_finetuned would test this.
Everything runs on one laptop CPU (Intel Core Ultra X7 358H, Windows 11, WSL Ubuntu 24.04). No GPU, no cloud.
lerobot/smolvla_libero, SmolVLA fine-tuned on LIBERO. lerobot-eval, wrapped to time every model call. The paraphrase runs use a small rollout loop that swaps in a new instruction while keeping the same initial states and seeds.
Two flags matter (excerpt of the full command): one keeps the run from failing, the other keeps the success rate from dropping.
python baseline/benchmark_eval.py \
--policy.path=lerobot/smolvla_libero \
--policy.device=cpu \
--policy.n_action_steps=10 \
--env.type=libero \
--env.task=libero_object \
--rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'
--rename_map: the checkpoint expects cameras named camera1 and camera2, but LeRobot 0.6.1 does not apply the saved mapping automatically (--policy.n_action_steps=10: the checkpoint ships with 50, which replays a whole action chunk before looking again. In lerobot#4614 that cost 20 points on LIBERO-Object (84% to 64%).
The simulator s while the model thinks. A real arm would not.
Each SmolVLA call plans 0.5 s of motion but takes 1.15 s on this CPU. On a real arm, the ketchup episode above would stop and go: 8.8 s of motion becomes about 29.5 s. Success rates are unaffected because the simulator waits, but real-time control would need a GPU or asynchronous inference. I did not test either.
A month-long exploration of open-weight AI, VLA, and Physical AI.
dev-challenge/weekend-challenge/: Weekend Challenge project workspace devrelay/: DevRelay usage notes and reviewed agent session transcriptsmodels/: Shared local model weights for all projects; excluded from Git
| Component | Configuration |
|---|---|
| Host | Windows |
| CPU | Intel Core Ultra X7 358H; 16 logical CPUs visible in WSL |
| GPU | Intel Arc B390 GPU; Windows driver 32.0.101.8356 |
| RAM | 32 GB |
| Runtime | WSL with Ubuntu 24.04 LTS |
| Shell for project commands | Bash in Ubuntu |
| Python | 3.12.3 |
The project lives in dev-challenge/week-1, with run scripts and a results page for the baseline and for the paraphrase experiment, plus setup notes for WSL.
This project would not exist with a closed model, because the interesting part was looking inside.
The honest tradeoff: CPU inference is too slow to drive a real arm, and a small open model fine-tuned on one sentence per task does not understand casual English yet. But I can see exactly why, and that is what open gives me.
I built this with Claude Code over four days. The session below is a curated slice: the decisions, the setup bug, the results, and the failure analysis.
Note: I worked with the agent in Korean. This is a condensed English translation of the key moments, not a verbatim transcript.
Disclosure: I designed the experiments, ran every episode, and reviewed every result. An AI coding agent helped write scripts, draft the paraphrased instructions (which I reviewed), and draft this post.