cd /news/robotics/can-i-ask-a-robot-to-pack-my-picnic-… · home › topics › robotics › article
[ARTICLE · art-149220] src=dev.to ↗ pub= topic=robotics verified=true sentiment=· neutral

Can I Ask a Robot to Pack My Picnic in My Own Words?

A developer built a simulation loop that lets a user instruct a SmolVLA vision-language-action model to pack picnic items in natural language, then measured how well the model handles paraphrased commands. The model succeeded on 89 of 100 baseline episodes using the exact training template, but only 52 of 160 (33%) episodes when the same instructions were rephrased, with variants dropping "put ... in the basket" falling to 2/40. All runs were done on a single laptop CPU with no GPU or cloud.

by read5 min views1 publishedOct 11, 2026

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

The plan was simple: before heading out for a picnic, you tell a robot what to pack, step away from the screen, and go. The robot does the rest.

I built that loop in simulation with an open vision-language-action (VLA) model, SmolVLA. It reads two camera images, the arm's state, and a sentence like "pick up the butter and place it in the basket", and outputs continuous arm actions. The picnic items are the ten groceries of the LIBERO-Object benchmark: alphabet soup, cream cheese, salad dressing, bbq sauce, ketchup, tomato sauce, butter, milk, chocolate pudding, and orange juice.

Then I tested the part that decides whether you actually get to leave: can you ask in your own words?

Each row is one item. Left to right: Original (the exact training sentence), then A, B, C, and D, four ways a person might say it instead (listed in the table under Step 2). Each tile shows the first successful episode out of ten, or episode 0 when all ten failed. For a closer look at the videos, see the full-resolution versions on GitHub.

The training sentence worked every time. The paraphrases split in two: A and D, which keep "put ... in the basket", worked about half the time (22/40 and 26/40), while B and C almost never worked (2/40 each).

First, the baseline: all ten items with their original instructions, 10 episodes each from slightly different starting layouts.

89 of 100 episodes succeeded. BBQ sauce was the weakest (6/10); four items were perfect (10/10). That matches a public run of the same checkpoint in huggingface/lerobot#4614 (42/50, 84%), so the setup reproduces what others see.

I took the four items that scored 10/10, so any drop points at the wording and not the grasp, and gave each one four paraphrases:

Instruction Success
Original pick up the {item} and place it in the basket 40/40
A grab the {item} and put it in the basket 22/40
B {item} in the basket 2/40
C pack the {item} for our picnic 2/40
D could you put the {item} in the picnic basket? 26/40

Across all four paraphrases: 52/160 (33%). Every failure ran to the 280-step limit without placing the item.

Three things stood out:

If only the exact template works, the person stays at the screen rephrasing instead of heading out. For a Touch Grass project, that is the result that matters.

These are hypotheses; the runs cannot separate them.

train_expert_only: False, so the SmolVLM2 language layers were trained on those same template sentences and may have drifted from general English.lerobot/pi05_libero_finetuned would test this. Everything runs on one laptop CPU (Intel Core Ultra X7 358H, Windows 11, WSL Ubuntu 24.04). No GPU, no cloud.

lerobot/smolvla_libero, SmolVLA fine-tuned on LIBERO. lerobot-eval, wrapped to time every model call. The paraphrase runs use a small rollout loop that swaps in a new instruction while keeping the same initial states and seeds. Two flags matter (excerpt of the full command): one keeps the run from failing, the other keeps the success rate from dropping.

python baseline/benchmark_eval.py \
  --policy.path=lerobot/smolvla_libero \
  --policy.device=cpu \
  --policy.n_action_steps=10 \
  --env.type=libero \
  --env.task=libero_object \
  --rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'

--rename_map: the checkpoint expects cameras named camera1 and camera2, but LeRobot 0.6.1 does not apply the saved mapping automatically (--policy.n_action_steps=10: the checkpoint ships with 50, which replays a whole action chunk before looking again. In lerobot#4614 that cost 20 points on LIBERO-Object (84% to 64%). The simulator s while the model thinks. A real arm would not.

Each SmolVLA call plans 0.5 s of motion but takes 1.15 s on this CPU. On a real arm, the ketchup episode above would stop and go: 8.8 s of motion becomes about 29.5 s. Success rates are unaffected because the simulator waits, but real-time control would need a GPU or asynchronous inference. I did not test either.

A month-long exploration of open-weight AI, VLA, and Physical AI.

dev-challenge/weekend-challenge/: Weekend Challenge project workspace devrelay/: DevRelay usage notes and reviewed agent session transcriptsmodels/: Shared local model weights for all projects; excluded from Git

Component Configuration
Host Windows
CPU Intel Core Ultra X7 358H; 16 logical CPUs visible in WSL
GPU Intel Arc B390 GPU; Windows driver 32.0.101.8356
RAM 32 GB
Runtime WSL with Ubuntu 24.04 LTS
Shell for project commands Bash in Ubuntu
Python 3.12.3

The project lives in dev-challenge/week-1, with run scripts and a results page for the baseline and for the paraphrase experiment, plus setup notes for WSL.

This project would not exist with a closed model, because the interesting part was looking inside.

The honest tradeoff: CPU inference is too slow to drive a real arm, and a small open model fine-tuned on one sentence per task does not understand casual English yet. But I can see exactly why, and that is what open gives me.

I built this with Claude Code over four days. The session below is a curated slice: the decisions, the setup bug, the results, and the failure analysis.

Note: I worked with the agent in Korean. This is a condensed English translation of the key moments, not a verbatim transcript.

Disclosure: I designed the experiments, ran every episode, and reviewed every result. An AI coding agent helped write scripts, draft the paraphrased instructions (which I reviewed), and draft this post.

── more in #robotics 4 stories · sorted by recency
── more on @smolvla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-i-ask-a-robot-to…] indexed:0 read:5min 2026-10-11 · —