# Can I Ask a Robot to Pack My Picnic in My Own Words?

> Source: <https://dev.to/johyeongseob/can-i-ask-a-robot-to-pack-my-picnic-in-my-own-words-40f2>
> Published: 2026-10-11 16:49:25+00:00

*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*

The plan was simple: before heading out for a picnic, you tell a robot what to pack, step away from the screen, and go. The robot does the rest.

I built that loop in simulation with an open vision-language-action (VLA) model, [SmolVLA](https://huggingface.co/lerobot/smolvla_libero). It reads two camera images, the arm's state, and a sentence like "pick up the butter and place it in the basket", and outputs continuous arm actions. The picnic items are the ten groceries of the LIBERO-Object benchmark: alphabet soup, cream cheese, salad dressing, bbq sauce, ketchup, tomato sauce, butter, milk, chocolate pudding, and orange juice.

Then I tested the part that decides whether you actually get to leave: **can you ask in your own words?**

Each row is one item. Left to right: Original (the exact training sentence), then A, B, C, and D, four ways a person might say it instead (listed in the table under Step 2). Each tile shows the first successful episode out of ten, or episode 0 when all ten failed. For a closer look at the videos, see the full-resolution versions on [GitHub](https://github.com/johyeongseob/hacktoberfest-2026/tree/main/dev-challenge/week-1/paraphrases).

The training sentence worked every time. The paraphrases split in two: A and D, which keep "put ... in the basket", worked about half the time (22/40 and 26/40), while B and C almost never worked (2/40 each).

First, the baseline: all ten items with their original instructions, 10 episodes each from slightly different starting layouts.

**89 of 100 episodes succeeded.** BBQ sauce was the weakest (6/10); four items were perfect (10/10). That matches a public run of the same checkpoint in [huggingface/lerobot#4614](https://github.com/huggingface/lerobot/issues/4614) (42/50, 84%), so the setup reproduces what others see.

I took the four items that scored 10/10, so any drop points at the wording and not the grasp, and gave each one four paraphrases:

|  | Instruction | Success | 
|---|---|---|
| Original | pick up the {item} and place it in the basket | **40/40** | 
| A | grab the {item} and put it in the basket | 22/40 | 
| B | {item} in the basket | **2/40** | 
| C | pack the {item} for our picnic | **2/40** | 
| D | could you put the {item} in the picnic basket? | 26/40 | 

Across all four paraphrases: **52/160 (33%)**. Every failure ran to the 280-step limit without placing the item.

Three things stood out:

If only the exact template works, the person stays at the screen rephrasing instead of heading out. For a Touch Grass project, that is the result that matters.

These are hypotheses; the runs cannot separate them.

`train_expert_only: False`, so the SmolVLM2 language layers were trained on those same template sentences and may have drifted from general English.`lerobot/pi05_libero_finetuned` would test this.
Everything runs on one laptop CPU (Intel Core Ultra X7 358H, Windows 11, WSL Ubuntu 24.04). No GPU, no cloud.

`lerobot/smolvla_libero`, SmolVLA fine-tuned on LIBERO.` lerobot-eval`, wrapped to time every model call. The paraphrase runs use a small rollout loop that swaps in a new instruction while keeping the same initial states and seeds.
Two flags matter (excerpt of the full command): one keeps the run from failing, the other keeps the success rate from dropping.

```
python baseline/benchmark_eval.py \
  --policy.path=lerobot/smolvla_libero \
  --policy.device=cpu \
  --policy.n_action_steps=10 \
  --env.type=libero \
  --env.task=libero_object \
  --rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'
```

`--rename_map`: the checkpoint expects cameras named `camera1` and `camera2`, but LeRobot 0.6.1 does not apply the saved mapping automatically (`--policy.n_action_steps=10`: the checkpoint ships with 50, which replays a whole action chunk before looking again. In lerobot#4614 that cost 20 points on LIBERO-Object (84% to 64%).
The simulator pauses while the model thinks. A real arm would not.

Each SmolVLA call plans 0.5 s of motion but takes 1.15 s on this CPU. On a real arm, the ketchup episode above would stop and go: 8.8 s of motion becomes about 29.5 s. Success rates are unaffected because the simulator waits, but real-time control would need a GPU or asynchronous inference. I did not test either.

A month-long exploration of open-weight AI, VLA, and Physical AI.

`dev-challenge/weekend-challenge/`: Weekend Challenge project workspace` devrelay/`: DevRelay usage notes and reviewed agent session transcripts`models/`: Shared local model weights for all projects; excluded from Git
| Component | Configuration | 
|---|---|
| Host | Windows | 
| CPU | Intel Core Ultra X7 358H; 16 logical CPUs visible in WSL | 
| GPU | Intel Arc B390 GPU; Windows driver 32.0.101.8356 | 
| RAM | 32 GB | 
| Runtime | WSL with Ubuntu 24.04 LTS | 
| Shell for project commands | Bash in Ubuntu | 
| Python | 3.12.3 | 

The project lives in [`dev-challenge/week-1`](https://github.com/johyeongseob/hacktoberfest-2026/tree/main/dev-challenge/week-1), with run scripts and a results page for the baseline and for the paraphrase experiment, plus setup notes for WSL.

This project would not exist with a closed model, because the interesting part was looking inside.

The honest tradeoff: CPU inference is too slow to drive a real arm, and a small open model fine-tuned on one sentence per task does not understand casual English yet. But I can see exactly why, and that is what open gives me.

I built this with Claude Code over four days. The session below is a curated slice: the decisions, the setup bug, the results, and the failure analysis.

**Note:** I worked with the agent in Korean. This is a condensed English translation of the key moments, not a verbatim transcript.

**Disclosure:** I designed the experiments, ran every episode, and reviewed every result. An AI coding agent helped write scripts, draft the paraphrased instructions (which I reviewed), and draft this post.
