cd /news/computer-vision/where-did-this-go-a-local-ai-workben… · home › topics › computer-vision › article
[ARTICLE · art-144991] src=dev.to ↗ pub= topic=computer-vision verified=true sentiment=· neutral

Where Did This Go? A Local AI Workbench Reset Assistant for a Coworker

A developer built the Shared Workbench Reset Assistant, a local prototype that uses the open-weight SmolVLM2 2.2B Instruct vision-language model in Q8 GGUF format via a llama.cpp server to compare a reference photo of a shared workbench with a post-use photo and draft an editable cleanup checklist. The Streamlit app runs on CPU in WSL Ubuntu 24.04 on an Intel Core Ultra X7 358H laptop with 32 GB of RAM and requires no hosted inference API key. The developer reports that SmolVLM2 500M produced repetition, incorrect directions, or incomplete guidance, and that moving to 2.2B did not solve the restoration task reliably; the prototype has been demoed on public Tabletop Tidying Up dataset images but not yet tested with the coworker it was built for.

by read5 min views2 publishedOct 4, 2026

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

I built Shared Workbench Reset Assistant for one coworker who has trouble remembering where objects originally belonged on our shared workbench.

The idea is simple: keep a reference photo of the intended layout, compare it with a photo taken after use, and prepare a short checklist for putting things back. The reference photo provides a visual memory of the workspace.

The prototype uses a local open-weight vision-language model to draft observations and a cleanup checklist. Both remain editable, and the user must review them against the photos before exporting a result. I have completed one demo workflow, but I have not yet tested it with my coworker or collected their feedback.

This demo is a screenshot walkthrough of the local prototype. The three screens show the photo inputs, reviewed observations, and final reviewed checklist.

The app runs locally in a browser. For a reproducible demonstration, I used public Tabletop Tidying Up (TTU) dataset images, not photographs of our company workbench.

The C1 sample contains a red cup, a beverage can, and a remote control. I assigned the first frame as the reference and the last frame as the current layout; these are prototype roles, not verified official tidy/untidy labels.

Left: the layout to restore. Right: the starting layout after use.

The workflow is:

Observations corrected with coding-assistant help and checked against the photos.

Final checklist corrected with coding-assistant help and confirmed against the reference photo.

Follow the project README to prepare the Python environment, download the Q8 model and projector, and build llama-server. Then keep two WSL terminals open. Replace /path/to/hacktoberfest-2026 with your checkout path; the virtual environment path below follows the README setup.

In terminal 1, start the model server:

cd /path/to/hacktoberfest-2026/dev-challenge/weekend-challenge
bash start_model_server.sh

After it reports that it is listening on port 8080, start the app in terminal 2:

source ~/.venvs/hacktoberfest-2026/bin/activate
cd /path/to/hacktoberfest-2026/dev-challenge/weekend-challenge
python -m streamlit run app.py \
  --server.address 127.0.0.1 \
  --server.port 8501 \
  --browser.gatherUsageStats false

Open http://localhost:8501 in your Windows browser and select the C1 sample. Review and correct both observations and the checklist against the photos before exporting. This address works on the machine running the app; it is not a public demo link. No hosted inference API key is needed.

Project repository · Weekend project and setup

The main files are app.py for the Streamlit interface, workbench.py for local inference requests and image preparation, and start_model_server.sh for starting llama.cpp. Model weights are downloaded separately and excluded from Git.

I used SmolVLM2 2.2B Instruct in Q8 GGUF format, including its vision projector, through a local llama.cpp server. Streamlit provides the interface. The app runs on the CPU in WSL Ubuntu 24.04 on my Intel Core Ultra X7 358H laptop with 32 GB of RAM. No GPU or model fine-tuning was used.

I first tried SmolVLM2 500M with two images, five ordered images, and a supplied object inventory. The results contained repetition, incorrect directions, or incomplete guidance. Moving to 2.2B did not solve the restoration task reliably.

In the completed browser run, the model took 7.01 seconds to describe the reference and 6.74 seconds to describe the current photo. It misreported the current cup's handle direction and the remote's face orientation. Even after I corrected the observations with my coding assistant's help, the 5.09-second checklist response mixed descriptions of both layouts instead of giving movement instructions. These timings describe one run, not a benchmark.

I therefore corrected the final instructions with my coding assistant and reviewed them against the images. The export preserves the original model responses, reviewed observations, final checklist, and confirmation flags so I can review exactly what required correction. Expected-result annotations were not provided to the model.

The current result demonstrates the review and export workflow. It does not establish reliable automatic spatial reasoning or a reduction in cleanup time.

TTU was created by Hogun Kee, Wooseok Oh, and Songhwai Oh (2024). The source, preserved upstream license notice, citation, and licensing caveat are documented in the project README. The sample images are third-party demo data.

Local inference lets me analyze the photos without sending them to a hosted model API. The app normalizes uploaded images in memory and sends them to a fixed localhost endpoint. It does not save uploaded photos to the repository, and the result export contains text rather than embedded photos.

Open weights and an editable inference stack also let me compare models and inspect failures on my own laptop. I could change prompts, retain unsuccessful outputs, and add review steps without depending on a provider's hosted model. There is no hosted inference API fee, although hardware and electricity still have costs. I have not yet verified the complete app with the internet disconnected.

This weekend is the first step in my month-long exploration of VLA and Physical AI, alongside my separate contributions to Intel's Physical AI Studio. This prototype uses a VLM and does not integrate Physical AI Studio or execute robot actions. My next step is coworker feedback and better visual grounding before considering a robot policy.

I used Codex to help implement the app, debug experiments, correct demo observations and instructions, and draft this post. The local SmolVLM2 responses are preserved separately from those assisted corrections. The prototype still needs a coworker trial and broader evaluation.

Experiment results and the reviewed JSON export are kept locally and excluded from Git. The screenshots show the reviewed workflow; they are not a saved agent-session transcript.

── more in #computer-vision 4 stories · sorted by recency
── more on @smolvlm2 2.2b instruct 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/where-did-this-go-a-…] indexed:0 read:5min 2026-10-04 · —