cd /news/computer-vision/small-yet-assistive-spatially-aware-… · home › topics › computer-vision › article
[ARTICLE · art-139427] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

Researchers released Smol-VL-BLV, a 500M-parameter vision-language model for blind and low-vision users that combines teacher-student distillation with Group Relative Policy Optimization (GRPO) using a composite BLV reward for directional language, metric distances, and hazard detection. The model improves the Spatial score by 19.3% and the Social score by 14.8% over baseline, raises OCR-Bench by 101.5% and TextVQA accuracy by 44.2%, and runs entirely on-device at approximately 450 MB on a mid-range Android smartphone via mixed-precision quantization. The model, dataset, and code are publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.28757v1 Announce Type: new Abstract: An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/

── more in #computer-vision 4 stories · sorted by recency
── more on @smol-vl-blv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/small-yet-assistive-…] indexed:0 read:1min 2026-09-25 · —