cd /news/robotics/direction-scale-decomposition-in-act… · home › topics › robotics › article
[ARTICLE · art-139433] src=arxiv.org ↗ pub= topic=robotics verified=true sentiment=↑ positive

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Direction-Scale Decomposition (DSD), an action representation that splits translation and rotation increments into direction and scale components before tokenization, improves average success rates for discrete-token vision-language-action models on LIBERO with both uniform binning (BIN) and the B-spline tokenizer BEAST, according to an arXiv paper (arXiv:2609.28865v1). On SimplerEnv, DSD-BIN outperformed BIN by 10.3 percentage points in overall success rate under mixed-dataset training, and real-robot experiments showed gains both with and without robotics pretraining. The authors position DSD as a way to mitigate performance degradation when training on large and diverse dataset mixtures.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.28865v1 Announce Type: new Abstract: Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

── more in #robotics 4 stories · sorted by recency
── more on @direction-scale decomposition 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/direction-scale-deco…] indexed:0 read:1min 2026-09-25 · —