cd /news/robotics/cap-x-lms-first-physical-exam · home › topics › robotics › article
[ARTICLE · art-145701] src=capgym.github.io ↗ pub= topic=robotics verified=true sentiment=↑ positive

CAP-X: LMs' First Physical Exam

A new benchmark called CaP-Bench, integrated with LIBERO-PRO, Robosuite and BEHAVIOR, shows frontier large language models can generate executable robot control code zero-shot at over 30% average success, leaving a 56-point gap to human performance. On LIBERO-PRO's 30 perturbed manipulation tasks, state-of-the-art Vision-Language-Action models OpenVLA and π0 scored 0% while the best VLA, π0.5, reached 13%, and the training-free coding agent CaP-Agent0 reached 18%. Using CaP-RL reinforcement learning on the coding agent, a 7B Qwen 2.5 Coder model rose from 20% to 72% average success in simulation after 50 training iterations and transferred to a real Franka Emika robot, reaching 84% on cube lifting and 76% on cube stacking.

read2 min views1 publishedOct 5, 2026

Today's off-the-shelf LMs have incredible generalization, reasoning, and planning capabilities. Agentic harnesses in CaP-Agent0 unleash their potential in the physical world.

Click on each task to see the agent in action.

CaP-Bench provides the first comprehensive benchmark for evaluating how well large language model agents can write code to control robots. Integrated with hundreds of manipulation tasks across multiple robot learning benchmarks (LIBERO-PRO, Robosuite, BEHAVIOR), CaP-Bench tests both LLM and VLM models on their ability to generate executable robot control policies from natural language instructions.

Key Findings

1
Frontier models achieve meaningful zero-shot success on robotic manipulation

            Without any task-specific training, today's best frontier models can directly generate executable robot control
            code and achieve over 30% average success — a sharp contrast to the prior belief that only specially trained
            models (VLAs) can perform manipulation. Yet a 56-point gap to human performance remains, marking this as
            one of AI's most important open challenges.
          

 chart...

2
Training-free CaP-Agent0 outperforms state-of-the-art VLAs on perturbed tasks

            On LIBERO-PRO — 30 manipulation tasks with position and instruction perturbations — state-of-the-art
            Vision-Language-Action models (OpenVLA, π0) score 0% across the board.
            Even the best VLA (π0.5) reaches only 13% average success. CaP-Agent0, a training-free
            coding agent, achieves 18% without any task-specific training, demonstrating that
            code-generation agents generalize where end-to-end learned policies break down.
          

 chart...

3
CaP-RL: Post-training on code dramatically boosts robot performance — and transfers sim-to-real

            Using CaP-RL, we apply reinforcement learning with environment rewards directly on the coding agent.
            A 7B model (Qwen 2.5 Coder) jumps from 20% to 72% average success in simulation after
            just 50 training iterations. The learned policies transfer to a real Franka Emika robot with
            minimal sim-to-real gap — reaching 84% on cube lifting and 76% on cube stacking,
            approaching human expert performance.
          

 chart...

4
Higher abstraction boosts all models — and dramatically closes the gap for smaller ones

            As API abstraction increases from raw primitives (S4) to high-level pick-and-place (S1),
            all models improve substantially — but the gains are most pronounced for
            weaker and open-source models, whose compilation rates collapse at low abstraction levels.
            This suggests a promising path: pair a lightweight LM for high-level planning
            with a visual-motor policy (e.g., a VLA) that handles low-level control, letting even
            smaller models achieve strong task performance through the right division of labor.
          

 chart...
── more in #robotics 4 stories · sorted by recency
── more on @cap-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cap-x-lms-first-phys…] indexed:0 read:2min 2026-10-05 · —