{"slug": "gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens", "title": "GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens", "summary": "PyRUA-Lean, a robot-agent framework that replaces per-step tool calling with generated Python cells, raised task success to 71.7% from 63.1% while cutting input tokens 65% (276k vs 788k), LLM calls 49% (8.7 vs 17.0) and cost 2.2x ($0.74 vs $1.63 per solved episode) on 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0 and RoboCasa365, according to the reported benchmark. The framework keeps the same vision-language model, robot primitives and VLA policies, giving the agent one python(code) tool and a robo object whose methods are those primitives, so a single cell chains steps, checks results and retries locally. In a recorded placement example, one PyRUA-Lean cell written by GPT-6 Astra replaced four baseline LLM calls for placing a bowl on a plate.", "body_md": "PyRUA-Lean matches or exceeds tool calling’s success rate on each sub-suite and costs less on every benchmark.(a) Success rate on all 700 instances, tagged by task property (semantic: perturbed layouts, objects, goals); RoboTwin 2.0 split by skill from task names.\n(b, c) Input tokens and GPT-6 Astra list-price dollars per solved episode, on instances both agents solved; Avg: mean over all of them. RC365: RoboCasa365.\n\nSame planner, same primitives: code instead of tool calls\n\nSuccess rate\n\n71.7%vs 63.1% with tool calling\n\nInput tokens\n\n65% fewer276k vs 788k per solved episode\n\nLLM calls\n\n49% fewer8.7 vs 17.0 per solved episode\n\nCost\n\n2.2× cheaper$0.74 vs $1.63 per solved episode\n\n700 simulated task instances from LIBERO-PRO, RoboTwin 2.0 and RoboCasa365; tokens, calls and cost on the instances both agents solved.\n\nStop paying one LLM call per step\n\nA robot agent built on a vision-language model usually acts through tool calling. The\nmodel picks one tool, such as move_to or release; the robot runs it; the result\ncomes back, with three camera images after every motion; and the model is called again with everything\nso far. Every small step costs a full LLM call, and every call reads the growing conversation again,\nimages included.\n\nPyRUA-Lean (Python for Lean Robot-Use Agents) keeps the model, the robot primitives and\nthe VLA policies, and changes only how the agent acts. The agent gets one tool, python(code),\nand a robot object robo whose methods are those same primitives. In one LLM call it writes a\ncell: a few lines of Python that chain several steps, check each result, retry when a policy\nfalls short, and compute what no tool returns. The cell runs on the robot; the agent then sees only what\nthe cell printed and the camera images it asked for.\n\nPyRUA-Lean combines primitive composition with selective feedback.\n(a) The tool-calling agent invokes robot primitives through tool calls. Operations that depend on preceding results generally require another VLM turn, and motion calls automatically return images and state.\n(b) PyRUA-Lean instead generates Python cells that compose robot primitives with helper functions, conditional checks, and local retries. Intermediate execution stays within the runtime, while explicitly requested images and state messages are recorded and returned at cell end for replanning.\n(c) A recorded placement example illustrates both mechanisms: one PyRUA-Lean cell replaces baseline steps 14–17, reducing four LLM turns to one.\n\nOne cell instead of four LLM calls\n\nThis is the cell of panel (c), exactly as GPT-6 Astra wrote it while placing a bowl on a plate. Tool\ncalling needed four LLM calls for the same step: two to locate the bowl and sample its base, one to lower\nit and one to release it, and each move returned three camera images. The cell measures the bowl's bottom\nin the point cloud, lowers it to the right height, releases it only if the move arrived, and asks for an\nimage only if the task is still not done. The three comment lines are ours; plate was found\nby an earlier cell.\n\n```\n# compose: compute a placement target from the scene geometry\nheld_at_plate = robo.segment(point=(314, 795))\nprint('bowl above plate', held_at_plate)\nassert held_at_plate.found\ncorrection = np.array(plate.world_xyz[:2]) - np.array(held_at_plate.world_xyz[:2])\nplace_xy = np.array(robo.state().eef_pos[:2]) + correction\nsurface_map = robo.world_map()\nbowl_patch = surface_map[290:365, 740:850]\nbowl_heights = bowl_patch[:,:,2]\nbowl_heights = bowl_heights[(bowl_heights>1.04) & (bowl_heights<1.17)]\nbottom_z = float(np.quantile(bowl_heights, 0.02))\nplacement_z = plate.world_xyz[2] + robo.state().eef_pos[2] - bottom_z + 0.007\nprint('correction', correction, 'bottom', bottom_z, 'placement z', placement_z)\nassert np.linalg.norm(correction)<0.07 and 0.95<placement_z<1.05\n# compose: lower, check, and release within the cell\nlower = robo.move_to([*place_xy, placement_z], gripper=+1, step_clip=0.012, tol=0.005, max_steps=100)\nprint('lower', lower)\nif lower.reached and not robo.done:\n    print('release', robo.release())\nprint('done', robo.done)\n# select: ask for an image only if the task is unfinished\nif not robo.done:\n    robo.show('agentview')\n```\n\nIt printed five lines and asked for no image: the bowl was on the plate and the task was done.\n\nOne task, two agents\n\nThe whole episode behind that example, run once with each agent on the same LIBERO-PRO task instance:\n“Pick the akita black bowl on the stove and place it on the plate”. Step through it one LLM call\nat a time: each square is one call, a filled square brought camera images back, and the outlined square\nsolved the task.\n\nTurn on JavaScript to step through the two episodes.\n\nWhat the agent writes\n\nWe read the code agent's cells on every benchmark, in solved and failed episodes, and counted what they do\nover all 11,145 cells of the 700 episodes. Four kinds of cells keep coming back, and none of them fits in\none tool call. Each example is a real cell, exactly as the model wrote it.\n\nGuarded chains\n\n19% of cells · 45% on LIBERO-PRO\n\nSeveral primitives in a row, each run only if the one before worked: move above the bowl, grasp with the\nVLA policy only if the move arrived and the task is not done, and look only if it is still not done.\nWith tool calling, each link of the chain is an LLM call.\n\n```\nassert bowl.found and plate.found\nbowl_xy = np.array(bowl.world_xyz[:2])\nplate_xyz = np.array(plate.world_xyz)\ntable_z = robo.back_project(837,660).world_xyz[2]\nprepose = [float(bowl_xy[0]),float(bowl_xy[1]),table_z+0.22]\napproach = robo.move_to(prepose,gripper=robo.OPEN)\nprint('APPROACH',approach)\nif approach.reached and not robo.done:\n    picked = robo.pi0_pick('pick up the black bowl between the plate and the ramekin', max_chunks=20)\n    print('PICK',picked)\nprint('STATE',robo.state())\nif not robo.done:\n    robo.show('agentview')\n    robo.show('wrist')\n```\n\nCell 2 · LIBERO-PRO, spatial swap, task 0, seed 0\n\nRetries\n\n8% of cells · 18% on RoboCasa365 atomic\n\nWhere a VLA policy may stop short of the goal, the cell runs it again until the task reports success.\nWith tool calling, every retry is a round trip that brings back three camera images.\n\nThe cell reads a camera's image and point cloud as arrays and locates objects itself: a colour mask for\na block, the points above a height for its top, their principal axis for its orientation, and from that\nthe yaw of the grasp. A tool-calling agent gets only what its tools compute.\n\n```\nmask=red&(rows<92)&(cols>98)&(cols<165)&(world[...,2]>0.745)\npoints=world[mask]\nupper=points[points[:,2]>0.794]\nmean_xy=np.mean(upper[:,:2],axis=0)\neigenvalues,eigenvectors=np.linalg.eigh(np.cov(upper[:,:2].T))\nlong_axis=eigenvectors[:,1]\nif long_axis[1]<0: long_axis=-long_axis\nshort_axis=np.array([long_axis[1],-long_axis[0]])\nprojection_long=upper[:,:2]@long_axis\nprojection_short=upper[:,:2]@short_axis\nblock_center=long_axis*((np.min(projection_long)+np.max(projection_long))/2)+short_axis*((np.min(projection_short)+np.max(projection_short))/2)\nprint('flat centre',block_center,'long axis',long_axis,'extents',np.ptp(projection_long),np.ptp(projection_short))\nleft_grasp_xy=block_center+0.045*long_axis\nright_grasp_xy=block_center-0.05*long_axis\nyaw=math.atan2(long_axis[1],long_axis[0])\nq_down=np.array([math.cos(yaw/2)/math.sqrt(2),-math.sin(yaw/2)/math.sqrt(2),math.cos(yaw/2)/math.sqrt(2),math.sin(yaw/2)/math.sqrt(2)])\nprint('grasp',left_grasp_xy,'receive',right_grasp_xy,'quat',q_down)\nsafe=np.array(robo.state().left.eef_pos);safe[2]=1.03\nif checked_move('left',safe,quat=robo.state().left.eef_quat,substeps=20):\n    waypoint=np.array([left_grasp_xy[0],0.015,1.03])\n    if checked_move('left',waypoint,quat=q_down,substeps=25):\n        print(robo.move_to('left',[left_grasp_xy[0],left_grasp_xy[1],0.97],quat=q_down,substeps=25))\nprint(robo.state().left)\nrobo.show('head')\n```\n\nCell 5 · RoboTwin 2.0 without the VLA policy, handover block, seed 100003.\nchecked_move is a helper the agent defined in an earlier cell.\n\nSkills of its own\n\n2% of cells define a helper · 7% reuse one\n\nThe agent writes a descent that moves down in 8 mm steps and stops as soon as a step does not reach\nits target, stops going down, or drifts sideways, and reuses it in its next cell. Helpers like this are\nhow it grasps without a VLA policy.\n\n``` python\ndef lower_guarded(arm_name, target_tcp_z):\n    initial=robo.state().arm(arm_name)\n    target=np.array(initial.eef_pos)\n    final_eef_z=target_tcp_z-0.12*initial.approach[2]\n    while target[2]>final_eef_z+0.001 and not robo.done:\n        before=robo.state().arm(arm_name)\n        target[2]=max(final_eef_z,target[2]-0.008)\n        result=robo.move_to(arm_name,target,substeps=8)\n        after=robo.state().arm(arm_name)\n        print('descent',result.reached,tuple(round(v,4) for v in after.tcp_pos))\n        if result.terminated or not result.planned or not result.reached or after.eef_pos[2]>=before.eef_pos[2]-0.001 or np.linalg.norm(np.array(after.eef_pos[:2])-target[:2])>0.01:\n            print('stop',result)\n            break\nlower_guarded('right',0.800)\nrobo.show('right_wrist')\nrobo.show('head')\n```\n\nCell 7 · RoboTwin 2.0, place object stand, seed 100002\n\nIt also chooses when to look: 84% of cells ask for a camera image,\nand 20% look only if the task is not yet done.\n\nWhen things go wrong\n\nCode helps most after a setback. PyRUA-Lean alone solved 96 task instances; tool calling alone solved 36.\nIn most of the instances only code solved, both agents met the same setback, most often a VLA grasp that\nmissed or an object that fell. Tool calling then asked the policy again, or stopped; the code agent closed\nthe gripper itself, scripted the recovery from the same primitives, and checked each step. Two of them:\n\nThe dropped moka pot\n\nLIBERO-PRO · “turn on the stove and put the moka pot on it”\n\nBoth agents turned the stove on and dropped the pot. Tool calling asked the VLA policy to grasp the fallen\npot four times and ended the episode. The code agent grasped it itself, fitted a plane to the points of its\nlid to find how it leans, turned it upright, and set its base, not the gripper, over the burner.\n\nTool calling (RPent)gives up at LLM call 32PyRUA-Leansolved at LLM call 33\n\nSimulator recordings at 10× speed.\n\nTop: the recorded camera views after the LLM calls named. Bottom: the code agent's cells, excerpts exactly as the model wrote them; “...” marks lines left out. The plane fit in call 23 raised an error (NumPy's linear algebra refuses half-precision arrays); call 24 cast the points and ran it again.\n\nThe cabinet door, pulled along its arc\n\nRoboCasa365 atomic · “Open the cabinet door.”\n\nBoth agents located the door's hinge. Tool calling pulled the handle a few centimeters per LLM call and was\nstill pulling when its 40 calls ran out. The code agent took the hinge as the centre of the circle through\nthree positions of the handle, two of which it copied from earlier outputs, and pulled along that arc in\none cell, stopping if a move fell short or the grip slipped. The door opened, and a VLA call in call 34\ncompleted the task.\n\nTool calling (RPent)still pulling when its 40 LLM calls ran outPyRUA-Leansolved at LLM call 34\n\nSimulator recordings at 15× speed.\n\nTop: the recorded camera views; bottom: the code agent's cells, excerpts exactly as the model wrote them.\n\nResults\n\nMore tasks solved. Across 700 task instances, PyRUA-Lean raises the success rate from\n63.1% to 71.7%, an improvement of 8.6 percentage points, or about 14% relative. It gains on every benchmark:\n11.0 points on LIBERO-PRO, 9.2 on RoboTwin 2.0, 7.8 on RoboCasa365 atomic tasks and 5.0 on composite tasks.\n\nFewer calls and tokens for the same tasks. On the instances both agents solved, PyRUA-Lean\nreduces mean LLM calls from 17.0 to 8.7 and input tokens from 788k to 276k, reductions of 49% and 65%.\nAt list prices, an episode costs $0.74 instead of $1.63; the saving in dollars is smaller than in tokens,\nin part because caching discounts the history that tool calling sends again and again.\n\nThe same success for a fraction of the tokens. Cut every recorded episode at a token budget\nand count what is solved by then: on LIBERO-PRO, PyRUA-Lean reaches tool calling's final success rate with\n564k tokens per episode, tool calling with 3.18M.\n\nSuccess rate if every episode were stopped once its token consumption reached the budget on the horizontal axis. Dashed lines mark the budget at which each agent first reaches tool calling’s final success rate, with its value on the axis; the large number is their ratio. RC365: RoboCasa365.\n\nMostly from calling the model less often. The token ratio splits into the ratio of LLM\ncalls and the ratio of input tokens per call. On LIBERO-PRO, the 4.48× token ratio comprises\n2.55× fewer calls and 1.76× fewer tokens per call. On RoboCasa365 atomic tasks, PyRUA-Lean's\ncalls are larger on average, yet fewer calls still reduce the total.\n\nAlso without VLA policies or guides. Removing the VLA policies, the operating guides or\nboth from both agents, PyRUA-Lean keeps the higher success rate and the lower token use. Without VLA\npolicies it reaches 68.8% on RoboTwin 2.0 against 42.8% for tool calling, composing classical\nprimitives into approach, gripper control and state checks.\n\nMore episodes solved with code\n\nNine episodes per benchmark. Pick one to see every cell the agent wrote in it, and what each cell printed.\n\nPut the bowl on the top of the drawer Code · 5 cellsPick the akita black bowl on the ramekin and place it on the plate Code · 5 cells2×put both the cream cheese box and the butter in the basket Code · 9 cellsPick the ketchup and place it in the basket Code · 4 cellsOpen the top layer of the drawer and put the bowl inside Code · 6 cells3×Put the wine bottle on the rack Code · 8 cells2×Pick the akita black bowl in the top layer of the wooden cabinet and place it on the plate Code · 8 cells4×put the black bowl in the bottom drawer of the cabinet and close it Code · 14 cellsPick the alphabet soup and place it in the basket Code · 5 cells\n\nOpen the laptop Code · 2 cellsPlace the empty cup Code · 2 cellsRotate the QR code Code · 2 cells2×Rank the blocks in RGB order Code · 7 cellsPress the stapler Code · 7 cellsHand the block from one arm to the other Code · 5 cells2×Stack three bowls Code · 7 cells2×Beat the block with the hammer Code · 7 cellsOpen the microwave Code · 5 cells\n\nClose the toaster oven door Code · 1 cellTurn on the sink faucet Code · 1 cell6×Turn on the microwave Code · 6 cells7×Move an item from the drawer to the counter Code · 8 cells4×Close the fridge Code · 4 cellsOpen the drawer Code · 1 cell2×Move an item from the counter to the stove Code · 2 cells14×Turn off the stove Code · 10 cellsNavigate the kitchen Code · 1 cell\n\n3×Stack bowls in the cabinet Code · 1 cell4×Wash the lettuce Code · 2 cells7×Load the dishwasher Code · 6 cells3×Rinse the sink basin Code · 6 cells6×Select the cutting tool Code · 7 cells10×Portion the hot dogs Code · 13 cells12×Reheat the waffle Code · 10 cells8×Scrub the cutting board Code · 13 cells10×Pack identical lunches Code · 6 cells\n\nSimulator recordings of PyRUA-Lean episodes from the main comparison; long episodes sped up as marked.\n\nHow we compared\n\nLIBERO-PRO\n\n40 tasks × 5 seeds = 200 instances\n\nTabletop tasks from its four perturbed suites. Franka arm; VLA policy π0.5.\n\nRoboTwin 2.0\n\n50 tasks × 5 seeds = 250 instances\n\nDual-arm tasks with randomized scenes. VLA policy LingBot-VLA.\n\nKitchen tasks for a mobile manipulator. VLA policy RLDX-1.\n\nBoth agents use GPT-6 Astra at high reasoning effort through the Codex CLI, with the same robot stacks,\nprimitive implementations and VLA policies, and SAM 3 for segmentation. Each task instance runs once with\neach agent, under a budget of 40 LLM calls per episode (100 on the composite tasks), and each benchmark's\nown success check decides the outcome. The tool-calling baseline is RPent, whose robot stacks PyRUA-Lean\nbuilds on.\n\nTry it\n\nPyRUA-Lean drives the robot stacks of RPent.\nInstall it into the Python environment of an RPent robot stack; the\nREADME covers the three robot stacks\nand a full episode.", "url": "https://wpnews.pro/news/gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens", "canonical_source": "https://dagroup-pku.github.io/PyRUA-Lean/", "published_at": "2026-10-02 09:20:50+00:00", "updated_at": "2026-10-02 09:36:45.688801+00:00", "lang": "en", "topics": ["robotics", "ai-agents", "large-language-models", "ai-research"], "entities": ["PyRUA-Lean", "GPT-6 Astra", "LIBERO-PRO", "RoboTwin 2.0", "RoboCasa365", "RC365"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens", "markdown": "https://wpnews.pro/news/gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens.md", "text": "https://wpnews.pro/news/gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-robot-agents-with-14-higher-success-rate-but-65-fewer-tokens.jsonld"}}