Robot-Use Agents: Why the Harness Matters More Than the Model A developer built an MCP-based control layer that places a quadruped and a humanoid robot behind a single tool interface, arguing that the engineering burden for robot-use agents has shifted from the model to the harness — the tools, pre-action checks, and evals around it. Drawing on a YC episode with the founders of Waddle Labs and RoboCurve, the writeup notes that general-purpose models can now control robots with little robot-specific training, that in-context learning plateaued after roughly 20 to 40 examples in one guest's experiment, and that repeated tasks should be distilled into code or reusable skills. The accompanying FastMCP example runs against a simulated robot, validating targets against a workspace bounding box before dispatching them to a swappable adapter. General-purpose LLM agents can now control robots with little or no robot-specific training. The hard engineering work is moving from the model to the harness around it: the tools it can call, the checks that run before it acts, and the evals that tell you whether a change helped. I watched YC's episode Robot-Use Agents: Why General-Purpose Models May Win in Robotics https://www.youtube.com/watch?v=Jv5B5CEaPJI with the founders of Waddle Labs https://www.waddlelabs.ai/ and RoboCurve https://robocurve.org/ . Below are the four ideas that stayed with me, then a small harness you can run locally. I have built an MCP based control layer that puts a quadruped and a humanoid behind one tool interface, so this maps closely to my own work. The episode builds on MIT professor Phillip Isola's idea of robot-use agents: general-purpose models that control different robots, write policies, and learn new physical tasks with little or no robot-specific training. Earlier approaches such as RT-2 https://robotics-transformer2.github.io/ trained a vision-language model to output robot actions instead of text. The guests argue that today's models are strong enough that this kind of fine-tuning is often unnecessary. They can call tools and write code directly. Give the LLM a list of robot functions, such as picking up an object or moving to a pose, and let it write the plan. Google's Code as Policies https://code-as-policies.github.io/ work 2022 showed this can work with only a few examples in the prompt, because the model has already seen so much code that sequences steps. An LLM thinking at every step is slow. The guests' approach: when a task repeats, turn it into code or a reusable skill. Keep the model for the points of variation, such as recognising an unfamiliar object or deciding what to do after an action fails. They also said frontier LLM latency is improving about 2x per month. Treat that as their claim, not an established figure, but it explains why they expect real-time control to get closer. Dragging a cursor to orbit a CAD model teaches top-down, left-right and near-far. The guests believe this kind of data helps newer models reason about space. One practical tip from the episode: design robot tools so they resemble computer-use tools, and the model does better on physical tasks. One guest shared an experiment where in-context learning stopped improving after roughly 20 to 40 examples. So lessons from a deployment have to leave the context window: first as skills, eventually as weights. It also means you need evals, or you cannot tell whether a harness change helped. The pattern: the model asks, the harness checks, the adapter acts. This uses FastMCP https://gofastmcp.com/ and a simulated robot, so it runs without hardware. harness.py python from typing import Protocol from fastmcp import FastMCP mcp = FastMCP "robot-harness" class RobotAdapter Protocol : """One adapter per platform. Transport details stay inside.""" def move to self, x: float, y: float, z: float - None: ... def stop self - None: ... class SimAdapter: """Stand-in robot so the harness runs without hardware.""" def move to self, x: float, y: float, z: float - None: print f" sim moving to {x}, {y}, {z} " def stop self - None: print " sim stopped" Swap SimAdapter for a real adapter WebRTC, DDS, ROS, and so on . robot: RobotAdapter = SimAdapter WORKSPACE = {"x": -0.4, 0.4 , "y": -0.4, 0.4 , "z": 0.0, 0.5 } meters def in workspace x: float, y: float, z: float - bool: return all lo <= v <= hi for v, lo, hi in zip x, y, z , WORKSPACE.values @mcp.tool def move to x: float, y: float, z: float - str: """Move the end effector to a target in meters, in the robot base frame.""" if not in workspace x, y, z : return "rejected: target is outside the allowed workspace" robot.move to x, y, z return "ok" @mcp.tool def stop - str: """Stop all motion immediately.""" robot.stop return "stopped" if name == " main ": mcp.run try it.py calls the tools the same way an agent would, through a real MCP client: python import asyncio from fastmcp import Client from harness import mcp async def main - None: async with Client mcp as client: tools = await client.list tools print "tools:", t.name for t in tools ok = await client.call tool "move to", {"x": 0.2, "y": 0.1, "z": 0.3} print "inside workspace:", ok.data bad = await client.call tool "move to", {"x": 2.0, "y": 0.0, "z": 0.3} print "outside workspace:", bad.data asyncio.run main Run it: pip install fastmcp python try it.py Output: tools: 'move to', 'stop' sim moving to 0.2, 0.1, 0.3 inside workspace: ok outside workspace: rejected: target is outside the allowed workspace Three things to notice: Physical AI engineering is starting to look a lot like agent engineering. Tool design, safety gates, and evals decide whether a good model becomes a useful robot. Which matters more right now, a better model or a better harness? I would like to hear your view in the comments.