{"slug": "robot-use-agents-why-the-harness-matters-more-than-the-model", "title": "Robot-Use Agents: Why the Harness Matters More Than the Model", "summary": "A developer built an MCP-based control layer that places a quadruped and a humanoid robot behind a single tool interface, arguing that the engineering burden for robot-use agents has shifted from the model to the harness — the tools, pre-action checks, and evals around it. Drawing on a YC episode with the founders of Waddle Labs and RoboCurve, the writeup notes that general-purpose models can now control robots with little robot-specific training, that in-context learning plateaued after roughly 20 to 40 examples in one guest's experiment, and that repeated tasks should be distilled into code or reusable skills. The accompanying FastMCP example runs against a simulated robot, validating targets against a workspace bounding box before dispatching them to a swappable adapter.", "body_md": "General-purpose LLM agents can now control robots with little or no robot-specific training. The hard engineering work is moving from the model to the harness around it: the tools it can call, the checks that run before it acts, and the evals that tell you whether a change helped.\n\nI watched YC's episode [Robot-Use Agents: Why General-Purpose Models May Win in Robotics](https://www.youtube.com/watch?v=Jv5B5CEaPJI) with the founders of [Waddle Labs](https://www.waddlelabs.ai/) and [RoboCurve](https://robocurve.org/). Below are the four ideas that stayed with me, then a small harness you can run locally. I have built an MCP based control layer that puts a quadruped and a humanoid behind one tool interface, so this maps closely to my own work.\n\nThe episode builds on MIT professor Phillip Isola's idea of robot-use agents: general-purpose models that control different robots, write policies, and learn new physical tasks with little or no robot-specific training.\n\nEarlier approaches such as [RT-2](https://robotics-transformer2.github.io/) trained a vision-language model to output robot actions instead of text. The guests argue that today's models are strong enough that this kind of fine-tuning is often unnecessary. They can call tools and write code directly.\n\nGive the LLM a list of robot functions, such as picking up an object or moving to a pose, and let it write the plan. Google's [Code as Policies](https://code-as-policies.github.io/) work (2022) showed this can work with only a few examples in the prompt, because the model has already seen so much code that sequences steps.\n\nAn LLM thinking at every step is slow. The guests' approach: when a task repeats, turn it into code or a reusable skill. Keep the model for the points of variation, such as recognising an unfamiliar object or deciding what to do after an action fails.\n\nThey also said frontier LLM latency is improving about 2x per month. Treat that as their claim, not an established figure, but it explains why they expect real-time control to get closer.\n\nDragging a cursor to orbit a CAD model teaches top-down, left-right and near-far. The guests believe this kind of data helps newer models reason about space. One practical tip from the episode: design robot tools so they resemble computer-use tools, and the model does better on physical tasks.\n\nOne guest shared an experiment where in-context learning stopped improving after roughly 20 to 40 examples. So lessons from a deployment have to leave the context window: first as skills, eventually as weights. It also means you need evals, or you cannot tell whether a harness change helped.\n\nThe pattern: the model asks, the harness checks, the adapter acts. This uses [FastMCP](https://gofastmcp.com/) and a simulated robot, so it runs without hardware.\n\n`harness.py`\n\n``` python\nfrom typing import Protocol\n\nfrom fastmcp import FastMCP\n\nmcp = FastMCP(\"robot-harness\")\n\nclass RobotAdapter(Protocol):\n    \"\"\"One adapter per platform. Transport details stay inside.\"\"\"\n\n    def move_to(self, x: float, y: float, z: float) -> None: ...\n    def stop(self) -> None: ...\n\nclass SimAdapter:\n    \"\"\"Stand-in robot so the harness runs without hardware.\"\"\"\n\n    def move_to(self, x: float, y: float, z: float) -> None:\n        print(f\"[sim] moving to ({x}, {y}, {z})\")\n\n    def stop(self) -> None:\n        print(\"[sim] stopped\")\n\n# Swap SimAdapter for a real adapter (WebRTC, DDS, ROS, and so on).\nrobot: RobotAdapter = SimAdapter()\n\nWORKSPACE = {\"x\": (-0.4, 0.4), \"y\": (-0.4, 0.4), \"z\": (0.0, 0.5)}  # meters\n\ndef in_workspace(x: float, y: float, z: float) -> bool:\n    return all(\n        lo <= v <= hi\n        for v, (lo, hi) in zip((x, y, z), WORKSPACE.values())\n    )\n\n@mcp.tool()\ndef move_to(x: float, y: float, z: float) -> str:\n    \"\"\"Move the end effector to a target in meters, in the robot base frame.\"\"\"\n    if not in_workspace(x, y, z):\n        return \"rejected: target is outside the allowed workspace\"\n    robot.move_to(x, y, z)\n    return \"ok\"\n\n@mcp.tool()\ndef stop() -> str:\n    \"\"\"Stop all motion immediately.\"\"\"\n    robot.stop()\n    return \"stopped\"\n\nif __name__ == \"__main__\":\n    mcp.run()\n```\n\n`try_it.py` calls the tools the same way an agent would, through a real MCP client:\n\n``` python\nimport asyncio\n\nfrom fastmcp import Client\n\nfrom harness import mcp\n\nasync def main() -> None:\n    async with Client(mcp) as client:\n        tools = await client.list_tools()\n        print(\"tools:\", [t.name for t in tools])\n\n        ok = await client.call_tool(\"move_to\", {\"x\": 0.2, \"y\": 0.1, \"z\": 0.3})\n        print(\"inside workspace:\", ok.data)\n\n        bad = await client.call_tool(\"move_to\", {\"x\": 2.0, \"y\": 0.0, \"z\": 0.3})\n        print(\"outside workspace:\", bad.data)\n\nasyncio.run(main())\n```\n\nRun it:\n\n```\npip install fastmcp\npython try_it.py\n```\n\nOutput:\n\n```\ntools: ['move_to', 'stop']\n[sim] moving to (0.2, 0.1, 0.3)\ninside workspace: ok\noutside workspace: rejected: target is outside the allowed workspace\n```\n\nThree things to notice:\n\nPhysical AI engineering is starting to look a lot like agent engineering. Tool design, safety gates, and evals decide whether a good model becomes a useful robot.\n\nWhich matters more right now, a better model or a better harness? I would like to hear your view in the comments.", "url": "https://wpnews.pro/news/robot-use-agents-why-the-harness-matters-more-than-the-model", "canonical_source": "https://dev.to/sravya_dangeti/robot-use-agents-why-the-harness-matters-more-than-the-model-2j3c", "published_at": "2026-10-11 11:42:20+00:00", "updated_at": "2026-10-11 11:51:51.467508+00:00", "lang": "en", "topics": ["robotics", "ai-agents", "agent-protocols", "large-language-models", "ai-tools"], "entities": ["Waddle Labs", "RoboCurve", "FastMCP", "Phillip Isola", "MIT", "Y Combinator", "RT-2", "Code as Policies"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/robot-use-agents-why-the-harness-matters-more-than-the-model", "markdown": "https://wpnews.pro/news/robot-use-agents-why-the-harness-matters-more-than-the-model.md", "text": "https://wpnews.pro/news/robot-use-agents-why-the-harness-matters-more-than-the-model.txt", "jsonld": "https://wpnews.pro/news/robot-use-agents-why-the-harness-matters-more-than-the-model.jsonld"}}