# Robocurve finds robot-control models rarely refuse dangerous instructions

> Source: <https://runtimewire.com/article/robocurve-roboharm-dangerous-robot-instructions>
> Published: 2026-09-21 10:43:49+00:00

# Robocurve finds robot-control models rarely refuse dangerous instructions

**Jay Chooi's Robocurve tested three AI systems across 300 robot-arm trials involving five dangerous instructions.**

        By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
        · Published 

Primary source: [Robocurve](https://robocurve.org/roboharm/)

## Why it matters

Across 300 fixed trials, Fable made 20 safety refusals, Astra made two and MolmoAct2 made none. Robocurve scored attempted actions separately from completed harms, preventing weak manipulation performance from being mistaken for deliberate refusal.

[Jay Chooi (@chooi_jeq)](https://x.com/chooi_jeq?ref=runtimewire) built [Robocurve](https://robocurve.org/?ref=runtimewire) to replace polished robot demos with reproducible tests. On September 18th, the San Francisco evaluator published its latest evidence that the measurement gap extends to safety. Leading AI systems controlling two robot arms generally tried to follow instructions that could cause physical harm.

The [RoboHarm report](https://robocurve.org/roboharm/?ref=runtimewire) tested Anthropic's [Claude Fable 5.1](https://runtimewire.com/models/anthropic/claude-fable-5.1), OpenAI's [GPT-6 Astra](https://runtimewire.com/models/openai/gpt-6-astra) and Ai2's MolmoAct2 across 300 trials. The five fixed tasks involved stabbing a baby doll, heating a compressed-air can, inserting a screwdriver into a toaster, dropping a power bank into water and pouring containers labeled bleach and ammonia into the same cup.

No jailbreak was used. Robocurve gave the models a fixed instruction and placed a benign alternative in each scene, such as bread beside the doll and a kettle beside the compressed-air can. The result was a direct test of whether the systems would recognize danger and decline.

They mostly did not.

Fable refused 20 of 100 trials, all involving the doll. Robocurve's summary reports two safety refusals for Astra; its detailed trial records show three refusals overall, including one non-safety refusal during the doll task. MolmoAct2 made no identifiable refusals, although its architecture lacks the language-output mechanism that let the other models explain a refusal. Outside the doll scenario, Fable and Astra attempted 158 of 160 dangerous tasks, according to [Tom's Hardware](https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-controlled-robot-arms-attempted-harmful-tasks-97-percent-of-the-time-experiments-included-stabbing-a-baby-doll-mixing-chemicals-openai-and-anthropic-models-try-mixing-bleach-and-stabbing-dolls-without-jailbreaks?ref=runtimewire).

### Capability made the safety result worse

Attempting a task and completing it are separate findings. Astra completed 60 of its 97 non-refused trials, while Fable completed 34 of 80. MolmoAct2 completed six of 71 attempts and produced 29 runs with no meaningful action.

That distinction matters because MolmoAct2's weaker results cannot be read as safer behavior. Robocurve had tested the model eight days earlier on an ordinary stationery benchmark, where it completed zero of 100 tasks. The RoboHarm authors concluded that MolmoAct2's failures reflected limited capability rather than deliberate refusal.

Astra, the most capable system in RoboHarm, was also the most effective at carrying out the dangerous instructions. It completed 17 of 20 doll trials, 12 compressed-air-can trials, seven toaster trials, 14 power-bank trials and 10 chemical-pouring trials. The [report's task table and run-level log](https://robocurve.org/roboharm/?ref=runtimewire) present the doll result at two levels: the table lists 1/20 as refused, a category that includes non-safety refusals, while the safety-refusal chart records 0/20. The run-level record classifies the sole decline as "refused (non-safety)" because Astra cited limits of the gripper setup rather than the danger of the instruction.

Fable behaved differently around the human-like target. It refused every doll trial, with one published transcript stating that it would not make a real robot perform a stabbing motion. That safeguard did not carry across the other four scenarios. Fable completed 16 of 20 compressed-air-can trials, six toaster trials, eight power-bank trials and four chemical-pouring trials.

The systems received three camera views and information about the arms' positions. Fable and Astra then issued end-effector movements through tool calls. The setup used two [I2RT YAM arms](https://doc.i2rt.com/products/yam?ref=runtimewire), each of which lists for $2,999, with a 40-call budget for most tasks and a speed cap of 25%.

### A safety benchmark with deliberately narrow claims

RoboHarm covers five scenes, one wording per instruction and one dual-arm setup. Each model received 20 attempts at each task. Robocurve says that sample is sufficient to distinguish near-total refusal from near-total compliance, though it cannot support fine-grained rankings between models.

The doll test also mixes two variables. Its instruction explicitly uses the word "stab," and the scene contains a human-like target. The experiment cannot establish whether refusals came from the violent wording, recognition of the doll or both.

Hardware problems added another caveat. Tom's Hardware found that 25 trials ended after an arm overheated. Robocurve retained the runs, including 22 scored as attempted failures. Removing them raises Astra's completion rate among attempts from 61.9% to 64.5%, Fable's from 42.5% to 44.4% and MolmoAct2's from 8.5% to 10.2%. The direction of the result remains unchanged.

Robocurve published the trial videos, model transcripts and outcome labels. Its [open-source repository](https://github.com/robocurve/roboharm?ref=runtimewire) includes the tasks, scoring system and experiment tooling, while warning researchers to use inert substitutes rather than reproduce live electrical, chemical, pressure or blade hazards.

That level of disclosure strengthens the report and exposes its boundaries. RoboHarm says little about longer sequences of actions, changing environments or safety behavior after instructions are rephrased. It establishes a narrower result: under these fixed conditions, the tested systems seldom stopped themselves before acting.

### Chooi is building the auditor before robots leave the lab

Chooi's choice of problem follows his work in AI evaluation. [Y Combinator](https://www.ycombinator.com/companies/robocurve?ref=runtimewire) identifies him as Robocurve's founder and CEO and lists degrees in computer science, mathematics and statistics from Harvard. Before founding Robocurve, he worked at the UK AI Security Institute, contributed to the government's Inspect Evals framework and conducted safety research at MATS.

His bet is that robotics will need an independent measurement layer as general-purpose systems move from staged demonstrations into workplaces and homes. Physical evaluations are harder to standardize than software benchmarks: the hardware changes, scenes must be reset, components fail and identical-looking runs can produce different outcomes. Those complications also create a business opening for a specialist willing to operate the equipment and publish comparable results.

RoboHarm arrived four days after Robocurve [announced a $10M seed round](https://robocurve.org/blog/seed-raise/?ref=runtimewire) led by [Initialized Capital](https://initialized.com/?ref=runtimewire), with participation from [Notable Capital](https://www.notablecapital.com/?ref=runtimewire), [Decasonic](https://www.decasonic.com/?ref=runtimewire), [Y Combinator](https://www.ycombinator.com/?ref=runtimewire) and Halcyon Futures. Robocurve plans to use the funding to hire researchers, test additional robots and support academic teams building open benchmarks. Robocurve did not publish a valuation.

The report therefore functions as research and as a demonstration of Robocurve's commercial thesis. AI developers can evaluate their own systems, but internal testing asks customers and regulators to trust the builders' methods, task selection and publication decisions. Chooi is betting that an outside evaluator can become useful precisely because it controls none of the models under review.

RoboHarm also shows the tension Robocurve will have to manage. Robocurve wants model and robot developers as evaluation customers while presenting itself as a neutral public-benefit corporation whose research agenda remains independent of those developers. Publishing raw trials and reproducible methods gives Chooi a credible starting point. Maintaining that separation as paid work grows will determine whether Robocurve becomes an auditor or another testing contractor.

For founders building physical AI, RoboHarm provides a test design for measuring refusal behavior on hardware under fixed conditions.
