Initialized Capital led the seed round for the YC-backed public benefit corporation, which publishes open-source robot evaluation tools and traces.
By [Ryan Merket](/author/ryan-merket)
· Published
Primary source: [X](https://x.com/chooi_jeq/status/2099545422876541001)
Why it matters #
Robotics labs still control most demos and evaluations. Robocurve is using venture capital to build an open third-party measurement layer before robot claims outrun the evidence.
Jay Chooi (@chooi_jeq), CEO and co-founder of Robocurve, said in a seven-post thread on X on September 14th that Robocurve raised a $10 million seed round to build an independent testing operation for AI models controlling physical robots.
Initialized Capital led the financing, with participation from Notable Capital, Decasonic, Y Combinator and Halcyon Futures. Robocurve said it will use the capital to expand its research staff, test additional robots and tasks, and fund academic groups developing open benchmarks.
Chooi graduated from Harvard with degrees in computer science, mathematics and statistics, and was elected a Rhodes Scholar in 2025. Before Robocurve, he worked on AI evaluation and safety research at the UK AI Security Institute and MATS Research. That background shaped Robocurve's core argument: robotics companies and AI labs can publish carefully selected demonstrations, while outsiders lack a consistent way to measure whether the same systems work across repeated trials and unfamiliar tasks.
Robocurve incorporated as a Delaware Public Benefit Corporation roughly three months before the round announcement. In its funding announcement, Robocurve said its public benefit mission requires it to evaluate frontier robotics systems independently and report the results publicly. Robocurve also says AI labs will not control its research agenda, methodology or published findings.
The structure provides a legal wrapper for the mission, while Robocurve's credibility will depend on its testing practices. Benchmark design, hardware setup, scoring rules and model access can materially change a robotics result. Publishing those details gives researchers a way to identify where a headline comparison is solid and where the experimental design limits the conclusion.
A benchmark business built on open traces
Robocurve's main software project, Inspect Robots, is an MIT-licensed evaluation framework that runs different policies, including large language models and vision-language-action models, across compatible robots and simulations. The framework records configuration data, model transcripts, grader scores and visualizations for each run.
The repository describes the software as being in early development and warns that its API may change between releases. It currently lists integrations for several physical platforms, including YAM bimanual arms, Franka arms, AgiBot's A2, Unitree's G1 and SO-ARM hardware, alongside simulation support.
Robocurve says Inspect Robots has recorded more than 97,000 package installs. It also claims its research received more than 6 million views during the first three months and that researchers from more than 200 institutions registered for its benchmark program. Those are company-reported distribution metrics. They measure attention and developer uptake rather than recurring revenue or paid evaluation work.
The early research shows why the underlying trial data matters. In a September 10th StationeryBench report, Robocurve compared OpenAI's GPT-6 Astra with Ai2's MolmoAct2 across five desk-based manipulation tasks and 200 total trials on YAM robot arms.
Astra fully completed seven of its 100 trials, while MolmoAct2 completed none. Astra achieved an average progress score of 46 out of 100, compared with 12 for MolmoAct2. The largest gap appeared in a marker task, where Astra removed and placed the cap successfully in five of 20 attempts. Neither model completed a paper-clip pouring task in 20 attempts.
Robocurve published limitations alongside the results. Human operators graded the trials while knowing which model was running, creating room for unconscious bias. MolmoAct2 was evaluated without task-specific fine-tuning, the models were not always tested on the same physical rig, and nearly half of MolmoAct2's trials never moved meaningfully from the starting position. Those qualifications narrow what can be inferred from the headline score, and their inclusion is central to Robocurve's pitch.
The round pays for physical repetition
Robot evaluations are harder to scale than software-only model benchmarks. Each physical trial can require hardware, cameras, object placement, safety controls and manual resets. Running enough repetitions to distinguish a reliable capability from a favorable demo consumes equipment time and staff hours.
Robocurve is allocating part of the seed capital to external researchers through a $500,000 open benchmark program. Robocurve says selected academic teams will receive $20,000 in cash, compute and materials funding, and a pair of YAM arms to build benchmarks containing at least 20 tasks. Participating teams are expected to evaluate at least three publicly accessible models and release the resulting work as open source.
Robocurve is also hiring a member of technical staff in San Francisco at a listed annual salary of $170,000 to $300,000. The job description calls for scaling the operation to hundreds of parallel real-world evaluations, developing low-latency inference infrastructure and extending Inspect Robots.
The $10 million round gives Chooi the budget to turn a young open-source project into a physical testing operation. The lasting asset will be trust in the measurements. That requires repeatable methods, complete traces and a willingness to publish results that model developers and robotics vendors would prefer to leave inside the demo reel.