Led by Gabriel Orlanski, LibraryDesignBench scores generated libraries by the correctness and simplicity of downstream code across 242 programming problems.
By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
· Published
Primary source: [Snorkel AI](https://x.com/SnorkelAI/status/2105366219675029941)
Why it matters #
As AI agents increasingly build on code written by other agents, LibraryDesignBench measures whether generated libraries make that handoff more reliable or simply add another layer to work around.
On September 29th, Gabriel Orlanski submitted a paper introducing LibraryDesignBench, a benchmark that asks a practical question about agent-written software: can one agent build a library that helps another agent write shorter, more correct code? Snorkel AI highlighted the project on October 1st.
Orlanski is a research fellow at Snorkel AI and a fourth-year Ph.D. student at the University of Wisconsin-Madison, where his advisers include Fred Sala, a co-author on the paper. His profile lists internships at X and Magic AI, as well as work on agents at Replit. The benchmark puts the software-engineering question into a testable form: evaluate a library by what other agents manage to do with it, rather than by whether its own code looks polished.
Score the handoff, not the library
In each task, a designer agent receives an open-ended specification describing what a library should do and a few example uses. It gets no required interface, method signatures, or tests. Three fixed implementer agents then use the resulting library to solve programming problems without seeing the hidden tests.
The benchmark covers 15 library-design tasks and 242 problems in Rust, Python, TypeScript, and Haskell. Tasks model libraries such as Rust's clap command-line parser and Python's pandas data-analysis package. The score combines the downstream solutions' hidden-test pass rate with their simplicity against reference solutions written using the established production libraries. The published score is a 0-to-100 composite, not a pass percentage.
That design catches a familiar failure mode in AI coding: a tool may compile and offer the requested functions, yet still leave later agents with awkward interfaces or extra work. In the authors' results, agent designers reproduce the abstractions of the human-written library in 11 of 15 tasks. But downstream agents often underuse the libraries they receive, reimplementing capabilities already present. The paper's analysis attributes much of the added code to rigid or difficult-to-use interfaces rather than missing features.
Snorkel's benchmark leaderboard puts the top listed run at 48.9, with an 86.6% pass rate; the human-written production-library reference scores 46.6, while the no-library reference scores 34.4. Those figures suggest agent-designed libraries can help on this test set, though the composite score also rewards simpler solutions. The result does not mean an AI-generated library is generally better than a human-built one.
The weaker results are as instructive as the headline ranking. One listed run scores below the no-library reference, and in Haskell an agent-written library scores below having no library in 70% of designer-task pairs. A poorly designed abstraction can therefore impose work instead of removing it. The leaderboard also shows that the execution harness changes results: the same model scores differently when paired with different coding-agent tools.
A research question with product stakes
Orlanski's project sits inside Snorkel AI's broader research program. Snorkel says it grew out of Stanford AI Lab and builds specialized data and environments for AI teams. Its research has increasingly focused on benchmarks and evaluations that probe failures conventional tests can miss. LibraryDesignBench extends that work from whether an agent can finish a coding task to whether it can create useful infrastructure for another agent.
That handoff is becoming a real engineering concern as AI-generated code accumulates in software projects. If agents build packages, APIs, or internal helpers that other agents must later navigate, library design becomes part of the coding system's performance, not just a matter of style. A clean interface could reduce repeated code and help downstream systems pass more tests; a confusing one can make an agent less effective even when the underlying functionality exists.
The paper also tests a possible improvement: give library designers more explicit, agent-focused guidance and let them use subagents to test their designs. The authors report that both approaches improve downstream scores and produce simpler programs. That result points toward design practices teams can test, rather than a claim that current agents have already solved reusable software architecture.
There are limits to what the benchmark establishes. Its 15 tasks span four languages, but do not represent every kind of software library or production environment. The score does not measure the generated library's security, runtime speed, maintainability, or correctness in isolation; it measures the code downstream agents produce with it. The researchers also built the benchmark with Snorkel colleagues and university collaborators, so wider replication will matter as the task set and outside evaluations grow.
Snorkel AI announced a $350 million Series E at a $3.5 billion valuation on September 22nd. RuntimeWire previously reported on the round and the company's data-service pivot in its coverage of Snorkel's financing. LibraryDesignBench is a research project, not a stated use of that financing; its relevance is that Snorkel is investing in the evaluation work that helps define what reliable agent performance should mean.
For Orlanski, the contribution is a way to measure an often-hidden part of agentic coding: whether the first agent leaves the next one with something worth using. The benchmark's early results show that the answer depends less on whether a library exists than on whether its interface fits the work agents actually have to do.