{"slug": "ruby-lean-a-ruby-semantics-with-a-type-soundness-proof", "title": "Ruby-lean: A Ruby semantics with a type soundness proof", "summary": "Software engineer Sam Xif built ruby-lean, an executable model of Ruby's semantics in the Lean proof assistant, together with a proof of type soundness for a small fragment of Sorbet's type system, according to the project write-up. The semantics is validated by differential testing against Ruby and passes a large swath of conformance tests, with source code published on GitHub and a playground available. Xif reported two lessons from the project: agents perform well on long-horizon tasks when the task definition is clear and fail unpredictably when it is not, and advanced AI working with proof assistants has the potential to change how software correctness is managed.", "body_md": "# ruby-lean: A Ruby semantics with a type soundness proof\n\nWritten for Software engineers and computer scientists, or technical hobbyists\n\n**Bottom line up front:** *I, with heavy augmentation from LLMs, have built an\nexecutable model of Ruby's semantics in Lean, together with a proof of type\nsoundness for a small fragment of Sorbet's type system built on the semantics.\nThe semantics is validated by differential testing against Ruby and passes a\nlarge swath of conformance tests. Through this project, I learned firstly that\nagents do great at long-horizon tasks if the task definition is clear; if it is\nnot clear, they mess up in unpredictable ways. Secondly, I learned that advanced\nAI, via its ability to work with proof assistants, has the potential to\nrevolutionize how we manage the correctness of software.*\n\nCheck out the `ruby-lean` [playground](https://samx.io/ruby-lean) and see it in action. For the\ntechnically inclined, see the\n[Technical Appendix](2026-09-26-ruby-lean-technical-appendix.html). View the\nsource code on [GitHub](https://github.com/sam-xif/ruby-lean).\n\n## Introduction: \"Semantics Done Quick\"[#](#introduction-semantics-done-quick)\n\nSuppose you have a program in your language of choice, and you want to prove\nthat it is correct. Proof here means *for all inputs*, of which there could be\ninfinitely many. No amount of unit tests can satisfy that obligation.[<sup>1</sup>](#tooltip-def-1)\nSo, you need a mathematical argument. This is where\n[formal methods](https://en.wikipedia.org/wiki/Formal_methods) shines.\n\nHow might you mount a mathematical argument for the correctness of a program,\nthough? First, you need to define correctness. A simple definition is \"this\nprogram terminates and returns a value.\" Now, you need to define *meaning*. To\nsee why, take the program `\"2\" + 2`. The meaning of this is ambiguous because it\ndepends on the programming language. If you've ever written JavaScript, you may\nrecognize that this program evaluates to `\"22\"`. In\n[Ruby](https://www.ruby-lang.org/en/), this program throws an exception. These\noutcomes hint at the *computational intent* of the statement in each language.\nIn JavaScript, the meaning is approximately \"concatenate the string '2' with the\nresult of coercing 2 into a string.\" In Ruby, the meaning is \"attempt to call\nthe string's `+` operator with 2 as an operand, which attempts to coerce 2 using\nthe `to_str` method.\" This fails as the code for `String#+` executes, because\nthe `Integer` class has no `to_str` method. This definition of meaning, given by\ncomputational intent, is known as a programming language's\n[*semantics*](<https://en.wikipedia.org/wiki/Semantics_(programming_languages)>).[<sup>2</sup>](#tooltip-def-2)\nWhen the semantics is loaded into a proof assistant like\n[Lean](https://lean-lang.org/), we can mechanically reason about programs and\ntheir correctness!\n\nArmed with a definition of correctness and a semantics, consider a program like\n`\"2\" + x`, where `x` is some input variable. In JavaScript, this program is\ncorrect for any `x` except specific cases like `x = Symbol()`. In Ruby, this\nprogram is correct for any `x` that can be coerced into a string (i.e., `x` has\na `to_str` method). To make the Ruby program correct for all inputs, we can make\nthe `to_str` check explicit:\n\n```\n\"2\" + x if x.respond_to?(:to_str)\n```\n\nThe astute reader might notice that even this is not correct! Evaluating\n`x.to_str` could still raise an exception or not terminate.\n\nThese example programs are dead simple, and yet they still have many edge cases\nthat are hard to reason about. This is why modeling the semantics of entire\nprogramming languages has historically taken multiple PhD-years of effort.\n[Mike Dodds](https://mikedodds.org/) proposed\n[Semantics Done Quick](https://oath.tech/pub/2026/05/semantics-done-quick/) out\nof the belief that advanced AI should be able to greatly accelerate the\nconstruction of useful semantics. I say \"useful\" because results about semantics\nmake it to academic conferences but don't find broad use in industry.\n\nWhy? It's not that semantics are inherently useless. Quite the opposite: they\nare *generally* useful because we can vary our definition of correctness to suit\nthe task at hand. They can be used, for instance, to prove *soundness of type\nchecking*, which states that if a type checker—like\n[Sorbet](https://sorbet.org/), which is widely used for Ruby—accepts a program,\nthen the typed parts of the program are truly free of type errors. This has\nimmediate business value: it provably rules out certain uncaught exceptions in\nproduction code. Semantics can additionally be used to prove robustness against\nmalicious inputs so that software can be relied on in security-critical\ncontexts.\n\nRather, semantics are not adopted because:\n\n1. they often have severe limitations, such as not being able to model the complex parts of a given language, which tend to be the most useful; and\n2. authoring proofs against them requires technical expertise in formal methods and in the specific nature of the semantics.\n\nAdvanced AI shows promise in addressing both of these impediments.\n\nThis summer, I worked on this project as a fellow in the\n[Apart Research](https://apartresearch.com/research)\n[Secure Program Synthesis](https://www.lesswrong.com/posts/8wtrLoDPyCfMLuHkt/how-to-solve-secure-program-synthesis)\nfellowship. As an outcome, I am pleased to announce `ruby-lean`, my semantics of\nRuby in Lean 4. `ruby-lean` is an executable semantics, meaning that it can\nexecute real Ruby code. It is modeled as a\n[CESK machine](https://en.wikipedia.org/wiki/CEK_Machine#CESK_machine). To\ndemonstrate that this semantics is indeed *useful*, I have also developed a\nsystem of type judgments based on Sorbet and proven it sound. I defined\nsoundness above, but to reiterate what this concretely means here: if a program\npasses the `ruby-lean` type validator, then it is completely free of a certain\nfamily of type-related exceptions. **Both of the aforementioned impediments to\nwidespread use of this semantics are effectively addressed:** the semantics\nmodels complex features of Ruby like\n[`method_missing`](https://noelrappin.com/2023/10/better-know-a-ruby-thing-10-method_missing/)\nand\n[eigenclass reopening](https://suchdevblog.com/lessons/ExplainingRubySingletonClass.html),\nand a substantial theorem about a type system has been written and proven.\n\nI can probably count on my hands the number of lines of actual Lean code I wrote by hand. Frontier AI breezed through a lot of this project, but when it came to the most difficult part of proving a hard property against the semantics, AI struggled, and I intervened by crystallizing the task definition, the theorem statements, and the proof design approach. Admittedly, it was somewhat refreshing to find a task that AI does not immediately excel at, and I came away with some learnings about how to best steer AI on complex tasks like these.\n\nIn sum, my results here illustrate that:\n\n1. Semantics *can* be built quickly (~2.5 months of part-time human labor +\n   agents).\n2. Semantics with AI-authored proofs provide a foundation to scale correctness claims up to all programs in ways that fuzzing or other empirical methods will never be able to.\n\n## `ruby-lean` and its type validator[#](#ruby-lean-and-its-type-validator)\n\nA live demo [playground](https://samx.io/ruby-lean) is available. It is seeded with many\nexample Ruby programs. Play around with it!\n\nHere are the five stages of the program analysis pipeline:\n\n1. Run Sorbet on a program with specific flags, so that Sorbet emits annotation information.\n2. Strip the program's annotations, since Sorbet annotations are real syntax that's outside of what the Lean model supports today.\n3. Desugar[<sup>3</sup>](#tooltip-def-3) the program, producing an s-expression that builds programs\n   from a core set of Ruby primitives.\n4. Run an *untrusted* certificate emitter (a companion program written in Ruby)\n   that proposes a type*derivation* for the program. A derivation is like a\n   conjecture about what types the methods and variables have in a program.\n5. Run the *trusted* validator, given as`validateD p d` in`ruby-lean` 's code.\n   This validates the type derivation against the program. The validator returns\n   true if the derivation accurately types the program according to its internal\n   rules, and we have proven a theorem that demonstrates the soundness of this\n   validation procedure.\n\nHere's how the pieces fit together, and where the trust boundary lies:\n\n```\nflowchart TB\n    subgraph untrusted[\"Untrusted\"]\n        corpus[(\"Corpus of typed<br/>Ruby programs\")]\n        anyprog[/\"Any Ruby program\"/]\n        prog([\"Program p\"])\n        sorbet[\"Sorbet\"]\n        emitter[\"Certificate emitter\"]\n        cruby[\"CRuby 4.0.5\"]\n    end\n\n    subgraph tcb[\"Trusted computing base\"]\n        subgraph tcbruby[\"Ruby\"]\n            rubypad[\" \"]\n            sigstrip[\"sig_strip\"]\n            desugarer[\"Desugarer\"]\n        end\n        subgraph tcblean[\"Lean\"]\n            validator[\"Type Validator<br/>validateD p d\"]\n            proof[\"Soundness proof\"]\n            semantics[\"Semantics<br/>(CESK machine)\"]\n        end\n    end\n\n    corpus --> prog\n    anyprog --> prog\n    prog --> sigstrip\n    sigstrip -->|\"sig_strip(p)\"| desugarer\n    prog --> sorbet\n    sorbet -->|\"type information in p\"| emitter\n    desugarer -->|\"core program\"| validator\n    emitter -->|\"derivation d\"| validator\n    proof -.->|\"proves sound\"| validator\n    proof -.->|\"over\"| semantics\n    cruby -.->|\"differential<br/>testing\"| semantics\n    cruby -.->|\"differential<br/>testing\"| desugarer\n    emitter ~~~ sigstrip\n    sigstrip ~~~ proof\n    validator -->|\"true\"| verdict([\"p has no type errors\"])\n\n    style rubypad fill:none,stroke:none\n    style untrusted stroke-dasharray: 6 4\n    style tcb stroke-width: 3px\n```\n\n## \"Why should I trust this?\"[#](#why-should-i-trust-this)\n\nMy semantics is validated against Ruby by differential testing. I have\nimplemented a multi-pronged conformance suite, where each prong generates test\ncases with a different methodology. This is explained in more detail in the\n[Technical Appendix](2026-09-26-ruby-lean-technical-appendix.html#phase-2-the-semantics).\n\nIf you have doubts, see for yourself: I have released a [playground](https://samx.io/ruby-lean)\nwhere you can run Ruby code in original Ruby alongside the Lean semantics, via\ncompiled artifacts in WebAssembly. The\n[code](https://github.com/sam-xif/ruby-lean) is open-source as well.\n\nI am not pretending that the semantics is perfect. In fact, I recently\ndiscovered, during a differential testing campaign, an instance of\nnon-conformance related to the example Ruby program given in the\n[intro](#introduction-semantics-done-quick). I have not yet patched this, so you\ncan run this program to convince yourself that CRuby and `ruby-lean` are two\ndistinct models of the language:\n\n```\nclass WithToStr\n  def to_str = \"ok\"\nend\n\nputs \"hi\" + WithToStr.new\n\n# ruby-lean: TypeError: no implicit conversion of WithToStr into String\n# CRuby: hiok\n```\n\nThe behavioral difference is that original Ruby automatically coerces the\nright-hand side of a string's `+` via its `to_str` method. The `ruby-lean`\nsemantics has not captured this behavior yet. When this is patched, I'll make a\nnote of it here.\n\nIn the type judgments and the soundness proof, I guarded against\n[proof-slop](https://www.lesswrong.com/posts/rhAPh3YzhPoBNpgHg/lies-damned-lies-and-proofs-formal-methods-are-not-slopless)\nby carefully auditing the end-to-end theorem statement. This theorem, given in\nthe\n[Technical Appendix](2026-09-26-ruby-lean-technical-appendix.html#the-soundness-theorem),\nis an easy-to-interpret statement that directly relates the validator to the\nstuck-freedom property.\n\nThe Lean kernel still has to be trusted, and soundness bugs in it have been\nfound in the past, but I have no reason to believe that agents exploited any of\nthem. I took great care to make sure that the tasks the AIs were given were\nachievable. I did not give them impossible goal statements and gave them\nemergency exits from goal pursuit that they could use if needed (a suggestion\nborrowed from\n[Mike Dodds](https://oath.tech/pub/2026/08/tentative-advice-building-with-ai/)).\n\n## \"Why Ruby?\"[#](#why-ruby)\n\nI chose Ruby for a few reasons:\n\n1. **It isn't Python.** Python has been[treated](https://dl.acm.org/doi/abs/10.1145/2661088.2661101)[rather](https://arxiv.org/abs/1610.08476)[extensively](https://dl.acm.org/doi/abs/10.1145/2544173.2509536) in the\n   literature. It amusingly appears to be a popular topic[for](https://www.researchgate.net/publication/213877472_An_executable_operational_semantics_for_Python)[master's](https://arxiv.org/abs/2109.03139)[theses](https://www.ideals.illinois.edu/items/45257) .[<sup>4</sup>](#tooltip-def-4) Ruby has[some](https://link.springer.com/chapter/10.1007/978-3-319-12736-1_5)[prior](https://www.cs.umd.edu/~mwh/papers/ril.pdf)[art](https://www.cs.umd.edu/projects/PL/druby/papers/druby-oops09.pdf) , and\n   I used it as inspiration for certain parts of the semantics, but in general\n   it seems that Ruby is more \"out of distribution\" for frontier AI than Python.\n2. **It has high-profile users in industry.** Notably,[Stripe](https://stripe.com/) has one of the largest Ruby codebases in the\n   world. Stripe processed[1.6% of global GDP](https://stripe.com/newsroom/news/stripe-2025-update) in\n   2025, so this Ruby codebase can be seen as a critical piece of global\n   infrastructure. Beyond Stripe,[Homebrew](https://brew.sh/) is written in\n   Ruby.[Ruby on Rails](https://rubyonrails.org/) is a mainstay in web\n   development. Shopify is a prominent user of Ruby on Rails, and was funding[academic research on Ruby](https://shopify.engineering/shopify-ruby-at-scale-research-investment) as of several years ago.\n3. **It's complicated.** Ruby has features, absent in Python, that complicate\n   static analysis. The example that immediately comes to mind is that of[*blocks*](https://tech.stonecharioteer.com/posts/2025/ruby-blocks/) . In\n   Ruby, it's possible to pass a block to a function. A block is sort of like a\n   lambda, except it has different scoping rules and multiple ways that it can\n   return. I believe there's a solid chance that Ruby is \"semantics-complete.\"\n   That is, if we can solve Ruby, we can solve semantics for every other\n   language.[<sup>5</sup>](#tooltip-def-5)\n4. **It has a *de facto* type checker.** Sorbet was created at Stripe, and it is\n   used at Stripe and beyond. Sorbet is unsound by construction, allowing`T.unsafe(...)` as an escape hatch. However, we are particularly interested\n   in the maximal*sound fragment* of Sorbet that we can model. Any soundness\n   bug in Sorbet, where it claims a typed program is safe when it is not, would\n   be immediately relevant to Ruby/Sorbet users.\n\n## Cost of verification[#](#cost-of-verification)\n\nFollowing\n[Quinn's](https://www.lesswrong.com/posts/SG82BkTDQDAjANRWj/please-measure-verification-burden)\nadvice, I report the verification burden here.\n\nThe work to obtain this result unfolded over the course of ~2.5 months, at 10–20 hours per week of human labor, with $9,000 of token budget provided by Apart Research.\n\nThe model used to author most of the code was Claude Opus 5. I occasionally used Claude Fable 5 as a consultant, but I shied away from using it for grindy sessions because I found that it burned through tokens far faster than I intended, without much of a speedup in the ladder climb. The final push towards the soundness proof over a nontrivial fragment of Ruby that is presented here was grinded with the new GPT 6 Astra model, which I found did a very good job at a more reasonable cost, but this may also be due to the more rigorous ratchet discipline I imposed in my most recent attempt at growing a sound type system.\n\n## Limitations[#](#limitations)\n\nThis work is limited by the fragments of Ruby semantics and Ruby types that are\ncovered. The semantics has a ways to go before it models the long tail of\nesoteric Ruby features, and the type system still needs to be grown to cover\nfrequently used Ruby constructs like blocks. **In case it is not clear: the\nsemantics and the type system cover two different sets of Ruby constructs.\nBlocks are well supported by the *semantics*, but not yet covered by the *type\nsystem*.** As a rule of thumb, it is much harder to admit a construct to the\ntype system than to the semantics, because of the proof required.\n\nThe semantics also does not model Ruby programs' interaction with the\nsurrounding system context, like the `RubyGems` package manager, the operating\nsystem, or the network. Modeling these interactions will be essential for\nindustrial-grade reasoning.\n\nAside from the limitations of the project as it's currently scoped, there are several interesting open research questions in the science of growing these semantics and steering agents to complete long-horizon proof tasks within them.\n\nSpecifically,\n\n1. I have not developed a theory of how best to steer LLMs to obtain useful semantics. This will require time and funding to run more controlled experiments and benchmarks.\n2. I did not run any comparison between different LLMs/harnesses for the tasks\n   here, except for a brief stint playing around with GLM 5.2 and GLM 5.3, to no\n   avail. I expect that performance will vary significantly with respect to\n   model choice, reasoning effort, choice of agent harness, and prompting\n   strategy. `ruby-lean` provides a nice environment for evaluating how agents\n   can reason about a complex logical system. A benchmark could look like a set\n   of properties about Ruby and its types, where the agent's job is to prove or\n   disprove them.\n\nIf you are interested in working on/funding this project or its future directions, let me know.\n\n## Future work[#](#future-work)\n\nMy immediate next steps are:\n\n1. Grind the semantics to cover any reasonable Ruby program.\n2. Grind the type system enough to run the type validator on a repo in the wild (Homebrew is my initial target). At a minimum, this includes reasoning about the types of blocks, flow sensitivity, and class inheritance with mixins.\n3. Attempt to find real Sorbet unsoundness (i.e., not just unsoundness introduced by deliberate escape hatches).\n\nBeyond these, something I would like to explore more is leveraging the semantics\nfor adversarial synthesis of\n[deserialization attack chains](https://www.elttam.com/blog/ruby-4-0-universal-rce-deserialization-gadget-chain).\nThis is a recurring bug class in Ruby and probably in other languages;\ndeserialization of uncontrolled input is a huge attack surface.\n\n## Learnings and final thoughts[#](#learnings-and-final-thoughts)\n\nWant to build something similar? Take away these learnings so that hopefully you don't have to tear down your work multiple times like I did :).\n\n1. ALWAYS have a clear idea of what you want the agent to do, unless you're deliberately exploring and okay with potentially throwing out whatever the agent gives you.\n2. Having a ratchet discipline, where an agent can only make progress by pushing some metric up, is important.\n3. Do not be wishy-washy in prompts. Be direct about what you're asking for and what the definition of done is.\n4. Clean up the code regularly to remove the buildup of agent-authored cruft. This is something I did not do, and now I'm paying for it.\n5. Deep thinking about the logical structure of semantics (or whatever you're trying to do) is still necessary. Relying too heavily on agents at times made me more confused and biased me towards certain nonsensical ways of thinking about the problem. I made real progress at inflection points where I decided to let go and rebuild my mental model from first principles.\n6. This is a technical detail, but one you can include in your prompts to your\n   agents: try to prove lemmata that allow for *decomposed reasoning* . For\n   example, a continuation stack decomposition lemma proved very useful for\n   expediting several proofs.[<sup>6</sup>](#tooltip-def-6)\n7. Commit regularly and save agent transcripts with `/export` to a journal\n   folder for posterity. I have had agents search them to find context from past\n   discussions, and this has been helpful at times.\n\nThis post is less of an announcement of a finalized piece of work than a check-in on one that is very much in progress, so I'm sure I'll have more learnings to report soon.\n\nUltimately, I think you should take away the elegance of this pattern of using\nan agent to formalize something with respect to some black-box oracle. It works\nquite well, and it can even be applied to specific software systems or libraries\n(e.g., the Ruby on Rails framework) instead of plain programming languages. This\ntechnique is useful anywhere you may want to abstract away the messy details of\nthe implementation and reason about the higher-level behavior.[<sup>7</sup>](#tooltip-def-7) You\njust need to build confidence that the model is faithful by developing a really\nstrong test suite.\n\nFinally, I have simply been blown away by how good LLMs are at most engineering tasks, but they are only as good as the task definition you give them. Sadly, this makes me more afraid than I used to be about the deployment of AI agents at scale on tasks where their performance is not being properly evaluated or not evaluable in the first place...\n\n## Acknowledgements[#](#acknowledgements)\n\nThank you, Eitan Sprejer and the Apart Research team, for your support in this project.\n\nThank you, Mike Dodds, for your mentorship and feedback on this piece. Thank you, Max von Hippel, Victor Arsenescu, and Dana Wensberg for your helpful comments as well.\n\nAnd thank you for reading. Want to get involved? Reach out at\n[s.xifaras999@gmail.com](mailto:s.xifaras999@gmail.com)!\n\n1. \"Program testing can be used to show the presence of bugs, but never to show their absence!\" —Edsger Dijkstra [↩](#tooltip-ref-1)\n2. A fun lightning talk by Gary Bernhardt on the weirdness of the semantics of these two languages can be found [here](https://www.destroyallsoftware.com/talks/wat) .[↩](#tooltip-ref-2)\n3. In the \"syntactic sugar\" sense. Most programming languages can be projected to simpler subsets of themselves. The notion of \"desugaring\" was advocated by Krishnamurthi, Lerner, and Elberty in [*The Next 700 Semantics: A Research Challenge*](https://par.nsf.gov/servlets/purl/10124984) .[↩](#tooltip-ref-3)\n4. I find this especially amusing because I, too, wished to do this for my master's thesis with Pete Manolios, but we eventually steered away from the idea on the premise that it would be too much of a lift without much practical value. We settled on [finding bugs in Python programs with fuzzing informed by their type annotations instead](https://samx.io/papers/thesis.pdf) .[↩](#tooltip-ref-4)\n5. I could be completely wrong here. I invite those who know much more than I do in the field of programming languages to confirm or deny. [↩](#tooltip-ref-5)\n6. The lemma is `run_pushK` in the[Technical Appendix](2026-09-26-ruby-lean-technical-appendix.html#lemmata) , and it reads: running the current machine state under continuation stack K is the same as running the machine state under an*empty* continuation stack until it produces an answer, then delivering that answer to continuation stack K and continuing the run.[↩](#tooltip-ref-6)\n7. The other concern, proving conformance of the implementation to the semantics, can be handled separately, with a clean interface between them. [↩](#tooltip-ref-7)", "url": "https://wpnews.pro/news/ruby-lean-a-ruby-semantics-with-a-type-soundness-proof", "canonical_source": "https://samx.io/blog/topics/devlog/2026-09-26-ruby-lean.html", "published_at": "2026-09-28 02:13:30+00:00", "updated_at": "2026-09-28 02:49:26.191296+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "developer-tools", "ai-agents"], "entities": ["Sam Xif", "ruby-lean", "Lean", "Ruby", "Sorbet", "GitHub", "Mike Dodds"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ruby-lean-a-ruby-semantics-with-a-type-soundness-proof", "markdown": "https://wpnews.pro/news/ruby-lean-a-ruby-semantics-with-a-type-soundness-proof.md", "text": "https://wpnews.pro/news/ruby-lean-a-ruby-semantics-with-a-type-soundness-proof.txt", "jsonld": "https://wpnews.pro/news/ruby-lean-a-ruby-semantics-with-a-type-soundness-proof.jsonld"}}