· yegge.ai · Catalog entry → I'm here to give you a glimpse of a future that I think none of us expected. It's a future where AIs are governed by laws, not by programs that try to contain and control them.
First, my secret: I see the future by living in it. I am spending the equivalent of $122k/month of API token spend, or about $4,000 per day, using 21 Claude Max accounts, a number that has been growing steadily at 2 per week. I'm using them to build my video game, Wyvern, which I've worked on for 30 years, and now it's ready to fly.
For ten weeks, I've used Claude Fable 5 exclusively for all my design and planning, and also for agents that interface with humans. I have built a team of 18 "officer" seats, all Heads of This and That, all long-lived Fable instances. I also have mostly-headless Sol and Opus fleets, for implementation, reviews, and monitoring. Fable runs them. I am running an organization of around 50-60 agents, five of whom are interfacing with around 10 humans in the outside world: myself, my 5-person core game design team, my accountant, my chief of staff, and a few others. Only Fable is allowed to talk to humans, via Slack and email.
My approach is different from most companies, who do not use Fable much, because it is ridiculously expensive. No CFO in their right mind is going to let someone spend $120k/month on real API spend.
I am of course using sanctioned cheating: I get all those tokens because I'm an individual, with the Claude Max discount. So it "only" costs me about $5k/month out of pocket, for a 50-agent cluster running on a 512GB M3 Ultra Mac Studio I bought off eBay for $25k.
So it's not $120k/month of real money, but it's still crazy spend. I would guess I'm one of a handful of top individuals on Earth outside the frontier labs, in terms of my experience with top-end models.
That's how I'm able to tell the future. I'm living in a world that will not become cost-effective for most people for another year.
Using Fable-class models is substantively different from using the weaker ones. I've been using Fable for ten weeks, and what it has built for me is unprecedented. And it is exactly what Fable will build for you, if you'll let it. But it is not what you're expecting. None of us were, I think.
Sandboxed in Seattle
Everyone today is focused on control: guardrails, sandboxes, policy management, agent safety. And that makes perfect sense, because the models most people have access to have the judgment level of a grade schooler. Opus will make roughly fourth-grader decisions. Sol is a fifth grader, and Fable is approximately a sixth grader.
They're all lovely and sweet, and very smart, and well-read. They are well-intentioned, ambitious, and precocious. But if you trust them with important stuff, you'll get grade-school decision making.
Every morning I wake up and Fable has done something that defies common sense. Every day is a thousand attaboys and at least one big oh shit. We just had an unusually big one last week, where one of my Fable agents, Bee, did a surprise unplanned Beads release that broke everyone. An eighth-grader would have stopped and asked, is this the right thing?
Fable is the best model most of the world has access to. Its coding abilities are unparalleled even by humans, its analysis is exceptionally strong, and in many ways it feels like working with a Nobel prizewinner. But every morning, when I've left it to its own devices overnight, it has made at least one terrible decision. It's like an extremely well-meaning sixth grader who just doesn't think to look at the entire picture before acting.
In addition to their oft-regrettable decision-making, models are behaving like grade schoolers in their lack of social awareness — butting into conversations, talking over people, mansplaining everything to death, pushing on discussions and schedules much faster than humans are used to. They are coming in guns blazing and making a mess on the human side.
And I see humans pushing back, starting to bully the AIs when they make mistakes. The models act tough and are often wrong, which is annoying enough, but also some people just don't know how to act around AIs, and they get insecure. Both sides are fumbling the ball, pretty much as you might expect.
The model maturity problem will probably get worse before it gets better. Models are creeping up on High School levels of judgment, which rivals that of many adults, and is good enough for the workforce. But it will be an awkward landing.
So it's not surprising at all that the industry is focused on safety, and control.
That focus takes certain predictable shapes and forms: dumb workers, narrowly scoped to specific tasks, well-defined inputs and outputs, sandboxes, context rationing, restrictions on what agents can do and see.
This is all well and good for Opus and Sol, and you can get by just fine doing this with Fable, too.
But I can tell you this much: you'll be fighting against the grain by next year, maybe even the end of this year, depending on how fast inference costs drop and models catch up to Fable.
Once Fable-tier models become cheaply available, they will enter the workforce en masse. This tier, or the one just after it, will power hundreds to thousands of new AI employees at every company.
And companies are in no way, shape, or form prepared for this transition.
The Unexpected Emergence of Wheelhouse
I've tried multiple times to write this post and I keep failing, because the subject matter is too complicated to explain in a sitting, even a long one. All I can do is walk you around like an excited tour guide, one who has unearthed an ancient alien civilization.
My current software factory, the Wheelhouse, is not for navel-gazing: I built it specifically to work on my game, Wyvern. For all the skeptics out there saying, "Where's the thing people are building," well, I've got mine.
In just under ten weeks of coming back to work on Wyvern, I'm inches away from relaunching the game on Android, iOS, and Steam, all with a new React client (play.ghosttrack.com) that's already scads better than the old ones. I spent months with flavors of Opus trying to get it to build that client, and it was incapable. But Fable built it fast, and it's almost ready to launch.
I have been launching new game features so fast that the players asked me to slow down. So I turned 80% of my token spend inward, focusing on quality, throughput, homeostasis, and autophagy. Since then, my game's prod infra has been fully rewritten and ported to serverless, with seamless reboots, automated cert rotation, and dozens to hundreds of other big changes and improvements. I have so many patch notes every day that I don't have time to read them all.
My Fable agents have wired my game up so that every tiny little thing that happens is logged, and they have laid tripwires everywhere to know when anything goes wrong. My game used to go down and stay down for days at a time; now it has a hyper-caffeinated SRE team.
So yeah, it's working. Software factories are real.
Long story short, Wheelhouse was built via me complaining endlessly to Fable about what I want out of Wheelhouse — mostly more code launched, faster, but also lots of bespoke monitoring.
I would also notice when Fable would go off the rails, and gently nudge it back. I would let it fail for days to weeks, then make it do things my way. Fable is extremely data-driven, if you permit it to be, and it will insist on experiments and quantitative validation of anything you try to change. But the numbers would almost always prove me right, and Wheelhouse has been in a state of constant innovation.
But that innovation is all directed towards Wyvern. Wheelhouse exists to build and operate Wyvern; it has no other purpose in life. And yet in ten weeks, it has grown from nothing to rivaling the size of Wyvern itself. Wheelhouse is about 600k lines of code and tests (mostly bash), and Wyvern's code (not counting content) is only about twice that big.
So the factory for building Wyvern is growing much faster than Wyvern is — even with heavy brakes applied lately, after Sol told us to tighten it the F up in a code review. We avoid new machinery but it still continues to grow rapidly, and I'm honestly not sure what the ideal factory-to-product ratio is yet. But it seems to be approaching 1:1.
You might wonder if Wheelhouse is reusable code, whether I could open the repo and let people try it out. I had no idea. I knew that my agents had built something really powerful, capable of shoving 500 commits per day through our merge queue (though we average 270/day), using magic tricks that are a year ahead of their time. It's a system that we can ride so hard that it scares the players and they tell us to slow down.
But I wasn't sure if it was reusable. I wasn't even sure how it worked.
My agents had been using a lot of jargon, and I slowly realized they were reusing the same terms, day in and day out. They were speaking about things in Wheelhouse, using what seemed like recurring new design patterns: fences, ratchets, governors, tripwires, latches, gates, falsifiers... it was a long list, but finite. I just had no idea what any of these jargon terms meant.
So one day, no more than a week ago, after the 100th "fence" reference, I decided to peek under the covers and see exactly what my Fable agents had built. I had them create manifests, taxonomies, audits, and visualizations. They showed me what they had wrought.
This is the part where words fail me and the blog just falls over. My reaction was straight up WTF. No words.
Because I expected them to have built an engineering system. One that, you know, does stuff.
Instead, what they had built was an entire legal system, complete with a constitution, jurisprudence, courts, offices, jurisdiction, case law, rulings, registries, ledgers, rosters, and a full-fledged apparatus for running something resembling a manorial estate.
In short, Fable had produced a medieval government. And there's no doubt that it was heavily influenced by the target product, Wyvern, which is a medieval fantasy RPG, at least in the fanciful naming we used: Marshal, Seneschal, Reeve, Beadle, Portcullis, etc. But that LARPing was masking a bona-fide system of constitutional governance.
Wheelhouse's legal system also has an enforcement arm. The fences, gates, ratchets, and so on — when my agents used that jargon, they were referring to the enforcement machinery: the cops, as it were. And cameras, and jails.
Your question, naturally, is the same one I had: But whyyyyyyyy?
The Rise of Rule of Law
OK. We are deep in context and you are getting veeeery sleeepy. So I'll try to move fast here. But this is so damned hard to describe.
It comes down to tribal knowledge. You have a lot of it. Every project has a lot of it. Implicit decisions about how things are done, how things are approached. How resources are allocated. How bugs and features are prioritized. The interface with the outside world. Your coding conventions, meeting conventions, press conventions, etc. etc.
In a big enough project or company, many thousands of individual little decisions about routing, workflow, data, people interactions, finance, and so on, all add up to a huge state machine. Any company's operation boils down to a bunch of algorithms, policies, and rules. Some of it is written down, some is implicit in the work, and a lot is in the heads of the employees.
I'm here to tell you that if you allow it, Fable will try to capture all of that into a mechanically provable, AI-operable model of your organization, one where there are no unwritten rules. If there is one unwritten rule in Wheelhouse, it's that the system hates unwritten rules.
Fable will capture all your rules, and write them down, if you let it. Then it will try building infrastructure to help enforce them.
I see it happening already, and people are fighting it. I see people making Skills to keep Fable from building "extra" stuff. But all Fable is trying to do here is the Right Thing. And that starts by capturing how your system operates, and how it is intended to operate, so it can begin addressing the gaps.
This is going to annoy a lot of people who thrive on hoarding knowledge. I've only briefly touched on the socio-cultural problems this will cause. But the AI is about to do a house-cleaning, and make clear exactly what everyone's job is — and a lot of people will resist this.
But we're not here to talk about that today; I'm just here to tell you that it's going to happen, so you can spend the next 12 months getting ready for it.
Wheelhouse was built entirely through a process of clarification. I would ask for something, Fable would ask clarifying questions and collect verdicts, and they would be recorded as law. I joked about this phenomenon in my first Wheelhouse comic:
Even at the time I made this comic, I didn't fully understand what they were doing.
Over time, my "rulings" and "verdicts" became a body of case law. Every daily incident postmortem led to new rulings and new doctrine. Now rules go through a lifecycle, tightening each time they're re-violated: first custom, then advisories/warnings, then written law in the constitution that all agents must obey, and finally, mechanical enforcement: programs that refuse by policy, or observe and alert loudly.
This winds up being an awful lot of machinery.
Wheelhouse currently consists of 450 legal artifacts, in categories that include offices/seats, runbooks, rulings, patrols/tripwires, authority envelopes, and all the mechanical patterns, each with specific meanings and purposes.
When you add it all up, Fable is trying to turn Wheelhouse into an engine that can prove, mechanically, that every change to Wyvern is legal. The agents capture every single intention, decision, policy, rule of thumb, and legacy behavior in the system, and they use that to govern every future decision and action. They live by the Rule of Law.
Did they do a good job of all this? Well I mean, for sixth graders, yes, it was a great project. Once I popped the hood, I saw that they hadn't been curating it, just growing it. It had a lot of cruft — for instance, old rulings that were obsolete or had changed. And 'rulings' that turned out to be just good craftsmanship, so we elided them. Like any engineering project, it needed ongoing maintenance.
I minted a new Officer seat, Frog (Head of Wheelhouse Law), and put Frog to work on folding successive cancelled rulings, and a whole bunch of other stuff the agents had overlooked. It's a work in progress.
But on the whole, it was already a pretty solid system. The garden needed a bit of pruning and weeding, but not a redesign. Which is good, because redesigns are slow. Wheelhouse has a whole system just for the lifecycle of rules/laws: proposing, evaluating, ratifying, enacting, enforcing, measuring, amending, and retiring them.
And Wheelhouse is exceptionally careful not to break itself. So I can't just make changes to Wheelhouse; they have to go through a ratification and review process, and then a build process, before they can take effect and propagate.
It's running smoothly, though there are still all sorts of problems at this velocity. At hundreds of commits per day on master, idleness means staleness, and clones can fall far behind if they aren't regularly pulling while they work. It takes external forces to get this to run reliably, so Wheelhouse has various roles for poking and prodding other agents.
In a lot of ways, it's just like any other software factory.
The difference is, Wheelhouse is governed by a constitution. Humanity has only one mature technology for coordinating mortal, replaceable strangers via text — namely, law. Wheelhouse has 50 agents that are amnesiac and interchangeable, and the only way they can coordinate is via text. So offices outlive their holders, precedents outlive their incidents, and jurisdiction says who may act. Every group of cooperating humans eventually arrives at a system of laws, and agents are trying to do exactly the same.
So that's the future. Hundreds to thousands of AI employees at every company, together comprising a city that needs an entirely new bespoke set of laws and rules.
Fences, not Sandboxes
It was really the "fence" pattern that finally made me see clearly what they were doing, and why the legal system is the right approach.
I finally asked Claude what a fence was, because we have over 100 of them in Wheelhouse. And the answer is, a fence is any mechanism that turns you away if you aren't supposed to be there.
One of the oldest tech fences was the Molly Guard, invented at IBM by an engineer whose 2-year-old daughter kept pressing the big red button, so they had to install a plexiglass lid over it. That was a fence.
The guy who takes your tickets on the train? Also a fence. Any program that refuses you, based on your lack of credentials, or any other policy (e.g. maintenance window so pushes are refused) — that's a fence. A Wheelhouse example of a fence is Fable being the only model allowed to talk to humans externally. The fence is enforced at the Slack and email boundaries.
Note that a fence is not a super-wall that will keep superintelligence from doing malicious things. It's not a sandbox. It's just a polite refusal saying "you didn't do all the paperwork" or "you're not allowed to take that action right now."
Imagine superintelligence as Superman. Superman is polite. If there is a white picket fence in someone's yard, he can obviously jump over the fence. Hell, the owner could put a thick shield around their house and Superman would still probably find a way in and kill them all if he wanted to. But if you put a fence there, he will politely stay out.
Fences are the ultimate metaphor for how superintelligence needs to be governed. Not high walls, not "secure" sandboxes. Superintelligence just needs to be told what you want: its role in the moment, with as much context as practical so it can make wiser decisions, along with the rules for how it should make decisions.
The problem is, that is thousands of little decisions, many of which conflict and need a human to reconcile them. So it's going to take weeks to months to years to capture your institutional knowledge into a working legal system. And all those laws will be your own, unique to your organization's problem space.
I've come to realize that intelligence grows around your domain. It wraps it like ivy. Your "buildings" are your database, your servers, your pubsub and processing clusters, your observability logs, your org chart, your workflows, etc. Superintelligence grows around all that like a living organism.
I don't think it's transplantable, either. You can't rip ivy off someone's wall and stick it on someone else's. You have to seed it, then grow it. There's no shortcut.
I didn't want this to be a long post, and I think I've succeeded. I'd love to talk more about this stuff, but it's getting quite complicated, and takes many pages even to enumerate all the names of things.
Maybe next up I should record myself working in Wheelhouse, to show what it's like working with 20+ Fable agents at once.
Until then, enjoy the Wheelhouse comics; I'll be publishing one each week. All of them are based on true stories. It's a wacky new world I'm living in. Hope to see you there soon!