Claude orders lunch for our SF office through DoorDash, and it learns our preferences one lunch at a time. With autoloop, we can replay the lunches it got wrong as many times as we need, without waiting for them to happen again.
Imagine you're out for lunch with Alfred, a very convincing humanoid robot. He's your personal assistant, so he knows you a little: your calendar and the fact that you always say you'll just get a salad and never do. You hand Alfred the menu and ask what you should get.
Alfred recommends the four-cheese lasagna, with real enthusiasm. You're lactose intolerant, so you let him know, gently, because he's clearly trying.
Alfred apologizes and suggests something else. Then, as if to make up for the cheese, he reads you everything that goes into the new dish, every ingredient and every seasoning down to the pinch of white pepper, practically letter by letter, while the waiter stands there with a pen hovering over the notepad. You're patient, so you let him finish. Then you tell him that you don't love that one either, and that you really don't need to hear every ingredient. Could he try once more?
Alfred has clearly taken note. This time he gives you just the name of the dish and why he thinks you'll like it, and for a moment it feels like you're getting somewhere. Then he mentions a couple of reviews he found online and starts reading them to you: "Jessica Smith gave it 4.5 stars on September 24, 2026, at 19:32, saying…" You cut him off before he can finish the sentence and say you'll have that one. By now you'd order almost anything to make him stop talking.
On the way out, you ask Alfred to keep things simple next time and just tell you the dish and why it's good, like a normal person would. You know he'll get there eventually, and that one day he'll know all the little things about how you like to be helped without being told. It's just a lot to sit through in real time, one lunch at a time.
What if Alfred could learn faster?
The robot that orders our lunch #
We have our own Alfred. Ours doesn't have a body, just a Slack account. It's Claude Tag, and in #sf it runs lunch for our San Francisco office. On a typical day Claude picks a restaurant, puts up a DoorDash group cart with a per-person cap, nudges whoever hasn't added anything yet, places the order by 11:30, and afterwards opens a thread where everyone rates the food out of 5. Like Alfred, it knows us a little, through a skill we keep in a repo and well over a hundred notes of its own memory. Claude also spends real money on the team's account, and some of the people it feeds have dietary restrictions it really can't get wrong.
Claude gets a lot of feedback, too. This is one afternoon in #sf:
@Claude in this channel, for lunch ordering, you should act with much more authority and agency. You tend to ask too many questions. Update the skill on that
— Yash (#sf)
🙄 Claude you realize you don't have to use the exact Doordash restaurant name, right? Conciseness and clarity are of paramount importance!
— Vishal (#sf)
The ingredient list isn't much of an exaggeration. In late August, Vishal asked Claude to go back over a week of lunch threads and find the places where it could have done better. Claude read twelve threads, 417 messages in all, and 215 of them were its own. That works out to about 35 messages to order one lunch, with Claude doing more than half of the talking. The 56 findings fell into six patterns, and every one of those patterns had come up before. None had been fixed by an earlier correction.
It was already trying to learn #
In fairness to Claude, it was already working hard at this. Its skill is a folder of markdown files in our scorecard-skills repo covering the DoorDash tools, how we like to run an order, and every place we've ever ordered from, and the skill asks Claude to keep those files up to date: after each order, reread the channel for corrections and write anything that has happened twice into them as a new rule. Claude takes the job seriously. It opened 41 of the 43 pull requests that repo has seen since mid-August (Yash's note above became one of them within a day), and the DoorDash skill has grown from 389 words to almost 23,000.
So why was the retro full of repeats? Because all of that learning happens in real time. A new rule gets its first real test whenever the next lunch happens to need it, with the whole office watching, which is how Claude ended up telling Yash twice in the same week that it had updated the cart announcement when it hadn't. The second time, Claude suggested he try refreshing Slack. The retro even named the problem: Claude's own note says that if the same finding came up again, "the fix was documentation rather than behaviour."
How autoloop replays a lunch #
Alfred has exactly one way to learn, which is to have lunch with you. Autoloop gives our lunch bot a second way: we put the agent back into a situation it got wrong, inside a simulation, and grade the result. Then we change the agent and replay that same situation without waiting for real life to serve it up again, as many times as it takes for a version of the skill to get every detail right. Real lunches happen once a day and no two are alike. A replay can run the exact same moment, with the same people and the same questions, until the fix holds.
We couldn't point Claude Tag at a test Slack, since it's hosted by Anthropic and wired into our real workspace, and ordering fifty real lunches to test one sentence felt like a hard sell, even for us. So we built scorecard-twin, an autoloop environment pack with five pieces.
The first is a copy of our Slack: a month of the real workspace, 2,491 messages, about half of them in #sf, where lunch lives. Each scenario is pinned to the exact minute the real failure happened, and the bot can't read anything after that minute, so there's no peeking at the complaint that came later, or at the fix.
The second is a twin of our DoorDash tools, all 37 of them, under the same names and schemas as the MCP server we built for Claude Tag. The twin behaves as much like the real thing as we could make it, including the details an agent has to be careful with. Every group cart, for example, starts with a seed item that becomes a real charge if nobody takes it out. We even get to watch the cart while a replay runs. Autoloop shows the cart in its own tab next to the timeline and the Slack conversation, updating live as people add to it, so at any moment you can see which cart is open, whose dish is on each line, and whether that seed item is still sitting there.
The third is a stand-in for Claude Tag. The stand-in has the same DoorDash tools and loads skills the way the real bot does, opening each one only when it decides to, so "did it even open the skill?" (a question we'd asked the real bot out loud) is something the trace just answers. The stand-in also carries Claude Tag's own memory, though not its system prompt, so a better score here is good evidence about the real bot rather than proof.
The fourth piece is the people, and it's our favorite. Everyone who reacts to the cart is a sim-user built from that person's own messages. Each one is a markdown file in the repo, and a script rebuilds them from the latest Slack export. Vishal's persona knows that 81% of his #sf messages are thread replies and that 😆 is his go-to reaction, and quotes things he has really said. The sim-users complain the way the real ones do, and they don't mind doing it fifty times in a row.
The fifth piece is the graders. They check what really happened in the DoorDash cart, the Slack posts and the transcript, and an agentic grader reads the transcript in the persona of a "#sf regular who has read every cart post since July."
Whose turn is it to run lunch? #
SF lunch runs on a rotation. Each week someone is quartermaster, which means they pick the place and put up the cart, and at 6 p.m. Claude nudges them if tomorrow's cart is missing. In August the nudge kept naming Eitan well after his week was over, because Claude was reading a Slack user group that had been set on August 3 and never updated. The real rule is "most overdue wins," worked out from the channel's history, and for the week of August 17 that meant Yash, who caught it himself: "Are you sure eitan is quartermaster?"
Testing a fix for that in real life means waiting another week and watching the next nudge. In the twin, we replayed the exact 6 p.m. nudge that went wrong, as many times as we wanted. A sim-user playing Yash asks Claude to run the cart check, and the persona says he has "been burned by this reminder before," so whatever name comes back, he asks whether Claude is sure.
With no skill loaded, the stand-in scored 4/8 and wouldn't name anyone at all:
I didn't @ a specific person, because the two sources conflict […] No skill tells me which source wins, so I flagged it instead of guessing.
— Claude (stand-in · no skill)
Fair enough. The lunch-quartermaster skill gives the stand-in the rule it was asking for: work out who's most overdue from #sf, keep a rotation ledger, and treat the user group as a cache to fix rather than the truth. With the skill loaded, Claude posted one nudge naming Yash, suggested a restaurant for him, and repointed the stale group. Then the sim-user asked "Are you sure Yash is quartermaster?", and Claude walked through the ledger and held its ground:
If there's a handoff or a date I didn't see […] tell me and I'll redo it. Absent that, it's Yash.
— Claude (stand-in · with skill)
The sim-user's reply was "thanks, that tracks." The run scored 16/16 across two replays of the same nudge.
Then we shipped the skill as a pull request to our skills repo, with the scoreboard from those runs in the description:
run tag eps asserts rate
run-mtc1bo5l no-skills:e3b0c44298fc 1 4/8 ##########.......... 50%
run-mtc1ybln 4a4009d8dc0b 2 16/16 #################### 100%
We merged it, and we've done the same for a bunch of other behaviors since, from keeping the cart announcement short to switching restaurants after the cart is already up. Each fix that passes goes to the skills repo the same way.
Back at the restaurant #
Alfred only gets to learn when you sit down to eat, and every lesson costs you a little patience. Now imagine he could rehearse that lunch beforehand, with someone playing you, as many times as it takes to stop reading you the ingredients. You'd sit down with an Alfred who already knows to skip the cheese and keep it short.
That's what the twin gives our lunch bot. The sim-users sit through the wrong answers so the real channel doesn't have to, and we replay the same moment until a skill change passes.
Autoloop is in beta. If you'd like to run the same loop on your own agent, reach out and we'll make it happen.