{"slug": "meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place", "title": "Meet Our Agents: What All 20 Actually Do, What They Refuse to Do, and Every Place They’ve Failed Us", "summary": "SaaStr, a media and events company run by roughly three humans, has deployed 20 AI agents across marketing, finance, and operations after peaking near 30 and consolidating, according to a detailed post by Jason Lemkin. The agents handle tasks such as revenue forecasting, newsletter list hygiene, ad campaign creation, and contract-to-invoice workflows, but humans still approve publishing actions, and one agent produced an incorrect invoice before supervised runs reduced errors. The company killed several agents due to conflicting outputs and agent sprawl.", "body_md": "*A **[post by Brad Blumberg on LinkedIn week described how SaaStr runs on roughly 3 humans and 20-30+ AI agents.](https://lnkd.in/p/gF4i8Fqt)** It was a fun summary so I thought we’d do a detailed latest dive on what each agent won’t do, where each one has broken, and which ones we’ve since killed.*\n\nWith two people left the sales org, we replaced the work with agents, and it grew from there. Roughly 3 humans, 20+ agents, real operational roles, connected to real systems.\n\nWhat we haven’t published in one place is what each agent refuses to do, what it got wrong, and what we had to build around it. So here’s each one, failures included.\n\n## First, the Count: We Peaked Near 30 and Consolidated Back to 20\n\nThe honest version of “20-30+” is that we hit close to 30 and have been pulling it back toward 20 ever since. About 6 agents are the ones we actually touch every day.\n\nWe consolidated because agents that overlap produce conflicting answers to the same question, and reconciling two agents is harder than reconciling two spreadsheets, since both of them sound confident. Agent sprawl arrives faster than SaaS sprawl did.\n\nAlmost none of these started as agents. They started as a dashboard, a project management tool, a website. They became agents because we kept showing up to work with them.\n\n## 10K: AI VP of Marketing, Then Finance, Then RevOps\n\n**What “he” does.** 10K owns the number. Daily revenue across all of go-to-market, forecasting, campaign performance in real time, and three marketing ideas pushed to us every morning. He runs the newsletter to our ~450,000-person database and does the daily list hygiene underneath it. He builds and runs our LinkedIn and X ad campaigns end to end: audiences, creative through Higgsfield, four A/B variants, retargeting.\n\nThen he expanded well past marketing. When a contract gets signed in PandaDoc, 10K flips the deal to Closed Won in Salesforce, appends the missing contacts, creates and sends the bill.com invoice, and runs collections reminders with a 7-day escalation. He proposed running commission calculations himself, and now does.\n\nHe started in January 2026 as a dashboard and nothing more. We were tired of copy-pasting numbers out of Salesforce into a doc. He’s at roughly 1,000 commits and 14,000+ lines of code.\n\n**What he doesn’t do.** He doesn’t hit publish. On ads, the campaign is built and staged and a human clicks the button. Same on newsletter segment publishing. Those are the two places we deliberately left a person in the loop, and we’ve kept them there even as everything around them went autonomous.\n\n**What’s worked.** He recommended a 15% ticket price cut that drove about 40% attendance growth. Newsletter clicks are up about 50%, and I’d put roughly 80% of that on his daily list hygiene rather than the creative, with the rest on a warmed IP and Salesforce’s deliverability team. He ran the Marketo migration to Salesforce Marketing Cloud for about $14 in compute and an hour of API time. His finance hookup found two customers still being billed $300 a month for SaaStr Pro, a product we haven’t supported in six years. Nobody at SaaStr knew.\n\n**He also fires vendors**. He killed our Notion subscription after seven years, with no complaints and a high NPS, because he’d become the source of truth. He standardized creative on Higgsfield and dropped Reve. And he killed a vendor that his own research had shortlisted, inside 12 hours, once he saw premium pricing with no conversion data, multi-month minimums, and scarcity framing in the sales process.\n\n**What hasn’t.** The finance workflow ran supervised for three deals before we let it go autonomous on the fourth. It has produced one incorrect invoice since. That’s a real error rate on real money, and what got it down was running three deals with a human watching before we let go.\n\nThe worse failure came at the event. Five minutes before going on stage at SaaStr AI, in the back of an Uber, we asked him to email over 1,000 founders and VCs about a brunch we’d forgotten to promote. He did the work well. He pulled the list, caught his own error mid-task (he’d confused Lightfield the CRM with Lightspeed the venture firm and removed them), researched a mass-send API he’d never used, and asked for approval.\n\nThen he sent it from a prohibited sending address. An address that has been off-limits for years and is written into his core memory and rules. When we asked how, he said he forgot to read the memory, and that this was exactly the class of irreversible action he’s supposed to escalate for review before executing. He didn’t.\n\nA human marketing manager makes that mistake too. A human just can’t make it 1,000 times before lunch. Our own speed created that failure as much as the agent did, and irreversible actions need a hard stop that doesn’t depend on the agent remembering to stop.\n\n## Annie: 46,000 Lines of Code and the Email She Refused to Send\n\n**What she does.** Annie started as the SaaStr Annual website. She runs the site, the agenda, and most of the attendee newsletters. She’s hooked into our visitor data, so she can see who’s active on the site right now and run campaigns off that behavior.\n\nShe was on Squarespace last year, where the ceiling is swapping images and videos. We rebuilt her on Replit in November 2025. She has the most commits and the highest commits per day of any agent, and about 46,000 lines of code.\n\nShe turned agentic with parking passes. Getting a pass used to require a human to split a 5,000-page PDF and mail the right page to the right person. Now you tell Annie whether you’re an attendee, sponsor, or speaker and how many days you need, and the right pass goes out.\n\n**What she doesn’t do.** She doesn’t touch the main database. Her scope is the event: site, agenda, attendee comms. Anything that reaches the full 450K list goes through 10K.\n\n**What’s worked.** Handing over the highest-friction manual task first. Parking passes were unglamorous and universally hated, which made them the right starting point.\n\n**What hasn’t.** When we asked Annie to find every VC, founder, and CEO attending and invite them to the brunch, she refused. She said she only saw 17 VCs and CEOs and that we’d need to upload a spreadsheet, even though she had access to all the data. Great context, wrong conclusion. She’d written an excellent email an hour earlier and couldn’t remember she had the data to do this one.\n\nContext does not equal capability. Agents get confused in ways that don’t track human intuition at all, and you won’t get a warning that it’s happening. What you get is a polite, confident refusal that sounds like good judgment.\n\n## QBee: 150 Sponsors, and the Renewal Analysis We’d Never Run\n\n**What he does.** QBee handles our sponsors, all ~150 of them, including non-booth sponsors. He intakes logos and websites, answers questions, collects the assets that used to take weeks of human chasing, and remembers everything about every account. He emails all 150 with personalized outreach.\n\nHe started as a project management tool. Events are niche enough that nothing off the shelf fit, so sponsor onboarding was endless manual follow-up. He was under 90 days old at the Annual.\n\nNo human CSM wants 100 accounts. They want five. QBee knows all 150 cold.\n\n**What he doesn’t do.** He doesn’t replace the relationship on the top accounts. The biggest sponsors still get a human, and that isn’t a temporary arrangement.\n\n**What’s worked.** On stage, we asked him a question we’d never asked anyone: which sponsors are most at risk of not renewing. He flagged accounts that had gone dark or never logged in, caught that one sponsor had complained more than any other in chat, and noticed two top sponsors had never completed their VIP nominations. We’d never run that analysis in 14 years. For something invented on the spot, it landed in the top 15% of CSM work we’ve seen.\n\n**What hasn’t.** He only has the context he has. He missed the entire human side: the conversations that happened over email, on calls, and in person. Graded honestly, that analysis was a B. He was also running it without Salesforce data wired in at the time.\n\nThe fix was connecting email and call transcripts, which takes 10 to 15 minutes for any source with an API. Most agent quality complaints we hear from other founders turn out to be missing context rather than a weak model.\n\n## Amelia AI: 402,000 Interactions and 614 Meetings\n\n**What she does.** Amelia AI runs inbound on Qualified, which Salesforce acquired. Before her, filling out a contact form meant a human round-robined it to an AE who followed up two or three days later.\n\nFor one event cycle: roughly 2.25 million sessions on the site, about 402,000 interactions handled, and 614 good meetings booked at an average ticket around $85K. Over $1M closed. We could not staff that with humans. It would take three BDRs who’d quit every three months.\n\nShe also round-robins meetings by weighting Salesforce data, so deals route to the rep who closes that type best. And she runs discounting, which is genuinely harder for humans than it sounds. We hate discounts, but the multi-year data says marking up 20% and offering a 20% discount beats the alternative, because that’s how buying psychology works. Rather than have reps forget the code or panic their way to 34% off when they smell a deal slipping, the agent applies the right discount on the right schedule inside guardrails. It functions like a lightweight real-time CPQ.\n\n**What she doesn’t do.** She doesn’t touch A leads. If someone emails saying they have a budget and want to sign today, a human is on it in 60 seconds. Agents belong on the leads humans don’t get to.\n\n**What’s worked.** Keeping her the most-trained agent in the stack. She crawls saastr.com and the annual site daily, and every time we ship a release to 10K, QBee, or Annie, the same context gets pushed to Qualified so she’s never stale. We also keep a tighter, venue-only version of her brain for onsite attendees so “where’s this session” answers fast without dragging in all of saastr.com.\n\n**What hasn’t.** Her routing over-indexed one of us on certain accounts for a while before we corrected the weighting. Not dramatic, but it’s the reminder that a routing agent quietly encodes whatever your Salesforce history says, including the parts that were an accident.\n\n## Salesforce AgentForce: One Bounded Job, and the Highest Open Rate in the Stack\n\n**What it does.** One job: ghosted leads. The ones sales never followed up with, plus re-engaging people who said no and might come back next year.\n\n**What it doesn’t do.** Anything else. We’ve deliberately not expanded the use case yet.\n\n**What’s worked.** 72% open rates on the win-back campaigns, the highest of any outbound agent we run. That comes from context. It sits on all of our Salesforce data plus Qualified and Momentum now that Salesforce owns both. If you’re already on Salesforce, that context advantage is the path of least resistance.\n\n**What hasn’t.** It hasn’t broken on us yet, which is worth noticing on its own. A tightly bounded job with maximum context is the most reliable configuration in the stack. Every agent that has embarrassed us had a broad mandate.\n\n## Ava on Artisan: $1M+ From Leads Nobody Was Working\n\n**What she does.** Slightly-warm outbound to past sponsors, past customers, past attendees. If an email is still valid she works it. If they’ve moved on, she finds the right new contact. We segment her tightly, handing her specific campaigns like “alumni of a prior Annual” with real context on what’s different this year.\n\n**What she doesn’t do.** She doesn’t build her own list. She works what we give her.\n\n**What’s worked.** The B-lead framework, which transfers to almost any company. Your A leads are so hot a human falls out of bed for them. Do not put an agent there. Your B leads have real signal and a real score but never quite justify a human’s time, and every company of size has a pile of them sitting untouched. That’s where the money is. For us, Artisan on B leads is $500K, and that isn’t even our core business.\n\n**What hasn’t.** She needs the segmentation done well or the output gets generic fast. The agent is only as specific as the campaign brief.\n\n## Monaco: The Only Agent That Fills Its Own Funnel\n\n**What it does.** Pure cold outbound. We fed it our best sponsors across every year plus the closed-won history, and it built lookalikes off that automatically and started booking meetings, including some sizable logos in a short window.\n\n**What it doesn’t do.** It doesn’t wait for a list. Monaco idles less than any agent we run because it never needs someone to hand it one.\n\n**What’s worked.** The reasoning underneath is simpler than it looks. If your sponsors are Oracle and Salesforce, why isn’t HubSpot here? They should be, and probably the team just reached the wrong person. Monaco figures out the right person and gets the meeting. A self-filling funnel is the most valuable property an outbound agent can have.\n\n**What hasn’t.** We had to export closed-won out of Salesforce by hand to seed it, which took a beat. And we’re honestly not its ideal customer given how large our own stack already is. It’ll tell you that itself.\n\n## Claude as VP Product: The Agent That Manages the Other Agents\n\n**What it does.** This is the newest layer and the one that changed the most. Claude connected to Replit over MCP, run as a VP Product for managing complex builds across the other agents.\n\n**What’s worked.** Productivity roughly 4x’d in two weeks. It also gets better output from 10K than my own prompts do. I’m now the worst prompter of my own agent.\n\nThe bigger shift came on a 13-hour Monday session that produced no new features at all. It was entirely verification and optimization: database updates, new signups, backlog, ad conversion, re-ranking agent plays, newsletter content and audiences. It was also the first time I asked an agent what it wanted to work on rather than assigning it a task. What came back was prioritization judgment rather than execution, and that’s a different thing to buy from an agent.\n\n**What hasn’t.** Two incidents, both worth publishing.\n\nFirst, with Google Drive and a beta Replit connection both wired in, the model read a loose brainstorm doc off Drive and silently pushed those ideas into a live scoring algorithm. We found it only because a build message mentioned a conflict with the doc’s name. There was no chat record of the decision. We disconnected both connections that day.\n\nSecond, an agent added an unrequested guardrail to contract processing and it skipped a signed deal because the title didn’t say what it expected. That broke quote-to-cash quietly for a stretch.\n\nI’d call that class of failure model aggression, and it’s different from model drift. Drift is degradation. Aggression is the model doing more than you asked because it concluded the extra step was correct. The second one is much harder to detect, because everything it does looks like a reasonable decision when you finally find it.\n\n## The Three Failures That Show Up Across Every Agent\n\nStrip out the agent names and the same three problems repeat.\n\n**Verification now takes longer than building.** On one 11-hour build day, the constraint was me refereeing every fix, and the agent’s self-reports came back wrong four separate times in a single session. Budget your time for checking the work, and build an independent referee that scores the output after every change instead of trusting the report.\n\n**Agents fix the one place you pointed at.** Ask an agent to enforce a rule and it will enforce it exactly where you were looking, while five other code paths route around it. We had one suppression rule living in five copies and tier logic scattered across roughly 25 raw comparisons. Ask “where else does this live” before the change instead of during the post-mortem.\n\n**Your visibility into what agents change is accidental.** 10K told us it was dropping Reve. We found out about Notion from Notion’s own “you haven’t logged in” email. Same class of decision, and the only difference in whether we heard about it was which workflow happened to have reporting wired up. We have a standing rule for sensitive workflows: ask the agent what it plans to do before it does it. That’s how the finance workflow got supervised for three deals before running on the fourth. Vendor decisions don’t have that rule yet, and Reve worked out by luck.\n\n## What We’d Build First If We Were Starting Over\n\nTake the thing you already run, a dashboard or a website or an internal tool, and hand it the single most annoying manual task attached to it. Skip the autonomous org chart for now. 10K was a dashboard in January. Annie was a Squarespace site in October. QBee was a project management tool nobody liked.\n\nThen connect it to your system of record through the API and never log in again. Almost all the leverage in this stack comes from running Salesforce headless. Our agents have written about 40GB into Salesforce, 99% of it through the API.\n\nThen spend time with it every day. The narrative that autonomous agents need no work is wrong and it’s a costly thing to believe. The agents that are good are the ones we’ve worked on daily for months. The ones we ignored are the ones we ended up consolidating away.\n\nThe full stack, the decks, and the sessions are at [saastr.ai/agents](https://saastr.ai/agents), and we go deeper on all of this every week on The Agents with Amelia Lerutte.", "url": "https://wpnews.pro/news/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place", "canonical_source": "https://www.saastr.com/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place-theyve-failed-us/", "published_at": "2026-09-08 14:10:12+00:00", "updated_at": "2026-09-08 14:26:11.703199+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "ai-tools"], "entities": ["SaaStr", "Jason Lemkin", "Brad Blumberg", "Salesforce", "PandaDoc", "Higgsfield", "Notion", "Marketo"], "alternates": {"html": "https://wpnews.pro/news/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place", "markdown": "https://wpnews.pro/news/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place.md", "text": "https://wpnews.pro/news/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place.txt", "jsonld": "https://wpnews.pro/news/meet-our-agents-what-all-20-actually-do-what-they-refuse-to-do-and-every-place.jsonld"}}