cd /news/ai-agents/how-we-scale-our-codebase · home › topics › ai-agents › article
[ARTICLE · art-144197] src=sageox.ai ↗ pub= topic=ai-agents verified=true sentiment=· neutral

How we scale our codebase

SageOx, founded in January by @tensorport443, @rsnodgrass and @milkanabrace, runs a set of cloud agents it calls its "beehive" to maintain a codebase spanning a web application, hardware, a CLI, a desktop app and an upcoming mobile app. The agents include Bugsy Loggins, which files bugs from production logs and ignored 31 logs that failed its bar; Verity TestAuditor, which finds testing gaps; Paul Bunyan, which implements issues and addresses review comments from @greptile and @coderabbitai; Whittle Lessmore, which reduces code duplication; RIP, which tags agent-created pull requests for a merge queue via @trunkio or flags them needs-human; and Beekeeper, which checks daily that the agents are running. SageOx says the agents are seeded with team context from SageOx, including meeting decisions and designs, making them more capable than general cloud agents.

by read4 min views5 publishedSep 22, 2026
How we scale our codebase
Image: Sageox (auto-discovered)

SageOx started in January by @tensorport443, @rsnodgrass and @milkanabrace with a vision to enable a small and high performance team with shared context knowledge across all the surfaces that makes the team move faster.

We quickly realized we need to cover a lot of surface and stitch them together to build a cohesive experience for teams to deliver impact even faster. We fully embrace AI and we have come to realize that moving faster is important but it also comes with code quality issues.

What are the general issues we see?

  • LLMs creates a lots of bugs (humans do too) which eventually gets surfaced at unexpected times
  • Code duplication: I can't even tell you how bad the problem of code duplication is. You can see it in my previous post https://x.com/shrimalmadhur/status/2096391539195048442?s=20
  • Useful test coverage: LLMs can create test coverage but they create so much sometimes that they also skip useful ones.
  • Stale code comments and docs: LLMs are bad at update their own comments, docs and then docs drift from code a lot.

Ok now multiply this across so many surfaces we have - a fully functioning web application, a dot hardware, a CLI, a desktop app, a mobile app (coming really soon - like whenever Apple approves it). This is so much code 9 humans can't possibly clean up!!

Enter our cloud agents aka our beehive (Well it's named after Buzz by Block where we observer our agents)

We have started running these agents who act as our coworkers. Before we see what these agents do, one important distinction I want to make it these agents are way more powerful than your general cloud agents because whenever they spin up they are seeded with all the team context from SageOx so they already know what the team is doing and have access all the meeting decisions, designs etc in near realtime.

Bugsy Loggins #

Bugsy looks at our production logs and files bugs based on if it's a real issue or flue. You can see that it is ignoring 31 logs because it did not pass its bar.

Verity TestAuditor #

Verity looks at the code and figures out what are the gaps in our testing. Have we covered all our important use cases, edge cases etc and files an issue.

Paul Bunyan #

Bunyan takes issues from bugsy and verity and implements them. It also looks for Github tags with "bunyan" and takes them up to. This agent has its own lifecycle. It wakes up at a defined cadence and then takes up any new issues or looks at previous PR it created and addresses any code review comments and pushes the changes. Currently we use @greptile and @coderabbitai for reviews. So Bunyan replies to them to keep them updated so comments can be resolved.

Whittle Lessmore #

Whittle takes care of code duplication. It goes through a certain folder or sometimes across packages to see where we can reduce our code duplication. If it finds a valid code duplication then it will create a fix and push a PR out for review. Same as Bunyan, it will address review comments and will make sure it's in a mergeable state.

RIP #

RIP is the final stage where it looks at all the PR created by agents and classifies them. If RIP thinks that this PR is good to merge and will be okay going to production then it will tag it with a merge queue tag (We are experimenting with @trunkio right now) so that a merge queue takes them, runs the CI with the latest code and merges this. If RIP thinks that this PR requires human attention, it will put a needs-human tag so we can take a look and make any changes or clear it to be merged by RIP.

Beekeeper #

Beekeeper is the leader of our hive. It wakes up everyday and checks in with everyone to make sure they are up and running.

I hope you are still with me. If you are, here are some cool stats from our agents.

Now this is the just the tip of the iceberg, We are brewing more cloud agents to do a lot of stuff like:

  • Release management
  • Agent that can constantly test our hardware product
  • Docs update
  • Agents for suggesting architectural improvements to our systems
  • Onboarding agent to onboard a new joiner
  • and many more....

So how can you get these capabilities? We have open sourced our Agent Toolkit so you can deploy agents in your infrastructure. In addition, you can sign up at https://sageox.ai if you want to supercharge your agents with all the context from your team and other agents.

If you are interested in talking to us, please reach out to us here!

── more in #ai-agents 4 stories · sorted by recency
── more on @sageox 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-we-scale-our-cod…] indexed:0 read:4min 2026-09-22 · —