{"slug": "i-let-ai-agents-write-most-of-a-live-trading-codebase", "title": "I let AI agents write most of a live-trading codebase", "summary": "A solo engineer running the live-trading system TopSet used AI coding agents including Claude Code, Cursor and Copilot to produce roughly 1,600 pull requests over 21 months, gating every merge on a green build of about 7,500 tests across 331 test modules plus lint, type checking, migration rollback and terraform validation. The engineer reads model-training and execution-path changes directly and treats any snapshot or end-to-end test diff as a stop, because agents will regenerate snapshots to force a build green. For unverified calculations, 161 report scripts dump the pipeline to CSV for hand-checking of return per trade, slippage and annualized return.", "body_md": "# I let AI agents write most of a live-trading codebase. Here's the gate that made it safe.\n\nFor the last 21 months I have been the only engineer on TopSet, a system that trades real money. It picks a portfolio with machine-learning models, rebalances on a schedule, and executes through live broker APIs on infrastructure I own end to end. It is my own capital, not anyone else’s.\n\nAgents wrote most of the code. Claude Code, Cursor, Copilot. About 1,600 pull requests in those 21 months, which is more than I could have typed in three times the span.\n\nThat leaves you with a choice everyone adopting these tools runs into. You can read every line the agent produces, in which case you have bought yourself a slower version of writing it yourself. Or you can skim, merge, and move fast, in which case you accumulate code that looks right and isn’t. On a system that places orders with real money, the second option ends with trades you didn’t intend.\n\nI took a third option: make something other than my attention the thing that decides what ships. Tests became the gate. Nothing merges or deploys without a green build, and my own reading moved up a level, to design, to test coverage, and to the handful of results I refuse to let change quietly.\n\nThat worked. It also failed in ways I did not expect, and the failures are the interesting part, so I’ll spend most of this post there.\n\n## What I actually read\n\n“Tests are the gate” is not the same as “I stopped looking.” What changed is where I spend attention.\n\nMost pull requests get a light review. Research code I skim. Changes to the parts that decide what to buy or how to trade it, the model training and the execution paths, I read properly, looking for whether the approach is reasonable rather than whether every line is correct. The tests cover the lines. I cover the idea.\n\nThen there is one rule I do not bend. If a change moves a snapshot, or shifts the result of an end-to-end test, that is a stop. Snapshots and integration results are supposed to be stable, so a diff in one of them means either the change did something I didn’t intend, or it did something I did intend and I now owe an explanation. An agent will happily regenerate a snapshot to make a build go green. That is the single most dangerous thing these tools do, and it is why the snapshot files live in the repo where a diff is visible in review.\n\nFor anything involving a calculation I have not verified before, the tests are not enough either. I have the agent write a report script, and there are now 161 of them, that dumps the pipeline to CSV. I open the CSV in a spreadsheet and check the arithmetic by hand: return per trade, slippage, annualized return. Then I spot-check inputs, picking a few prices and going back to the raw data directory to confirm those were the prices the system actually used.\n\nIt is slow and it is not clever, and it has caught things nothing else would. A test asserts that a function returns what the author believed it should return. A spreadsheet asks whether the number is true.\n\n## What the gate is\n\nEvery pull request runs:\n\n- lint (ruff) and type checking (mypy)\n- about 7,500 tests across 331 test modules\n- an assertion that the test database is empty afterward\n- a migration rollback-and-reapply, all the way to base and back\n- `terraform validate` against both AWS accounts\n\nIf any of that fails, nothing merges. The build is the merge gate, and it is not advisory.\n\nOnce a week, on a schedule, a heavier set runs:\n\n- two end-to-end rebalancing runs against a mock broker\n- the same integration test 100 times in a row, looking for flakes\n- nine deterministic model-training snapshot runs, pinned inside Docker\n- two leakage checks on the training pipeline\n\nRunning the same integration test 100 times only makes sense because the test is not the same each time. The mock broker fills orders probabilistically: a 20% chance per 10-second cycle, 70% of those complete and the rest partial at a random size, at random prices. Every iteration produces a different interleaving of fills, partial fills and timing. A hundred runs is a search for race conditions, not a hundred repeats of one path.\n\nWhat it finds are real concurrency bugs wearing flaky-test costumes. Worker threads leaking between tests. A queued task waking up after teardown had already deleted the row it wanted. A data loader that wasn’t thread-safe and corrupted a shared index under load. Each one showed up first as “that test fails sometimes.”\n\nThe weekly set exists because a flaky gate is worse than no gate. If a test fails once every thirty runs, and twenty agent PRs land in a week, that one test goes red most weeks. Have a few like it and something red is always on screen, and you start ignoring red. Hunting flakes is not hygiene here. It is what keeps the gate meaningful.\n\n## The tests that matter are about money, not functions\n\nUnit tests are cheap and agents write them happily. They are not what I trust.\n\nWhat I trust are end-to-end tests that drive a full rebalancing against a mock broker, with the awkward parts turned on. They’re named after the ways money goes wrong:\n\n- orders that have to be resubmitted\n- a rebalancing that resumes from the middle, with partial fills already on the books\n- a resume that has both buys and sells outstanding\n- smart orders cancelled mid-flight, on the sell side and then the buy side\n- deposits and repeating withdrawals arriving while a rebalancing is in progress\n\nThese are exactly the control flows a language model gets plausibly wrong. Interrupted work, partial state, money arriving in the middle of an operation: the code for each of these reads fine in isolation. It’s the interaction that bites.\n\nHere is one that made it past review and got caught by the constraint at the bottom\nof the stack. When the last pending buy of a rebalancing had its amount adjusted for\npending withdrawals, the adjusted value could land on exactly zero. The existing\nguard raised an error on a negative amount, but zero isn’t negative. The\naffordability check, `available_cash >= amount`, was `0 >= 0`, so it passed. The\norder was promoted, and the database rejected it: `amount > 0` is a check constraint\non that table.\n\nThe interesting part is what that did to the retry. The integrity error fired inside the same transaction that carried the completing order’s status update, so the rollback discarded that too. The scheduler came back around, found identical inputs, and did it again. Not a self-healing retry: a stuck rebalancing.\n\nThe fix took three attempts, and the second one is the useful story. Gating the\npromotion on `amount > 0` looks correct and does nothing, because the code had\nalready assigned the adjusted value to a live ORM object. That marked it dirty, and\nthe next autoflush in the same transaction persisted the zero regardless of whether\nthe order was ever promoted. The condition guarded the decision, not the mutation.\n\nI only know that because the failure was reproduced live against `main` before any\nfix was written, and the first fix was re-run against that same repro rather than\nchecked by reading the code. When an agent is writing the patch, “I read it and it\nlooks right” is the single least reliable signal available to you.\n\n## Three ways a green build lied to me\n\nThis is the part I’d want to read, so here are the three worst cases, all real, all from this codebase.\n\n### 1. The test asserted a convention that does not exist\n\nCorporate actions have to be applied to historical prices, and stock dividends are\nthe fiddly ones. My broker’s API returns a `rate` for each. The code treated that\nrate as a bare fraction: a 5% stock dividend arrives as `0.05`.\n\nIt isn’t. It’s the ratio of new shares to old. Every one of the 101 stock-dividend records across 43 symbols in my data has a rate of at least 1.0016. Not one is a bare fraction. The convention the code assumed does not occur in the data at all.\n\nThe unit test passed, because the test asserted the same wrong convention with a\nfixture value of `rate: 0.05`, a number that cannot appear in production. The test\nand the code were wrong in exactly the same way, which is what you get when both are\nwritten in the same sitting from the same misunderstanding.\n\nThe effect was that each stock dividend roughly halved the adjusted price, and it compounded across events. One symbol with nine of them came out deflated by about 490×, turning a real price history into a fake ramp of more than a thousand-fold over the period. And because every price stayed between a few cents and a few hundred dollars, a check for implausible price levels would not have flagged it either.\n\nThe fix was a corrected fixture, plus a regression test that runs a nine-event sequence and asserts the cumulative adjustment factor stays sane. The lesson I took: a fixture is a claim about the world, and nobody reviews fixtures.\n\n### 2. The test passed on NaN\n\nIn a research tool, an agent added a second way to select a model configuration and reported the result in the PR: 100% agreement with the existing selector, at every sample size tried.\n\nThat number was not real. The tool built its pool from a source that was missing the columns the default scoring version needs, so every score came out NaN. With all scores NaN, the selector fell through to a fixed tie-break order, which of course agrees with itself 100% of the time at every sample size.\n\nThe test that was supposed to cover this had been passing on those same NaN scores for as long as it had existed. It never checked anything. Green, meaningless, and quoted in a PR description as a finding.\n\nThe follow-up PR made the condition a hard error rather than a silent NaN, pinned the test to a scoring version where the columns exist, and added guards. It also flagged that an older result produced the same way might carry the same artifact. That is the right instinct: when you find a lie, go and check what else it told you.\n\n### 3. The change did nothing at all\n\nAn agent wired a new training target through what looked like the right code path. The first screen ran clean and produced numbers bit-for-bit identical to the baseline. Every metric, to four decimal places.\n\nIdentical results are not a pass. They’re a smell. Two different configurations producing byte-equal output usually means one of them didn’t happen.\n\nIt hadn’t. The runner always sets a parameter that routes execution to a completely separate function, which builds its own training target and never looked at the new field. Unpickling both runs and diffing the actual selections showed 238 of 238 identical picks for the test year. After the real fix, 0 of 238 were identical, and a regression test now asserts the two modes produce different output.\n\nNothing failed. The compute was spent. The result would have gone into the research log as “no effect,” which is a conclusion, and it would have been wrong.\n\n### The shape of all three\n\nIn every case the build was green, and the number was wrong. Tests prove that the code does what the tests say. They say nothing about whether what the tests say is true, whether the code under test ran at all, or whether the data going in is what you think.\n\nThat is the agent-era failure mode, and it is worse with agents than without them, for a simple reason: an agent writes the code and the test together, from one reading of the problem. If the reading is wrong, both artifacts are wrong, and they agree. A human pair would at least have had two readings.\n\nWhat I do about it now, concretely:\n\n- Distrust results that are identical, unchanged, or perfect. Check the output, not the exit code.\n- Treat fixtures as claims. If a fixture value can’t occur in production, that’s a bug in the test.\n- Make impossible states loud. NaN, missing data, and empty inputs should raise, not flow through and produce a number.\n- When something is found to be wrong, go back and check what else that thing told you.\n\n## The reviewer that isn’t a test\n\nSome mistakes are not code-shaped and no test will ever catch them.\n\nSo I added a second kind of review: a deliberately skeptical reviewer persona, run against a finished result rather than a diff, with instructions to argue that the result is not real. I iterated on the persona with Claude until it stopped being agreeable.\n\nIt found things I am glad it found. That a configuration had been tuned by repeatedly checking results on data I had promised myself I would only look at once, so my holdout wasn’t a holdout any more. That a cost model was running with zero financing cost while the strategy it priced used nearly four times leverage.\n\nNeither is a bug. Both are the kind of error that invalidates a conclusion, and no test suite in the world was going to tell me.\n\n## The gate after the gate\n\nCI stops caring the moment something merges. After that the system runs unattended, on a schedule, against a live broker, and the question changes from “is this code correct” to “is this thing behaving right now.”\n\nEvery ERROR the system logs in production becomes a CloudWatch alarm that posts to a Discord channel on my phone. That is the whole alerting stack, and for a one-person operation it is enough: I am not running a rotation, I just need to know.\n\nThe loop from there is the same one I use for everything else. Point an agent at the production logs, hand it a local copy of the database from around the failure, ask for a diagnosis and a fix with a test. The logs and the data are usually enough for it to find the problem, and the test it writes is the thing that stops the problem coming back. Several of the bugs in this post arrived that way rather than through CI.\n\nWorth saying plainly: the tests did not catch those. The tests caught the next occurrence.\n\n## What stays human\n\nArchitecture. Test design. What “done” means. Merge judgment when the tests pass and something still feels off.\n\nAnd one more thing that turned out to matter: the repo’s own instruction file. Every time an agent reached a wrong conclusion in a way another agent would repeat, I wrote the correction down there. One rule says that any claim of data leakage must name the specific date boundary it violates, because two different agents had asserted leakage from reasoning alone, and both were wrong. That file is now a list of the mistakes this codebase invites. It’s the closest thing I have to institutional memory for a team of one.\n\n## What it cost\n\nThe per-PR suite takes about ten minutes. That is the tax on every change, and it is the number I would defend hardest: ten minutes is short enough that I never route around it, and long enough to run tests that touch a database and a fake broker.\n\nThe machine bill is every one of the 3,000 GitHub Actions minutes my plan includes, plus roughly $20 a month for extra minutes on top. The agents themselves are Claude Pro and Cursor Pro, $20 each. I also paid $40 a month for Copilot until GitHub moved it to per-token billing, which would have raised my bill more than tenfold, so that work moved to Claude Code and Cursor. So the tools that wrote most of this codebase cost tens of dollars a month, like the compute spent checking them, and neither is what the project actually cost.\n\nWhat it actually cost is fixtures and flakes. Writing a mock broker that fails in realistic ways, maintaining a self-contained dataset so model training can be tested at all, and then the long tail of chasing tests that fail once in thirty runs. None of that is glamorous and all of it is load-bearing. Every hour I skipped there came back later as an hour of not being able to trust a result.\n\nAnd sometimes the gate itself is the thing that’s broken. The 100-run soak lives in a\nshell script rather than inline in CI, because the inline version ran under a POSIX\nshell where `{1..100}` isn’t brace expansion. It was a literal string. The loop ran\nexactly once, and reported success, for longer than I’d like to admit.\n\nThe tests are not where I over-built. Every gate in this post exists because something broke, or nearly did, and the existing checks let it through. I would add all of them again.\n\nWhere I over-built is everything the tests had to cover.\n\nI started TopSet expecting to turn it into a product, so I built it for customers who never existed. A bot lifecycle with eleven kinds of change request, each one a queued, locked, asynchronous workflow object, so that deposits and withdrawals and stops could be requested safely by someone who wasn’t me. Two broker integrations instead of one. A staging environment that mirrors production, behind a VPN. Cross-account backups with compliance-grade retention locks on a database that holds my own trades.\n\nFor one person running their own accounts, several of those could have been a script I ran by hand while watching it.\n\nThat is the part that compounds, and it is why I am putting it in a post about testing. Every one of those surfaces is something the gate has to cover. The change requests need end-to-end tests for money arriving mid-rebalance. Two brokers means two of everything, including the mocks. The staging environment is a second place for migrations to fail. I did not pay for the product shape once, in the building. I have paid for it every week since, in the size of the harness that has to keep it honest.\n\nIf I were starting again for a book of one, I would build the trading and the research properly and leave the product machinery until somebody asked for it.\n\n## If you’re rolling this out to a team\n\nI’ve done this alone, with tests I wrote and one person’s judgment. That is not the same problem as a team of fifteen, and I’d be careful about anyone who tells you otherwise. But some of it transfers, and the ordering I’d insist on comes from a team I did run: at a healthcare company, requiring automated tests on every PR is what took active medium and high bugs down 72% and weekly hotfixes from 7 to 1.5. That was before agents. It’s the same first move.\n\n1. **Fix the gate before you raise throughput.** Coverage on the paths that lose\nmoney or data, and no tolerated flakes. Agents multiply whatever your CI already\nis.\n2. **The author owns the change, whoever typed it.** “The agent wrote it” is not a\nreview comment, and it is not a defense in a postmortem.\n3. **Review the result, not the diff.** For anything that produces a number, ask\nwhat would look different if the change had done nothing.\n4. **Measure escaped defects and hotfix rate.** Not lines, not PR counts. Agents\nmake those metrics meaningless overnight.\n\nThe part I would not change: the gate is what let one person move at this speed on a system where mistakes cost real money. The part I underestimated: a gate is a claim about correctness, and claims need checking too.\n\n*This describes software I run on my own accounts. It is not investment advice and I\ndo not offer investment services.*", "url": "https://wpnews.pro/news/i-let-ai-agents-write-most-of-a-live-trading-codebase", "canonical_source": "https://redgeoff.com/posts/ai-agents-test-gate/", "published_at": "2026-09-29 14:34:48+00:00", "updated_at": "2026-09-29 14:47:59.123320+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["TopSet", "Claude Code", "Cursor", "Copilot", "AWS", "ruff", "mypy", "Terraform"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-let-ai-agents-write-most-of-a-live-trading-codebase", "markdown": "https://wpnews.pro/news/i-let-ai-agents-write-most-of-a-live-trading-codebase.md", "text": "https://wpnews.pro/news/i-let-ai-agents-write-most-of-a-live-trading-codebase.txt", "jsonld": "https://wpnews.pro/news/i-let-ai-agents-write-most-of-a-live-trading-codebase.jsonld"}}