# After the Sprint: A 72-Hour Build Retrospective - Spoiler: It Wasn't Secure

> Source: <https://dev.to/earlgreyhot1701d/after-the-sprint-a-72-hour-build-retrospective-surprise-it-wasnt-secure-4bd>
> Published: 2026-09-08 01:28:58+00:00

I built [Charitas Clew](https://charitas-clew.web.app/) for a 72-hour DEV Weekend Challenge. It worked. I submitted it. And then I did something slightly inconvenient: I kept poking at it.

Apparently one day is enough historical distance from an event to call

what follows a "retrospective," so here we are. lol 

Charitas Clew takes the sort of bureaucratic notice that can ruin an

otherwise perfectly good afternoon---a benefits letter, utility notice,

court document, hospital bill---and uses Gemini to turn it into plainer

language, dates that may need attention, concrete next steps, and a

speaking script for the phone call someone may need to make next.

The challenge version did those things. It was live. It wasn't

collapsing under its own weight. I had built what I intended to build in

a weekend and submitted it before the deadline.

Then came the uncomfortable question:

**Working according to whom?**

I didn't do the retrospective alone.

My role throughout this project has been the same role I usually occupy

in AI-assisted development: define what the product should do, direct

the agents doing implementation work, interrogate the results, and

decide whether the evidence is good enough to accept.

For the post-submission hardening, I worked with two AI systems in

different roles. **Antigravity using Gemini 3.8 Flash** worked directly against the codebase: implementing changes, running tests, inspecting deployment behavior, and---critically---opening a browser and exercising the live application. **[Dr. Kahlo](https://chatgpt.com/g/g-68af555e39808191a53fcd1ef6451fda-dr-kahlo), my custom ChatGPT code-review and QA assistant**, took the adversarial review side: questioning findings, challenging proposed fixes, catching overclaims, setting the next audit boundary, and repeatedly asking some variation of, "Yes, but what does that actually prove?"

I was the human in the loop deciding what got changed, what didn't, and

when the evidence was sufficient.

That distinction matters because this isn't a story about an AI

magically securing another AI's code.

It's a story about using agents to **challenge the work other agents helped produce**, while keeping a human responsible for the decisions.

The first problem was obvious: the sprint version had no automated test

suite. Before changing security behavior, Antigravity added a minimal

Node/Supertest regression harness so fixes could be checked against

existing behavior. The first 19 tests immediately uncovered an unrelated

malformed-input hang.

By the final audit, there were 143 passing tests.

That number becomes relevant later.

Not because 143 is magical.

Because production still managed to surprise us.

The sprint version had a prompt-injection filter.

It looked for phrases such as "ignore previous instructions," "disregard

prior rules," `system:`, and other strings commonly associated with

attempts to manipulate a model. The filter was there but basic, very basic. 

That sounds responsible until you remember what Charitas reads.

Bureaucracies also use phrases like "disregard all prior notices."

Documents contain instructions. Government paperwork refers to systems.

A legitimate notice can look suspicious to a keyword filter while a

mildly creative attacker can simply phrase an instruction differently.

So during the audit, we stopped asking, "How can we make the blacklist

smarter?" and asked a better question:

**Why is a blacklist responsible for this boundary at all?**

It wasn't.

The hardening removed the keyword gate and strengthened the structural

boundary instead. Application-controlled instructions stayed separate

from uploaded document content. The document was explicitly treated as

untrusted source material. Request fields were validated before entering

the model path.

The other half of that boundary was just as important.

Charitas already asked Gemini for structured JSON, but requesting a

schema from a model is not the same thing as validating what comes back.

Runtime validation was added so types, lengths, required fields,

action-step structure, deadline information, and unexpected properties

had to satisfy the application's rules before the response could reach

the browser.

The lesson wasn't that prompt injection had been "solved." We were

careful not to make that claim.

It was that **security controls should enforce boundaries, not recognize scary vocabulary.**

And model output is still input.

This was the finding that changed how I thought about the whole

exercise.

By this point, the application had a substantial automated suite. Input

validation had been tightened. Model output had a runtime contract.

Unsafe rendering had been removed. Error behavior, retries, timeouts,

deployment headers, and other boundaries had regression coverage.

Then Antigravity opened the deployed application.

This capability turned out to be one of the most valuable parts of the

hardening process because we weren't limited to asking what the source

code *should* do. Antigravity could use the live Firebase-hosted

application: load the page, submit notices, switch languages, test

uploads, generate results, and exercise sharing, printing, and

text-to-speech. We could also inspect the deployed headers and watch

what happened across the real Firebase-to-Cloud-Run request path.

Most of it worked exactly as expected.

The rate limiter did not.

The Express application had been configured with:

```
app.set('trust proxy', 1);
```

That setting depended on an assumption about how many trusted network

hops existed between the user and the application.

Our assumption was wrong.

The production request traveled through Firebase Hosting and Google

infrastructure before reaching Cloud Run. The resulting

forwarded-address chain meant Express could identify a proxy address as

the client instead of the actual originating user.

For an IP-based rate limiter, that's not a cosmetic error. Different

users can collide into a shared quota, while the same user's apparent

identity can vary across proxy paths.

The automated tests hadn't lied. They were correctly testing the network

model we had given them.

**The network model was wrong.**

Antigravity inspected the live behavior, the proxy trust logic was

narrowed around the actual deployment boundary, and the application was

redeployed. Live verification then tested multiple client paths and

spoofed `X-Forwarded-For` values to make sure we hadn't "fixed" one

problem by creating an easier spoofing path.

That produced my favorite lesson from the entire retrospective:

**Tests can prove behavior inside the world you modeled. Production can tell you that you modeled the wrong world.**

The final audit still records a limitation: rate-limit counters are held

in memory per Cloud Run instance rather than globally synchronized. For

the current scale of a small stateless application, we accepted that

instead of adding distributed infrastructure merely because we knew how.

Sometimes knowing what **not** to build is part of the review.

Charitas handles documents people may not want hanging around: court

notices, benefits letters, bills, housing paperwork, and other

potentially sensitive material.

During the privacy pass, the audit found that generated notice

information was being persisted in browser `localStorage`.

There had once been an idea for restoring a previous result. That

feature wasn't meaningfully part of the application anymore.

The storage was.

This created one of those moments where engineering sophistication can

become its own trap. We could have discussed encryption. We could have

created sessions. We could have added a database and retention policies.

We could have transformed a small stateless application into a

significantly larger security problem in the name of solving the smaller

one.

Instead, the review asked:

**Why are we keeping this data at all?**

There wasn't a good answer.

So Antigravity removed the persistence. Sensitive notice content now

remains in application memory for the active session rather than being

restored from persistent browser storage. The application also cleans up

the legacy storage key from earlier versions. Only a non-sensitive

language preference remains persistent.

That is considerably less impressive on an architecture diagram.

It is also the design I trust more.

**Sometimes the strongest data-protection feature is not having the data.**

Those were the three findings worth telling as stories. They weren't the

only things we found.

The complete hardening review covered the application from browser

rendering through model invocation and deployment behavior. By the final

audit, all 17 findings from the original review had an explicit

disposition rather than quietly disappearing from a checklist.

A few of those changes were technically small but important.

Model-controlled values stopped reaching unsafe HTML rendering paths.

Uploads gained stricter MIME, base64, size, and file-signature

validation. Transient frontend failures stopped masquerading as

permanent shutdowns. Security headers were aligned across Firebase

Hosting and the Express backend. Secret handling was verified against

the deployed environment rather than inferred from the repository.

One change wasn't strictly a security fix at all.

Charitas extracts dates from notices, but a date printed on a document

is not automatically *the legal deadline*. Deadlines can depend on

service dates, receipt dates, procedural rules, statutes, or facts that

aren't present in one uploaded page.

The hardening therefore changed the product language too. Extracted

dates are treated as evidence from the document, while users are

reminded to confirm actual deadlines with authoritative sources.

That was an important reminder that hardening an AI product isn't only

about preventing malicious behavior.

Sometimes you have to harden **what the product is allowed to claim**.

I'm going to resist answering that with "yes."

The final Antigravity audit gave Charitas Clew a 9.0/10 production-readiness score and a verdict of:

**APPROVED FOR CURRENT SCOPE.**

I care more about those last three words than I do about the number.

The application still has accepted limitations. Rate limiting is per

Cloud Run instance. The proxy-trust configuration carries maintenance

debt because the network ranges it relies on can change. Documents

necessarily cross an external model-inference boundary. There are

security improvements that could still be made.

Those aren't forgotten fixes. They're documented trade-offs.

Security work has no natural finish line. There is always another

scanner, dependency, service, abstraction, test, policy, or hypothetical

scale problem available to consume an afternoon. Eventually, "hardening"

a 72-hour project stops being responsible maintenance and starts

becoming an elaborate way to build a different application.

So there is no Phase 7.

I cannot believe I just wrote that sentence about something I started

building on Friday.

The sprint proved that Charitas could work.

The retrospective asked whether the boundaries around that working

product could survive harder questions. Some couldn't. We changed them.

Then we tested again---including against the deployed system rather than

only the system we imagined we had built.

That's a different standard from "it runs."

For an application people may trust with documents that matter, I think

it should be.

*Charitas Clew was built and hardened with substantial AI assistance. I directed the product, review criteria, trade-offs, and acceptance
decisions; Antigravity worked directly with the codebase and live
deployment; Dr. Kahlo, my custom ChatGPT QA and code-review assistant,
helped structure and challenge the post-submission audit. The hardening
described here occurred after the DEV Weekend Challenge submission and
is not represented as challenge-period work.*

**AI Assisted. Human Approved. Powered by NLP.**
