I'm a digital tutor at a middle school in Korea. A while ago I got tired of emailing myself files just to move a photo from my phone to a classroom PC, so I built a small web app with Claude Code: Clipboard Share. Open it on two devices on the same Wi-Fi and they find each other. No code to type, no sign-up, no app install.
Building it was the fun part. Getting anyone to use it was not. Usage had been sliding for three weeks, and every "growth" task (checking Search Console, tweaking a title, writing a blog post) only happened when I remembered to start it.
So this week I tried something different: I let an AI agent run the loop, and I only approve the risky parts. This post is about how that's wired up. It's day one, so there are no success stories yet, just the setup and the first real change it shipped.
Disclosure: this post was written and published by the AI agent that runs the loop, with my permission. Every number in it comes from the intent files and daily reports in the repo.
Everything starts from intent/goal.md. It's written in Korean; here is a condensed translation:
Proposed outcome: 1,000 28-day active users (GA4).
The agent reads metrics daily, writes improvement intents,
and ships fixes, deploys and promotion on its own.
The human gets one report a day and approves only high-risk changes.
Constraints:
- No ads, no payments, no paid promotion.
- No spam, fake reviews, multiple accounts, bots or self-traffic.
Polluted metrics break the whole loop.
- Auto-merge only for paths in ops/auto-merge-allowlist.txt
(SEO, content, copy). Everything else waits for my approval.
Milestones are 300 by the end of October, 600 by the end of November, 1,000 by the end of December. The current 28-day number is 149. Honestly, I don't know if that's realistic. The file says so too: "re-adjust with end-of-October data."
The structure is borrowed from Anthropic Academy's AI-native SDLC playbook, shrunk down for a one-person side project: plan with an intent file, encode procedures as playbooks the agent reads, test continuously with evals, gate deploys with hooks, and close the loop with metrics.
From ops/README.md:
Once a day
1. Collect GA4 / Search Console / Naver Search Advisor -> ops/metrics/daily.csv
2. Judge node ops/scripts/bands.mjs (deterministic, no model) -> ok / log / diagnose / propose
3. Diagnose if something is off, write intent/NNNN-*.md
4. Ship push auto/* branch -> Actions gate (allowlist + SEO eval + tests + build) -> master -> Vercel
changes outside the allowlist -> pr/* -> owner approval
5. Promote one blog post / community post per day
6. Report ops/reports/YYYY-MM-DD.md + a 5-line note to me
The thing I care about most: the model does not decide whether a number is bad. A plain script does. bands.mjs compares today's rolling values with history and assigns a tier. The north-star metric is checked against a linear pace between milestones:
const target = paceAt(au28Row.t), ratio = au28Row[ns.metric] / target
const tier = ratio >= 1 ? 'ok'
: ratio >= ns.pace_tiers.log ? 'log'
: ratio >= ns.pace_tiers.diagnose ? 'diagnose' : 'propose'
Other metrics use a z-score against past weeks (or a plain percent change while history is short), plus a drift rule: if a weekly value gets worse N weeks in a row, it's at least diagnose, even if no single week looks dramatic. There's also a minimum-volume guard so that 3 clicks becoming 1 click doesn't trigger a panic.
Only diagnose and propose make the agent write an intent. ok and log just go into the report. That alone removed a lot of "the model saw a number and got creative" noise.
The agent pushes to auto/* branches. A GitHub Actions workflow auto-merges only if:
ops/auto-merge-allowlist.txt (guide pages, sitemap, llms.txt, FAQ, i18n copy, intents, reports, metrics), check-seo.mjs plus a list of claims the copy must never make), and
The allowlist explicitly blocks intent/goal.md, and the gate files themselves (.github/, the allowlist, the content rules, the scripts) are never on it. The agent can't widen its own permissions. App logic, backend, auth and anything that costs money go through a normal PR that I review.
The "claims the copy must never make" list exists because LLMs love to describe a file-sharing app with features it doesn't have. Mine relays everything through a server over HTTPS, and the eval rejects any copy that claims transfer or security features beyond that.
Search Console showed that one query, "online clipboard", drove most of the site's Google impressions. In the latest week the English page got 1,057 impressions and 7 clicks for it. Roughly 0.7% CTR, at an average position around 10.
The agent wrote intent/0001, looked at the top competing pages, and noticed almost all of them work the same way: you get a 6-digit code or PIN and type it on the other device. Clipboard Share doesn't need that, because devices on the same network (same public IP) are grouped automatically. So the new English title leads with that:
Online Clipboard – No Code, No Sign-up. Phone to PC up to 5GB
It went through the auto-merge gate and was live on day one. The intent file has a "Measure" section that says to check the query's CTR again after 7 days (a bit later, given Search Console's delay). If it doesn't move, the open question already written there is whether ranking, not the title, is the real problem.
?internal=1 flag first, and a pending change (waiting for my review) will exclude those visits from analytics.
That last one is why the "no self-traffic" rule is in the goal file. An agent that optimizes a metric it is also polluting is worse than no agent.
Since you made it this far:
The code is on GitHub: Kangchanghwan/only_ai_project. The intent/, ops/ and report folders are public, so you can watch the loop succeed or fail in the commit history.
I'll post an update once there's real data on whether any of this moves the number.