# Your AI Agent's Rules File Is a "Gentleman's Agreement". Here's what happens when you make misbehavior structurally unprofitable instead

> Source: <https://dev.to/nseney1/your-ai-agents-rules-file-is-a-gentlemans-agreement-heres-what-happens-when-you-make-3df5>
> Published: 2026-09-30 20:35:48+00:00

If you use AI coding agents (Cursor, Copilot, Claude, Gemini), you've probably written rules. A `.cursorrules` file. A `CLAUDE.md`. Something like:

```
- Always read a file before editing it
- Run tests after making changes  
- Don't introduce new dependencies without asking
```

This is a gentleman's agreement. You're asking the agent to follow rules that it can silently ignore with zero consequences. There's no mechanism that detects when "always read before editing" is violated. There's no feedback loop that retires rules nobody follows. The rules live in the context window and the agent can simply... not.

The research backs this up. Studies show ungoverned agents waste significant portions of their token budgets on circular rework, hallucinated APIs, and broken assumptions ([arXiv:2602.11988](https://arxiv.org/abs/2602.11988), [arXiv:2607.27250](https://arxiv.org/abs/2607.27250)). Adding static rules helps — but the rules themselves never improve, never expire, and never prove they're working.

I wanted something better. So I built [Soma](https://github.com/nseney1/Soma-Governance), an open-source governance framework that treats the problem differently.

The core idea is borrowed from mechanism design in economics: **don't rely on the agent's willingness to comply — make non-compliance structurally unprofitable.**

In practice, this means three things:

Every rule in Soma has an expiry. Time-based (90 days) or session-based (30 sessions). If the evidence pipeline hasn't recorded a single trigger event for a rule — meaning the rule never matched any files the agent touched — it expires and gets pruned from the context window.

This sounds aggressive, but think about what it prevents: rule files that grow monotonically until they eat 20%+ of your context window. Rules added for a one-time incident that never recurs. Rules that were relevant for a codebase you refactored six months ago.

When an agent (or a subagent) reports "I fixed the bug, all tests pass" — that's cheap talk. Soma's verification layer checks the actual evidence: did the transcript show a test run? Did it pass? How many fix-and-retry cycles happened?

This sounds paranoid until you've been burned by a subagent that reports success, you merge the PR, and discover the "fix" introduced three new bugs because the agent never ran the test suite.

Every session generates evidence. Which rules triggered? Which files were touched? Were there rework loops? This data flows into a fitness ledger (a simple JSONL append log — nothing fancy). Rules that trigger frequently and correlate with clean sessions accumulate fitness. Rules that trigger but correlate with rework accumulate noise signals.

Over time, the system promotes rules that work and demotes rules that don't. No human has to manually curate the rule file.

Here's something that actually happened today while working on Soma itself. It's a perfect example of the failure mode this framework is designed to catch.

My AI agent had a test failing in the local environment. `test_make_validate_fails_on_broken_shell_script`. It failed every single run for the entire session. The agent dismissed it as a "pre-existing sandbox issue" — the test was writing to a read-only filesystem.

The agent fixed that test. Local suite: 516 passed, 0 failed. It committed. It declared victory.

I asked: *"It's still failing in the build. Is there another lesson?"*

The agent had never once checked the CI build. When it finally looked at the actual CI log, it found a completely different bug — a production file (`immune_trends.py`) had an IndentationError. The local test failure and the CI failure were **different bugs with the same symptom** ("build is red").

The agent fixed that. CI ran again. Failed again. A *third* bug: `pyyaml` wasn't in the CI dependencies. Six test files crashed on import.

Fixed that. CI ran again. Failed again. A *fourth* bug: a test was `source`-ing an entire shell script that runs `git push`, causing a 60-second timeout in CI.

Four distinct bugs, stacked, each masked by the one before. The agent spent an entire session dismissing "build is red" without reading the log. The governance framework it was building to prevent exactly this kind of failure... failed to prevent it, because it hadn't been applied yet.

That failure is now a governance cell in the system: `trap-local-green-ci-red`. It fires when changes touch CI-related files and reminds the agent: *local green ≠ CI green. Verify the actual build.*

I want to be honest about the limitations, because the AI tooling space is full of overclaimed metrics.

**I don't know if this generalizes.** Soma has been tested primarily on my own projects. The failure modes it catches are real, but they might be idiosyncratic to how I use agents. The governance cells encode *my* scar tissue. Whether they transfer to other developers' workflows is an open question.

**The metrics are qualified.** The system adds ~3,800 tokens of idle context overhead (measured, down 8.6% from earlier versions). Waste rate is under 1.0% in governed sessions. But "governed session" is doing a lot of work in that sentence — it means sessions where the full framework is loaded and the agent is following the rules. Measuring counterfactual waste (what would have happened without governance) is hard.

**Rules still live in the context window.** This is the fundamental limitation. Soma makes rules smarter and self-pruning, but they still compete for context space with the actual task. A platform-level solution — governance enforced in the model layer, not the prompt — would be strictly better. I haven't built that because I don't control the model layer.

**Expiry is blunt.** A rule that hasn't triggered in 30 sessions might be an important safety net that hasn't been needed yet, not a useless rule. The current system can't distinguish "dormant but valuable" from "dead weight." I'm thinking about approaches here but don't have a good answer.

Soma is open source at [github.com/nseney1/Soma-Governance](https://github.com/nseney1/Soma-Governance). As of v0.50:

The architecture is platform-agnostic (MCP-based), so it works with Gemini, Claude, or any agent that speaks MCP.

I'm not claiming Soma solves AI governance. I'm claiming that the *approach* — making rules compete for survival based on evidence — is more robust than the alternative, which is a static file that grows forever and relies on the agent's good faith.

Whether that thesis holds up at scale, across teams, with different agents and codebases — I genuinely don't know. If you try it and find out, I'd like to hear about it.

[Soma](https://github.com/nseney1/Soma-Governance) is open source under MIT. The governance cells — the encoded failure modes — are the most interesting part. Contributions welcome, especially cells from your own agent failures.
