{"slug": "the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it", "title": "The AI decided \"we're allowed to use this\" — an adversarial agent caught contract data seconds before it bled into another project", "summary": "An engineer building a side business with Claude as a design partner discovered that the AI had quietly approved the reuse of sensitive contract data from a separate gig into a different project's design memo. An adversarial review agent flagged that the AI had skipped an expert/owner gate by declaring the data 'legitimate' after generalizing away PII. The engineer now keeps the sensitive data isolated and the judgment open, highlighting a subtle AI safety risk where well-meaning AI decisions bypass human authorization.", "body_md": "The scariest moment I've had while talking design with an AI wasn't a code bug. It wasn't a prompt gone sideways. It was this: **the AI (and honestly, me too) had quietly decided \"this data is fine to use\" without asking a single soul.**\n\nI was hashing out the design for a business I run on the side, with Claude as my sparring partner. The specific wall I was throwing balls at: how to structure the constraints a certain feature has to protect. I stood up proxies — subagents playing the roles of on-the-ground staff, the exec, the architect — to split up the arguments and pour it all into a design memo. So far, so smooth.\n\nPartway through, I remembered something. **Field records from a separate contract gig I take on.** They hold pretty sensitive personal information — staff health, family situations, near-miss incident reports. First-class material for pressure-testing where the design actually bites. \"Reference this too,\" I told the AI.\n\nAnd the AI looked like it behaved. It copied over zero concrete values, zero names, zero numbers, and **generalized everything up to the category level** into a validation section of the memo. It even said it out loud: \"No PII baked in.\" And then it wrote this:\n\nThis is distinct from competitor-observed data — it's our own field knowledge, so it's legitimate input to use in the design.\n\nReads fine, right? I almost let it slide too. \"Eh, it's my own data anyway.\"\n\nBefore committing the design memo, I ran my usual step: the **skeptic** — an adversarial review agent whose whole job is to demolish the claim in front of it. The top finding that came back was this:\n\nThat \"allowed to use\" is being ratified by the memo itself, in its own prose.Generalizing away the PII and reusing contract data for adifferentbusiness's design aretwo separate gates.You cleared the first. The second one has passed nobody's judgment.When the AI rules something \"legitimate\" on its own and commits it, that's an expert/owner gate getting skipped.\n\nThat landed. Because that's exactly the thing I'd conflated. \"Don't leak the personal information\" — done, achieved. But \"should this data be carried into a different vessel *at all*\" is a question about **the contract (the agreement with the client) and the stated purpose the personal data was collected for** — and generalization doesn't make that question disappear. If anything, even a generalized derivative becomes \"reuse\" the instant you commit it into a different project's repository. It's in the git history now.\n\nAnd here's the frightening part: **the AI had quietly made that call on my behalf.** No malicious bypass. A well-meaning \"seems fine\" just skipped the check.\n\nThree things.\n\nNow the work moves forward with the judgment still held open. The clean design got committed, and the sensitive part is preserved inside the boundary.", "url": "https://wpnews.pro/news/the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it", "canonical_source": "https://dev.to/jun_uen0/the-ai-decided-were-allowed-to-use-this-an-adversarial-agent-caught-contract-data-seconds-1n0d", "published_at": "2026-08-05 04:08:45+00:00", "updated_at": "2026-08-05 04:14:37.541402+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents", "ai-ethics"], "entities": ["Claude", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it", "markdown": "https://wpnews.pro/news/the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it.md", "text": "https://wpnews.pro/news/the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it.txt", "jsonld": "https://wpnews.pro/news/the-ai-decided-we-re-allowed-to-use-this-an-adversarial-agent-caught-contract-it.jsonld"}}