cd /news/artificial-intelligence/ai-in-security-volume-from-the-model… · home › topics › artificial-intelligence › article
[ARTICLE · art-145651] src=collinsthoughts.substack.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI in Security: Volume From the Model, Trust From the Structure

A security engineering team built a four-stage pipeline that uses an LLM to triage scanner findings, a deep pass to prove code reachability, and a separate refuter model to challenge confirmed results, with human review for anything rated High. Across six batches the first pass ruled out 47% to 92% of scanner noise, but the deep pass overturned 10% to 28% of findings the first pass had called real, mostly because the model assumed code was deployed, recalled framework behavior instead of verifying it, inflated insider-only bugs, and trusted its own write-ups. The team's conclusion is that models supply volume while trust must come from structure: required reachability evidence, version-pinned framework checks, an internal risk matrix, and a second model tasked with disagreeing.

by read10 min views9 publishedOct 5, 2026
AI in Security: Volume From the Model, Trust From the Structure
Image: Collinsthoughts (auto-discovered)

AI is very good at the part of security that looks like reading. Hand a model a flagged line of code and it will trace the data flow, explain the bug and rate its severity, faster and more consistently than most people can. That part is real, and it is why every security team is wiring models into its pipeline.

It is also where the trouble starts. One of the most convincing findings a model ever gave us was a textbook missing authorization check in a GraphQL resolver, traced end to end and rated High. Every line of the reasoning held up. The resolver had never been registered with the framework, so it had never served a single request.

The model was right about the code and wrong about the system around it: what actually runs, who can reach it, and what the model itself should be allowed to touch. That gap is the whole story of using AI in security. The model gives you volume. Trust comes from the structure you build around it.

We hit that lesson three times, from three directions. AI reads our scanner output and decides what reaches engineers. Agents do work that used to take the team weeks. And our developers run their own agents, which hold tokens and reach production. Each direction paid off, and each one failed the same way until we built for it.

The rule that came out of all three is short:

  • Let the model carry the volume: the reading, the fan-out, the first draft.
  • Build structure that carries the trust: evidence it has to show, a second model whose job is to disagree, scripts that do the writing, and a wherever the risk is real.
  • Spend people on what the model can’t settle alone: does this run, who can reach it, and should this happen at all.

🔎 Direction 1: AI that reads the findings #

Our scanner flags broadly on purpose. Missing a real bug costs more than over-flagging, as long as the noise never reaches an engineer. A model sits between the scanner and the teams, and that’s where it earns its keep.

Each finding passes through four stages. A first pass reads it with the surrounding code and calls it real, noise or uncertain. A deep pass takes everything called real, clones the repository and tries to prove the code is reachable. A separate refuter model then tries to kill each confirmed finding. Anything rated High gets a person before it reaches a team.

The first pass delivers what everyone promises. Across six batches, it ruled out between 47% and 92% of what the scanner raised, and engineers stopped seeing noise. The spread is the interesting part: it depends on which rules fired. A batch dominated by rules that trip on scripts, tests or globally guarded routes is mostly noise.

The deep pass is where it got interesting. In most batches it overturned between 10% and 28% of what the first pass had called real, and almost every overturn came from one of four habits:

  • It assumes the code runs. It confirmed a textbook missing-authorization bug in a resolver that was never registered with the framework, a route behind a flag that’s off in production, and a package nobody imports. Its reading of the code was right every time. Nothing in a single file tells you whether that file is ever served, so we made reachability a required field with evidence: deployed, not deployed or unknown, plus the proof.
  • It recalls framework behavior instead of checking it. It decided a middleware covered only some routes when it covered the whole router, and reasoned about secret precedence for the wrong framework version until a dead string became a phantom High. Framework behavior now gets verified against the pinned version in the repository.
  • It inflates bugs only an insider can reach. CI injection that needs write access to the repository and container hardening with no way in both arrived as High. Severity now comes from our own risk matrix, capped by reachability.
  • It believes its own write-up. One impact statement claimed nineteen leaked fields. The refuter checked and found seven.

What the evidence looks like. Two small changes did most of the work. The first is a verdict record the model has to fill in, where reachability is a field with proof instead of an assumption. Here is the shape we use, with a made-up finding:

{
  "finding": "graphql-missing-authz  src/notes/resolver.ts:42",
  "verdict": "overturned",
  "confidence": "high",
  "reachability": "not_deployed",
  "evidence": [
    "NotesResolver is missing from the module's providers (notes.module.ts:12)",
    "updateNote does not appear in the generated schema.graphql"
  ],
  "severity": null
}

The second is the refuter’s instruction, which is short on purpose:

You are reviewing a confirmed finding. Your only job is to show it is wrong. Check, in order: is the entrypoint actually registered and deployed; did the first reviewer miss a control, such as global middleware, a guard, ORM scoping or upstream validation; is the severity higher than who can really reach it. Return the strongest reason it is wrong with file and line evidence, or “could not refute” with what you checked.

The severity habit tells you the most. On the findings that mattered, one user reading another user’s records or SQL injection reachable from a public endpoint, the model was reliably right. When we re-scored confirmed findings on our own matrix, 56% came down and only 2% went up, and every one that went up was data exposure. The model knows what a bug looks like. It doesn’t know your environment.

Mindset shift: the model answers "is this pattern a bug" well. Spend people on "does this code run, and who can reach it".

🤖 Direction 2: AI that does the team’s legwork #

Triage showed us the shape of the problem. The same shape came back when we pointed agents at our own work, with higher stakes, because now the agents could change things.

The upside was bigger than triage. A small team can now do work that used to need a much larger one, mostly by fanning agents out in parallel. We hunted one bug class across our backend repositories with one agent per repository and a second wave validating every hit. As a test, it re-found 100% of the instances we already knew about without being told where to look, and surfaced new ones worth acting on. We checked closed tickets the same way, one agent per ticket reading the code behind it. Most fixes had landed, and most of the ones that hadn’t were leaked secrets removed from the code but never rotated. Whole-backlog passes that used to take a quarter took a day, which turned “we have a backlog” into a number we could plan against. And when we replaced our ticket-tracker workflow with our own vulnerability management system, agents helped build it and now help run it.

The risks were about writing. Most of what we learned here, we learned by getting it slightly wrong first.

Agents analyze, scripts write. Agents produce verdicts as files, and a plain, deterministic script applies them to the system of record. An agent never writes there directly, so every change is logged, repeatable and reversible.

Gate bulk changes in code. We once described a two-unit pilot to an agent workflow and passed the pilot filter as a parameter. The parameter never arrived, and the workflow applied the entire run before anyone spot-checked. They were reversible, so the cost was cleanup time. A pilot is now a separate run with the ids written into it.

Normalize anything you combine. One stage reported confidence from 0 to 100 and another from 0 to 1, and the blended score sent every finding to human review until we noticed.

Treat transcripts as output. An agent that greps a repository can print a secret into its transcript as easily as into a terminal, so secrets in agent work follow the same rules as anywhere else: never echoed, never passed as an argument, rotated if one ever shows up.

Rules of engagement: agents can read anything and propose anything. Only code you can audit gets to change a real system.

🛡️ Direction 3: AI that developers use #

The third direction turns the question around. Everything above is about agents we run. Our developers run their own, and the same two questions apply: what is it assuming, and what is it allowed to do?

Developers adopted coding agents faster than our policies did, and that’s fine, because the agents are good. But an agent is more than autocomplete. It runs shell commands, holds tokens and reaches repositories, the data warehouse and the cloud through MCP servers, dozens of actions before it answers. Approving every tool up front fails, because the tools change weekly and blocking too much pushes people to personal accounts where you see nothing. So we’re rolling out an AI governance platform built around one goal: spend friction only where the risk is real, and keep everything else fast.

Graduate every control. Each control can run at four levels: monitor, alert, ask and block. We start at monitor, learn what normal looks like, and raise each control one level at a time. Going straight to block produces exceptions for everything.

Give agents a permission baseline. Most agent actions are harmless and stay allowed: reading code, running tests, local builds. A dozen are risky enough to ask the developer in the moment: package installs, infrastructure changes, pushing code. A few are denied outright: reading credential files, destructive shell commands, disabling security controls. “Ask” carries the weight, because it keeps a person in the loop exactly where it matters.

Replace standing access with just-in-time access. Block powerful MCP servers by default. A developer requests access with a reason and a scope, and an approver can narrow and time-box it from a Slack message, while a minimal envelope approves itself. Limited warehouse access is the default, and all tables for a few hours when an investigation needs it.

Publish a general-use tier. An approved set of models and servers anyone can use without asking is what lets you deny new and unknown platforms by default without stranding anyone.

Write data rules by destination. In healthcare, protected health information goes only to model endpoints covered by a business associate agreement, and secrets go nowhere. The check has to run before anything leaves the machine, on every turn of the agent loop, and cover tool results as well as what the developer typed.

Treat everything an agent reads as untrusted. Prompt injection arrives through issues, pull requests, READMEs, web pages and MCP responses. Inspect what comes back from tools, and watch for the chain no single step reveals: untrusted input, then sensitive access, then an export.

Put an identity on every session. Single sign-on makes every action auditable and lets policy differ by role. Personal accounts get flagged, people get steered to the managed one, and we block where we can enforce it.

Mindset shift: decide which dozen agent actions deserve a , and make everything else fast.

Bringing It Together #

In each direction, the model did the reading, the fan-out and the first draft well, and in each one the trust came from somewhere else. For triage it came from evidence fields and a refuter. For our own agents it came from deterministic writes and pilots written in code. For developers’ agents it came from graduated controls and a few well-placed s.

Use the AI for the volume. Use structure for the trust.

Get that structure right and a small security team does the work of a large one, while developers keep moving at the speed the tools make possible.

── more in #artificial-intelligence 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-in-security-volum…] indexed:0 read:10min 2026-10-05 · —