cd /news/ai-agents/a-sanity-check-for-ai-generated-cybe… · home › topics › ai-agents › article
[ARTICLE · art-144357] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

A Sanity Check for AI-Generated Cyber Attack Reconstructions

A developer built Cyber Autopsy, a cyber-forensics investigation system that uses Sanity as a structured content layer and an MCP-compatible server to let an AI reconstruct attack timelines from fragmented reports without presenting guesses as facts. The project models evidence, events, claims, relationships, incidents, and sources as linked documents, and labels conclusions as confirmed, inferred, attempted, or unknown. The repository includes a Python benchmark, Sanity Studio, importer, MCP server, and a forensic-themed web UI.

by read9 min views1 publishedOct 3, 2026

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

Cyber Autopsy is a cyber-forensics investigation room for asking structured

questions about an incident: what happened, what came before it, which evidence

supports it, what remains unknown, and what was knowable at a specific point in

time.

The project started as a Kaggle-style benchmark for evaluating whether an AI

can reconstruct an attack from fragmented reports without presenting guesses as

facts. I added Sanity as the structured content layer, an MCP-compatible

investigation server, and a forensic-themed web interface.

The core model preserves the relationships that make an investigation useful:

Evidence ──supports──> Event ──precedes/enables/causes──> Event
Evidence ──supports/contradicts──> Claim

The UI presents each case as a small investigation console. Investigators can

select a case, read its incident description, ask a question, and inspect the

result as a reconstruction, relationship graph, uncertainty list,

contradictory claims, and expandable evidence with source provenance.

Cybersecurity autopsy is an evidence-reconstruction problem. Analysts rarely

receive one perfect timeline; they receive authentication logs, endpoint

observations, threat reports, analyst notes, negative findings, and competing

claims. The difficult part is deciding how those pieces relate without

overstating what they prove.

Cyber Autopsy uses Sanity to make that reasoning inspectable. An analyst can

follow an event back to its supporting evidence, move through the attack graph,

check whether a claim is contradicted, and distinguish confirmed, inferred,

attempted, and unknown. This is useful for:

The result is closer to a digital incident autopsy than a generic search box:

the system preserves the chain of evidence, the gaps in the record, and the

reasoning boundaries around each conclusion.

Run the project locally:

npm run dev

Then open http://localhost:3000.

The repository contains the Python benchmark, Sanity Studio, importer, MCP

server, local UI, tests, and documentation.

Repository: https://github.com/ujjavala/cyber-autopsy

The most relevant implementation files are:

studio-cyber-autopsy/schemaTypes/index.ts  Sanity content model
scripts/sanity-seed.mjs                    repeatable JSONL importer
agent/sanity-context.mjs                   scoped GROQ investigation context
agent/mcp-server.mjs                       MCP tool surface
web/app.js                                 investigation UI behavior
web/style.css                              cyber-forensics console theme

Sanity gave the project a durable, queryable content layer for the benchmark's

investigation graph. Before the integration, the dataset was primarily a group

of local JSONL files. That was reproducible, but it made the content harder to

inspect, connect, and reuse across an agent, a Studio, and a browser UI.

With Sanity, the same investigation content is modeled as linked documents.

Evidence, events, claims, relationships, incidents, and sources can be queried

independently or traversed together. This helped in four practical ways:

Sanity therefore acts as the investigation knowledge base, while the context

layer acts as the reasoning boundary. The agent is not asked to remember an

incident or invent a timeline; it queries structured content and returns a

calibrated reconstruction.

Instead of storing one large incident document, I modeled the investigation as

separate document types. This matches the way a forensic analyst works: an

incident is the case context, evidence is an observation, an event is a

reconstruction, a claim is an assertion, and a relationship explains how two

events connect.

The separation gives each kind of information a clear validation and query

boundary. For example, an event has a status and confidence, while an

evidence document has a timestamp, observation type, entities, and source

references. A relationship has a source event, target event, relation type,

and supporting evidence.

The most important Sanity feature here is the reference field. Evidence does

not copy an event's text, and a relationship does not copy both event records.

They point to the canonical documents:

defineField({
  name: 'sourceEvent',
  type: 'reference',
  to: [{type: 'event'}],
  validation: (rule) => rule.required(),
})

defineField({
  name: 'evidence',
  type: 'array',
  of: [defineArrayMember({type: 'reference', to: [{type: 'evidence'}]})],
})

This turns Sanity into a navigable investigation graph. The agent can traverse

case -> incident -> event -> evidence -> source, then separately retrieve

relationship and claim documents for the same incident. In the UI, that

becomes a reconstruction with visible provenance instead of an unsupported

paragraph.

The Studio schema validates required identifiers and references, constrains

confidence to the range 0..1, and uses controlled options for evidence and

event status. The available status taxonomy is deliberately forensic:

confirmed · inferred · attempted · failed · unknown

Claims additionally support contradicted. These values are used in the UI to

separate established facts from uncertain or failed steps. This prevents a

missing observation from silently becoming a confirmed event and makes

calibrated uncertainty part of the content model.

The context layer uses GROQ projections to request only the fields required by

an investigation. The -> operator dereferences related documents so one

response can contain the case, incident summary, evidence, and source metadata

without the browser making a request for every individual record.

For example, the case query projects the incident and joins its evidence and

events:

*[_type == "investigationCase" && caseId == $caseId][0] {
  caseId, mode, temporalCutoff, framingActor,
  incident->{incidentId, title, description, actor, impact},
  evidence[]->{evidenceId, timestamp, type, description, status,
    source[]->{sourceId, title, publisher, url}},
  events[]->{eventId, description, timestamp, status,
    evidence[]->{evidenceId}}
}

This is useful for security analysis because the response remains structured.

The agent can filter on timestamps, statuses, evidence IDs, and relationship

types instead of trying to recover those distinctions from prose.

Queries are parameterized with caseId and incidentRef. The selected case

defines the visible evidence set, and the context derives a set of allowed

evidence IDs before returning events, relationships, or claims. A relationship

is included only when its supporting evidence and both endpoint events are in

scope.

This is how the implementation handles temporal safety. CASE-004 is not just

a label in the UI; its Sanity references define the evidence boundary. A

day-one question can also apply an additional D1 filter before the result is

rendered. The model cannot accidentally cite a later ransomware event as if it

were known on day one.

The Node context uses Sanity's Content API with the project ID, dataset, GROQ

query, and optional server-side token. The token is loaded from .env on the

server and is never exposed to the browser. The browser talks to the local web

server, while the local server talks to Sanity.

That separation gives the project a safer shape:

Browser UI → local investigation API → Sanity Content API
                                      ↑
                              server-side token

The same API boundary also makes failures explicit. A missing case produces a

clear error, while an empty evidence set is handled as an empty investigation

scope rather than causing a client-side length exception.

The project includes a standalone Sanity Studio configured for the same project

and production dataset. The Structure Tool provides the editing workspace

for the seven document types, with previews showing useful forensic identifiers

such as CASE-004, E-021, or an event description.

The Vision Tool is enabled for inspecting and testing GROQ directly in Studio.

That is useful when developing an investigation query: I can verify a case

projection, inspect references, and test a cutoff query against the real

dataset before wiring it into the MCP context.

The importer uses Sanity mutations in batches rather than hand-editing hundreds

of documents. It creates or replaces the normalized documents in dependency

order, then applies reverse-reference patches for evidence-to-event,

evidence-to-claim, and contradiction links.

The import is repeatable and testable:

npm run sanity:validate       # validate source inventory
node scripts/sanity-seed.mjs --dry-run
npm run sanity:seed           # write the content to Sanity

This gives the benchmark a reproducible migration path while keeping the live

knowledge base editable in Studio. It also means the dataset can be rebuilt if

the schema gains another forensic field later.

Sanity stores and retrieves the content; the MCP server exposes investigation

capabilities to an agent host. The server provides list_investigation_cases

and investigate_case(caseId, question). Internally, those tools call the same

Sanity context used by the web UI.

This division is intentional. Sanity is responsible for content modeling,

relationships, querying, and provenance. The context layer is responsible for

case scoping and graph filtering. The MCP layer is responsible for making those

grounded operations available to an agent. The agent can therefore reason over

retrieved evidence without receiving unrestricted access to the whole dataset.

Sanity stores seven document types:

incident: case narrative, actor, impact, confidence, and source referencessource: publisher, URL, publication date, type, and reliability evidence: an observation with timestamp, status, description, and provenanceevent: a reconstructed action or state with evidence references relationship: a typed link between events claim: an assertion that can be supported or contradicted investigationCase: a benchmark scope with mode, cutoff, and actor framing The schema is defined with Sanity's typed helpers:

defineField({
  name: 'supportsEvents',
  type: 'array',
  of: [defineArrayMember({type: 'reference', to: [{type: 'event'}]})],
})

defineField({
  name: 'relationship',
  type: 'string',
  options: {list: ['precedes', 'enables', 'causes', 'depends_on']},
})

IDs from the benchmark remain explicit and stable: INC-001, E-001, N01,

and CASE-004. This makes the imported content easy to inspect in Studio and

keeps evidence references understandable in the investigation output.

The JSONL files under data/ remain the reproducible source of truth. The

importer maps them into Sanity documents:

incidents.jsonl       → incident
sources.jsonl         → source
evidence.jsonl        → evidence
attack_graphs.nodes   → event + claim
attack_graphs.edges   → relationship
benchmark_cases.jsonl → investigationCase

The importer writes documents in dependency order and applies reverse evidence

references in a second pass because events and evidence point to each other.

npm run sanity:validate
node scripts/sanity-seed.mjs --dry-run
npm run sanity:seed

The current Sanity dataset contains 445 imported Cyber Autopsy documents in

production, plus 12 existing documents in the project.

agent/sanity-context.mjs uses scoped GROQ queries. It does not download the

entire dataset for every question:

*[_type == "investigationCase" && caseId == $caseId][0] {
  caseId, mode, temporalCutoff, framingActor,
  incident->{incidentId, title, actor, actorType},
  evidence[]->{evidenceId, timestamp, type, description, status,
    source[]->{sourceId, title, publisher, url}},
  events[]->{eventId, description, timestamp, status,
    evidence[]->{evidenceId}}
}

The context then resolves relationships and claims and filters them against the

case's allowed evidence IDs. That filtering is the temporal safety boundary:

CASE-004 can answer what was knowable at the end of day one without leaking

later ransomware evidence from the full case.

The MCP server exposes two tools:

list_investigation_cases
investigate_case(caseId, question)

The investigation function maps question intent to structured graph operations:

predecessors, enablers, supporting evidence, statuses, cutoff checks, actor

framing, and contradictions. The result is deterministic, grounded in Sanity

content, and designed to be passed to a model without requiring another model

API key in this repository.

Project ID: 41l9o4xn
Dataset: production
Organization ID: oqf9m6vy6

The public Sanity project details are intentionally included so the structured

content model can be inspected. The server-side SANITY_API_TOKEN stays in

.env, is ignored by Git, and is never sent to the browser.

The local agent is available through the MCP server:

npm run agent:mcp

The browser UI uses the same investigation context and exposes the workflow in

a more approachable way:

Useful sessions to capture for the final submission are:

CASE-001: How did the attacker move from initial access to ransomware deployment?
CASE-001: What evidence supports the Rclone event?
CASE-004: What could we have known at the end of day one?
CASE-011/012: What does the evidence establish regardless of actor framing?
CASE-001: Are there conflicting claims?

The current build has been checked with:

.venv/bin/pytest -q
npm run sanity:validate
npm --prefix studio-cyber-autopsy run build -- --no-auto-updates

The browser smoke checks cover the full investigation, the CASE-004 temporal

cutoff, case descriptions, evidence rendering, and zero browser console errors.

── more in #ai-agents 4 stories · sorted by recency
── more on @cyber autopsy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-sanity-check-for-a…] indexed:0 read:9min 2026-10-03 · —