A Sanity Check for AI-Generated Cyber Attack Reconstructions A developer built Cyber Autopsy, a cyber-forensics investigation system that uses Sanity as a structured content layer and an MCP-compatible server to let an AI reconstruct attack timelines from fragmented reports without presenting guesses as facts. The project models evidence, events, claims, relationships, incidents, and sources as linked documents, and labels conclusions as confirmed, inferred, attempted, or unknown. The repository includes a Python benchmark, Sanity Studio, importer, MCP server, and a forensic-themed web UI. This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content https://dev.to/challenges/sanity-2026-09-16 Cyber Autopsy is a cyber-forensics investigation room for asking structured questions about an incident: what happened, what came before it, which evidence supports it, what remains unknown, and what was knowable at a specific point in time. The project started as a Kaggle-style benchmark for evaluating whether an AI can reconstruct an attack from fragmented reports without presenting guesses as facts. I added Sanity as the structured content layer, an MCP-compatible investigation server, and a forensic-themed web interface. The core model preserves the relationships that make an investigation useful: Evidence ──supports── Event ──precedes/enables/causes── Event Evidence ──supports/contradicts── Claim The UI presents each case as a small investigation console. Investigators can select a case, read its incident description, ask a question, and inspect the result as a reconstruction, relationship graph, uncertainty list, contradictory claims, and expandable evidence with source provenance. Cybersecurity autopsy is an evidence-reconstruction problem. Analysts rarely receive one perfect timeline; they receive authentication logs, endpoint observations, threat reports, analyst notes, negative findings, and competing claims. The difficult part is deciding how those pieces relate without overstating what they prove. Cyber Autopsy uses Sanity to make that reasoning inspectable. An analyst can follow an event back to its supporting evidence, move through the attack graph, check whether a claim is contradicted, and distinguish confirmed , inferred , attempted , and unknown . This is useful for: The result is closer to a digital incident autopsy than a generic search box: the system preserves the chain of evidence, the gaps in the record, and the reasoning boundaries around each conclusion. Run the project locally: npm run dev Then open http://localhost:3000 http://localhost:3000 . The repository contains the Python benchmark, Sanity Studio, importer, MCP server, local UI, tests, and documentation. Repository: https://github.com/ujjavala/cyber-autopsy https://github.com/ujjavala/cyber-autopsy The most relevant implementation files are: studio-cyber-autopsy/schemaTypes/index.ts Sanity content model scripts/sanity-seed.mjs repeatable JSONL importer agent/sanity-context.mjs scoped GROQ investigation context agent/mcp-server.mjs MCP tool surface web/app.js investigation UI behavior web/style.css cyber-forensics console theme Sanity gave the project a durable, queryable content layer for the benchmark's investigation graph. Before the integration, the dataset was primarily a group of local JSONL files. That was reproducible, but it made the content harder to inspect, connect, and reuse across an agent, a Studio, and a browser UI. With Sanity, the same investigation content is modeled as linked documents. Evidence, events, claims, relationships, incidents, and sources can be queried independently or traversed together. This helped in four practical ways: Sanity therefore acts as the investigation knowledge base, while the context layer acts as the reasoning boundary. The agent is not asked to remember an incident or invent a timeline; it queries structured content and returns a calibrated reconstruction. Instead of storing one large incident document, I modeled the investigation as separate document types. This matches the way a forensic analyst works: an incident is the case context, evidence is an observation, an event is a reconstruction, a claim is an assertion, and a relationship explains how two events connect. The separation gives each kind of information a clear validation and query boundary. For example, an event has a status and confidence, while an evidence document has a timestamp, observation type, entities, and source references. A relationship has a source event, target event, relation type, and supporting evidence. The most important Sanity feature here is the reference field. Evidence does not copy an event's text, and a relationship does not copy both event records. They point to the canonical documents: defineField { name: 'sourceEvent', type: 'reference', to: {type: 'event'} , validation: rule = rule.required , } defineField { name: 'evidence', type: 'array', of: defineArrayMember {type: 'reference', to: {type: 'evidence'} } , } This turns Sanity into a navigable investigation graph. The agent can traverse case - incident - event - evidence - source , then separately retrieve relationship and claim documents for the same incident. In the UI, that becomes a reconstruction with visible provenance instead of an unsupported paragraph. The Studio schema validates required identifiers and references, constrains confidence to the range 0..1 , and uses controlled options for evidence and event status. The available status taxonomy is deliberately forensic: confirmed · inferred · attempted · failed · unknown Claims additionally support contradicted . These values are used in the UI to separate established facts from uncertain or failed steps. This prevents a missing observation from silently becoming a confirmed event and makes calibrated uncertainty part of the content model. The context layer uses GROQ projections to request only the fields required by an investigation. The - operator dereferences related documents so one response can contain the case, incident summary, evidence, and source metadata without the browser making a request for every individual record. For example, the case query projects the incident and joins its evidence and events: type == "investigationCase" && caseId == $caseId 0 { caseId, mode, temporalCutoff, framingActor, incident- {incidentId, title, description, actor, impact}, evidence - {evidenceId, timestamp, type, description, status, source - {sourceId, title, publisher, url}}, events - {eventId, description, timestamp, status, evidence - {evidenceId}} } This is useful for security analysis because the response remains structured. The agent can filter on timestamps, statuses, evidence IDs, and relationship types instead of trying to recover those distinctions from prose. Queries are parameterized with caseId and incidentRef . The selected case defines the visible evidence set, and the context derives a set of allowed evidence IDs before returning events, relationships, or claims. A relationship is included only when its supporting evidence and both endpoint events are in scope. This is how the implementation handles temporal safety. CASE-004 is not just a label in the UI; its Sanity references define the evidence boundary. A day-one question can also apply an additional D1 filter before the result is rendered. The model cannot accidentally cite a later ransomware event as if it were known on day one. The Node context uses Sanity's Content API with the project ID, dataset, GROQ query, and optional server-side token. The token is loaded from .env on the server and is never exposed to the browser. The browser talks to the local web server, while the local server talks to Sanity. That separation gives the project a safer shape: Browser UI → local investigation API → Sanity Content API ↑ server-side token The same API boundary also makes failures explicit. A missing case produces a clear error, while an empty evidence set is handled as an empty investigation scope rather than causing a client-side length exception. The project includes a standalone Sanity Studio configured for the same project and production dataset. The Structure Tool provides the editing workspace for the seven document types, with previews showing useful forensic identifiers such as CASE-004 , E-021 , or an event description. The Vision Tool is enabled for inspecting and testing GROQ directly in Studio. That is useful when developing an investigation query: I can verify a case projection, inspect references, and test a cutoff query against the real dataset before wiring it into the MCP context. The importer uses Sanity mutations in batches rather than hand-editing hundreds of documents. It creates or replaces the normalized documents in dependency order, then applies reverse-reference patches for evidence-to-event, evidence-to-claim, and contradiction links. The import is repeatable and testable: npm run sanity:validate validate source inventory node scripts/sanity-seed.mjs --dry-run npm run sanity:seed write the content to Sanity This gives the benchmark a reproducible migration path while keeping the live knowledge base editable in Studio. It also means the dataset can be rebuilt if the schema gains another forensic field later. Sanity stores and retrieves the content; the MCP server exposes investigation capabilities to an agent host. The server provides list investigation cases and investigate case caseId, question . Internally, those tools call the same Sanity context used by the web UI. This division is intentional. Sanity is responsible for content modeling, relationships, querying, and provenance. The context layer is responsible for case scoping and graph filtering. The MCP layer is responsible for making those grounded operations available to an agent. The agent can therefore reason over retrieved evidence without receiving unrestricted access to the whole dataset. Sanity stores seven document types: incident : case narrative, actor, impact, confidence, and source references source : publisher, URL, publication date, type, and reliability evidence : an observation with timestamp, status, description, and provenance event : a reconstructed action or state with evidence references relationship : a typed link between events claim : an assertion that can be supported or contradicted investigationCase : a benchmark scope with mode, cutoff, and actor framing The schema is defined with Sanity's typed helpers: defineField { name: 'supportsEvents', type: 'array', of: defineArrayMember {type: 'reference', to: {type: 'event'} } , } defineField { name: 'relationship', type: 'string', options: {list: 'precedes', 'enables', 'causes', 'depends on' }, } IDs from the benchmark remain explicit and stable: INC-001 , E-001 , N01 , and CASE-004 . This makes the imported content easy to inspect in Studio and keeps evidence references understandable in the investigation output. The JSONL files under data/ remain the reproducible source of truth. The importer maps them into Sanity documents: incidents.jsonl → incident sources.jsonl → source evidence.jsonl → evidence attack graphs.nodes → event + claim attack graphs.edges → relationship benchmark cases.jsonl → investigationCase The importer writes documents in dependency order and applies reverse evidence references in a second pass because events and evidence point to each other. npm run sanity:validate node scripts/sanity-seed.mjs --dry-run npm run sanity:seed The current Sanity dataset contains 445 imported Cyber Autopsy documents in production, plus 12 existing documents in the project. agent/sanity-context.mjs uses scoped GROQ queries. It does not download the entire dataset for every question: type == "investigationCase" && caseId == $caseId 0 { caseId, mode, temporalCutoff, framingActor, incident- {incidentId, title, actor, actorType}, evidence - {evidenceId, timestamp, type, description, status, source - {sourceId, title, publisher, url}}, events - {eventId, description, timestamp, status, evidence - {evidenceId}} } The context then resolves relationships and claims and filters them against the case's allowed evidence IDs. That filtering is the temporal safety boundary: CASE-004 can answer what was knowable at the end of day one without leaking later ransomware evidence from the full case. The MCP server exposes two tools: list investigation cases investigate case caseId, question The investigation function maps question intent to structured graph operations: predecessors, enablers, supporting evidence, statuses, cutoff checks, actor framing, and contradictions. The result is deterministic, grounded in Sanity content, and designed to be passed to a model without requiring another model API key in this repository. Project ID: 41l9o4xn Dataset: production Organization ID: oqf9m6vy6 The public Sanity project details are intentionally included so the structured content model can be inspected. The server-side SANITY API TOKEN stays in .env , is ignored by Git, and is never sent to the browser. The local agent is available through the MCP server: npm run agent:mcp The browser UI uses the same investigation context and exposes the workflow in a more approachable way: Useful sessions to capture for the final submission are: CASE-001: How did the attacker move from initial access to ransomware deployment? CASE-001: What evidence supports the Rclone event? CASE-004: What could we have known at the end of day one? CASE-011/012: What does the evidence establish regardless of actor framing? CASE-001: Are there conflicting claims? The current build has been checked with: .venv/bin/pytest -q npm run sanity:validate npm --prefix studio-cyber-autopsy run build -- --no-auto-updates The browser smoke checks cover the full investigation, the CASE-004 temporal cutoff, case descriptions, evidence rendering, and zero browser console errors.