{"slug": "a-sanity-check-for-ai-generated-cyber-attack-reconstructions", "title": "A Sanity Check for AI-Generated Cyber Attack Reconstructions", "summary": "A developer built Cyber Autopsy, a cyber-forensics investigation system that uses Sanity as a structured content layer and an MCP-compatible server to let an AI reconstruct attack timelines from fragmented reports without presenting guesses as facts. The project models evidence, events, claims, relationships, incidents, and sources as linked documents, and labels conclusions as confirmed, inferred, attempted, or unknown. The repository includes a Python benchmark, Sanity Studio, importer, MCP server, and a forensic-themed web UI.", "body_md": "*This is a submission for the [Sanity Challenge, Path One: Ship an Agent That Queries Real Content](https://dev.to/challenges/sanity-2026-09-16)*\n\nCyber Autopsy is a cyber-forensics investigation room for asking structured\n\nquestions about an incident: what happened, what came before it, which evidence\n\nsupports it, what remains unknown, and what was knowable at a specific point in\n\ntime.\n\nThe project started as a Kaggle-style benchmark for evaluating whether an AI\n\ncan reconstruct an attack from fragmented reports without presenting guesses as\n\nfacts. I added Sanity as the structured content layer, an MCP-compatible\n\ninvestigation server, and a forensic-themed web interface.\n\nThe core model preserves the relationships that make an investigation useful:\n\n```\nEvidence ──supports──> Event ──precedes/enables/causes──> Event\nEvidence ──supports/contradicts──> Claim\n```\n\nThe UI presents each case as a small investigation console. Investigators can\n\nselect a case, read its incident description, ask a question, and inspect the\n\nresult as a reconstruction, relationship graph, uncertainty list,\n\ncontradictory claims, and expandable evidence with source provenance.\n\nCybersecurity autopsy is an evidence-reconstruction problem. Analysts rarely\n\nreceive one perfect timeline; they receive authentication logs, endpoint\n\nobservations, threat reports, analyst notes, negative findings, and competing\n\nclaims. The difficult part is deciding how those pieces relate without\n\noverstating what they prove.\n\nCyber Autopsy uses Sanity to make that reasoning inspectable. An analyst can\n\nfollow an event back to its supporting evidence, move through the attack graph,\n\ncheck whether a claim is contradicted, and distinguish `confirmed`, `inferred`,\n\n`attempted`, and `unknown`. This is useful for:\n\nThe result is closer to a digital incident autopsy than a generic search box:\n\nthe system preserves the chain of evidence, the gaps in the record, and the\n\nreasoning boundaries around each conclusion.\n\nRun the project locally:\n\n```\nnpm run dev\n```\n\nThen open [http://localhost:3000](http://localhost:3000).\n\nThe repository contains the Python benchmark, Sanity Studio, importer, MCP\n\nserver, local UI, tests, and documentation.\n\nRepository: [https://github.com/ujjavala/cyber-autopsy](https://github.com/ujjavala/cyber-autopsy)\n\nThe most relevant implementation files are:\n\n```\nstudio-cyber-autopsy/schemaTypes/index.ts  Sanity content model\nscripts/sanity-seed.mjs                    repeatable JSONL importer\nagent/sanity-context.mjs                   scoped GROQ investigation context\nagent/mcp-server.mjs                       MCP tool surface\nweb/app.js                                 investigation UI behavior\nweb/style.css                              cyber-forensics console theme\n```\n\nSanity gave the project a durable, queryable content layer for the benchmark's\n\ninvestigation graph. Before the integration, the dataset was primarily a group\n\nof local JSONL files. That was reproducible, but it made the content harder to\n\ninspect, connect, and reuse across an agent, a Studio, and a browser UI.\n\nWith Sanity, the same investigation content is modeled as linked documents.\n\nEvidence, events, claims, relationships, incidents, and sources can be queried\n\nindependently or traversed together. This helped in four practical ways:\n\nSanity therefore acts as the investigation knowledge base, while the context\n\nlayer acts as the reasoning boundary. The agent is not asked to remember an\n\nincident or invent a timeline; it queries structured content and returns a\n\ncalibrated reconstruction.\n\nInstead of storing one large incident document, I modeled the investigation as\n\nseparate document types. This matches the way a forensic analyst works: an\n\nincident is the case context, evidence is an observation, an event is a\n\nreconstruction, a claim is an assertion, and a relationship explains how two\n\nevents connect.\n\nThe separation gives each kind of information a clear validation and query\n\nboundary. For example, an `event` has a status and confidence, while an\n\n`evidence` document has a timestamp, observation type, entities, and source\n\nreferences. A `relationship` has a source event, target event, relation type,\n\nand supporting evidence.\n\nThe most important Sanity feature here is the reference field. Evidence does\n\nnot copy an event's text, and a relationship does not copy both event records.\n\nThey point to the canonical documents:\n\n```\ndefineField({\n  name: 'sourceEvent',\n  type: 'reference',\n  to: [{type: 'event'}],\n  validation: (rule) => rule.required(),\n})\n\ndefineField({\n  name: 'evidence',\n  type: 'array',\n  of: [defineArrayMember({type: 'reference', to: [{type: 'evidence'}]})],\n})\n```\n\nThis turns Sanity into a navigable investigation graph. The agent can traverse\n\n`case -> incident -> event -> evidence -> source`, then separately retrieve\n\n`relationship` and `claim` documents for the same incident. In the UI, that\n\nbecomes a reconstruction with visible provenance instead of an unsupported\n\nparagraph.\n\nThe Studio schema validates required identifiers and references, constrains\n\nconfidence to the range `0..1`, and uses controlled options for evidence and\n\nevent status. The available status taxonomy is deliberately forensic:\n\n```\nconfirmed · inferred · attempted · failed · unknown\n```\n\nClaims additionally support `contradicted`. These values are used in the UI to\n\nseparate established facts from uncertain or failed steps. This prevents a\n\nmissing observation from silently becoming a confirmed event and makes\n\ncalibrated uncertainty part of the content model.\n\nThe context layer uses GROQ projections to request only the fields required by\n\nan investigation. The `->` operator dereferences related documents so one\n\nresponse can contain the case, incident summary, evidence, and source metadata\n\nwithout the browser making a request for every individual record.\n\nFor example, the case query projects the incident and joins its evidence and\n\nevents:\n\n```\n*[_type == \"investigationCase\" && caseId == $caseId][0] {\n  caseId, mode, temporalCutoff, framingActor,\n  incident->{incidentId, title, description, actor, impact},\n  evidence[]->{evidenceId, timestamp, type, description, status,\n    source[]->{sourceId, title, publisher, url}},\n  events[]->{eventId, description, timestamp, status,\n    evidence[]->{evidenceId}}\n}\n```\n\nThis is useful for security analysis because the response remains structured.\n\nThe agent can filter on timestamps, statuses, evidence IDs, and relationship\n\ntypes instead of trying to recover those distinctions from prose.\n\nQueries are parameterized with `caseId` and `incidentRef`. The selected case\n\ndefines the visible evidence set, and the context derives a set of allowed\n\nevidence IDs before returning events, relationships, or claims. A relationship\n\nis included only when its supporting evidence and both endpoint events are in\n\nscope.\n\nThis is how the implementation handles temporal safety. `CASE-004` is not just\n\na label in the UI; its Sanity references define the evidence boundary. A\n\nday-one question can also apply an additional D1 filter before the result is\n\nrendered. The model cannot accidentally cite a later ransomware event as if it\n\nwere known on day one.\n\nThe Node context uses Sanity's Content API with the project ID, dataset, GROQ\n\nquery, and optional server-side token. The token is loaded from `.env` on the\n\nserver and is never exposed to the browser. The browser talks to the local web\n\nserver, while the local server talks to Sanity.\n\nThat separation gives the project a safer shape:\n\n```\nBrowser UI → local investigation API → Sanity Content API\n                                      ↑\n                              server-side token\n```\n\nThe same API boundary also makes failures explicit. A missing case produces a\n\nclear error, while an empty evidence set is handled as an empty investigation\n\nscope rather than causing a client-side `length` exception.\n\nThe project includes a standalone Sanity Studio configured for the same project\n\nand `production` dataset. The Structure Tool provides the editing workspace\n\nfor the seven document types, with previews showing useful forensic identifiers\n\nsuch as `CASE-004`, `E-021`, or an event description.\n\nThe Vision Tool is enabled for inspecting and testing GROQ directly in Studio.\n\nThat is useful when developing an investigation query: I can verify a case\n\nprojection, inspect references, and test a cutoff query against the real\n\ndataset before wiring it into the MCP context.\n\nThe importer uses Sanity mutations in batches rather than hand-editing hundreds\n\nof documents. It creates or replaces the normalized documents in dependency\n\norder, then applies reverse-reference patches for evidence-to-event,\n\nevidence-to-claim, and contradiction links.\n\nThe import is repeatable and testable:\n\n```\nnpm run sanity:validate       # validate source inventory\nnode scripts/sanity-seed.mjs --dry-run\nnpm run sanity:seed           # write the content to Sanity\n```\n\nThis gives the benchmark a reproducible migration path while keeping the live\n\nknowledge base editable in Studio. It also means the dataset can be rebuilt if\n\nthe schema gains another forensic field later.\n\nSanity stores and retrieves the content; the MCP server exposes investigation\n\ncapabilities to an agent host. The server provides `list_investigation_cases`\n\nand `investigate_case(caseId, question)`. Internally, those tools call the same\n\nSanity context used by the web UI.\n\nThis division is intentional. Sanity is responsible for content modeling,\n\nrelationships, querying, and provenance. The context layer is responsible for\n\ncase scoping and graph filtering. The MCP layer is responsible for making those\n\ngrounded operations available to an agent. The agent can therefore reason over\n\nretrieved evidence without receiving unrestricted access to the whole dataset.\n\nSanity stores seven document types:\n\n`incident`: case narrative, actor, impact, confidence, and source references`source`: publisher, URL, publication date, type, and reliability` evidence`: an observation with timestamp, status, description, and provenance`event`: a reconstructed action or state with evidence references` relationship`: a typed link between events` claim`: an assertion that can be supported or contradicted` investigationCase`: a benchmark scope with mode, cutoff, and actor framing\nThe schema is defined with Sanity's typed helpers:\n\n```\ndefineField({\n  name: 'supportsEvents',\n  type: 'array',\n  of: [defineArrayMember({type: 'reference', to: [{type: 'event'}]})],\n})\n\ndefineField({\n  name: 'relationship',\n  type: 'string',\n  options: {list: ['precedes', 'enables', 'causes', 'depends_on']},\n})\n```\n\nIDs from the benchmark remain explicit and stable: `INC-001`, `E-001`, `N01`,\n\nand `CASE-004`. This makes the imported content easy to inspect in Studio and\n\nkeeps evidence references understandable in the investigation output.\n\nThe JSONL files under `data/` remain the reproducible source of truth. The\n\nimporter maps them into Sanity documents:\n\n```\nincidents.jsonl       → incident\nsources.jsonl         → source\nevidence.jsonl        → evidence\nattack_graphs.nodes   → event + claim\nattack_graphs.edges   → relationship\nbenchmark_cases.jsonl → investigationCase\n```\n\nThe importer writes documents in dependency order and applies reverse evidence\n\nreferences in a second pass because events and evidence point to each other.\n\n```\nnpm run sanity:validate\nnode scripts/sanity-seed.mjs --dry-run\nnpm run sanity:seed\n```\n\nThe current Sanity dataset contains 445 imported Cyber Autopsy documents in\n\nproduction, plus 12 existing documents in the project.\n\n`agent/sanity-context.mjs` uses scoped GROQ queries. It does not download the\n\nentire dataset for every question:\n\n```\n*[_type == \"investigationCase\" && caseId == $caseId][0] {\n  caseId, mode, temporalCutoff, framingActor,\n  incident->{incidentId, title, actor, actorType},\n  evidence[]->{evidenceId, timestamp, type, description, status,\n    source[]->{sourceId, title, publisher, url}},\n  events[]->{eventId, description, timestamp, status,\n    evidence[]->{evidenceId}}\n}\n```\n\nThe context then resolves relationships and claims and filters them against the\n\ncase's allowed evidence IDs. That filtering is the temporal safety boundary:\n\n`CASE-004` can answer what was knowable at the end of day one without leaking\n\nlater ransomware evidence from the full case.\n\nThe MCP server exposes two tools:\n\n```\nlist_investigation_cases\ninvestigate_case(caseId, question)\n```\n\nThe investigation function maps question intent to structured graph operations:\n\npredecessors, enablers, supporting evidence, statuses, cutoff checks, actor\n\nframing, and contradictions. The result is deterministic, grounded in Sanity\n\ncontent, and designed to be passed to a model without requiring another model\n\nAPI key in this repository.\n\n```\nProject ID: 41l9o4xn\nDataset: production\nOrganization ID: oqf9m6vy6\n```\n\nThe public Sanity project details are intentionally included so the structured\n\ncontent model can be inspected. The server-side `SANITY_API_TOKEN` stays in\n\n`.env`, is ignored by Git, and is never sent to the browser.\n\nThe local agent is available through the MCP server:\n\n```\nnpm run agent:mcp\n```\n\nThe browser UI uses the same investigation context and exposes the workflow in\n\na more approachable way:\n\nUseful sessions to capture for the final submission are:\n\n```\nCASE-001: How did the attacker move from initial access to ransomware deployment?\nCASE-001: What evidence supports the Rclone event?\nCASE-004: What could we have known at the end of day one?\nCASE-011/012: What does the evidence establish regardless of actor framing?\nCASE-001: Are there conflicting claims?\n```\n\nThe current build has been checked with:\n\n```\n.venv/bin/pytest -q\nnpm run sanity:validate\nnpm --prefix studio-cyber-autopsy run build -- --no-auto-updates\n```\n\nThe browser smoke checks cover the full investigation, the CASE-004 temporal\n\ncutoff, case descriptions, evidence rendering, and zero browser console errors.", "url": "https://wpnews.pro/news/a-sanity-check-for-ai-generated-cyber-attack-reconstructions", "canonical_source": "https://dev.to/ujja/a-sanity-check-for-ai-generated-cyber-attack-reconstructions-31bm", "published_at": "2026-10-03 08:31:17+00:00", "updated_at": "2026-10-03 08:38:06.889194+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-tools", "structured-data", "ai-safety"], "entities": ["Cyber Autopsy", "Sanity", "MCP", "GitHub", "Kaggle"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-sanity-check-for-ai-generated-cyber-attack-reconstructions", "markdown": "https://wpnews.pro/news/a-sanity-check-for-ai-generated-cyber-attack-reconstructions.md", "text": "https://wpnews.pro/news/a-sanity-check-for-ai-generated-cyber-attack-reconstructions.txt", "jsonld": "https://wpnews.pro/news/a-sanity-check-for-ai-generated-cyber-attack-reconstructions.jsonld"}}