I Asked 10 AI Models to Reconstruct Real Cyber Attacks A developer built Cyber Autopsy, a Kaggle benchmark that tests whether AI models can reconstruct reported cyber incidents from partial evidence while citing support for each claim and flagging what remains uncertain. The pilot evaluates ten models, including Gemini 3.7 Flash, GPT-5.6 Sol, Claude Opus 5 and Qwen 3 Coder 480B, across seven tasks derived from four public incident reports, scoring them with a deterministic Evidence-Grounded Reconstruction Score alongside event recall, causal-link quality, evidence attribution and hallucination measures. The author cautions the results reflect reconstructions of reported incidents, not live attack behavior, and that evidence quality varies between the DFIR Report ransomware case and vendor-reported AI-actor campaigns. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Cyber incident reports rarely arrive as a clean, complete timeline. They describe what investigators or security teams could observe, mix those observations with analysis, and leave some questions unanswered. That is exactly where an AI-generated summary can become risky: it may sound convincing while quietly turning an inference into a fact or filling a gap with an event that the evidence never established. I built Cyber Autopsy to test a narrower, practical question: when given pieces of a reported cyber incident, can a model reconstruct what happened while showing which evidence supports each claim and admitting what remains uncertain? The model must produce a timeline, connect events with causal or temporal links, cite evidence, and distinguish confirmed activity from inference, failed attempts, contradictions, and unknowns. A plausible attack story is not enough; unsupported certainty should count against it. I chose this problem because incident reconstruction depends on more than recognizing familiar attack techniques. An analyst needs to know whether a step was observed or inferred, whether an attempt actually succeeded, and how one event led to another. Those distinctions can disappear in a fluent summary. Measuring them separately makes it easier to see whether a model is recovering the evidence or merely telling a likely-sounding story. The first evaluation uses seven tasks built from four public reports. They include a detailed human-operated ransomware intrusion and vendor-reported campaigns involving AI-assisted activity. These are useful real-world case studies, but they are not a controlled contest between human and AI attackers: the reports differ in detail, evidence source, and corroboration. I also included two versions of one case with identical evidence but different actor framing, to see whether that wording changes the model's reconstruction. This is a pilot, not a claim that AI attackers are more or less capable than people. It evaluates model reconstructions of reported incidents, not live attack behavior. The scorer is deterministic and reports an Evidence-Grounded Reconstruction Score EGRS , alongside event recall and precision, causal-link quality, evidence attribution, status accuracy, uncertainty calibration, and hallucination-related measures. The temporal cutoff and framing pair are exploratory comparisons; only the framing pair holds the evidence fixed. The seven tasks are built from four public incident reports, not invented scenarios: These are real reported incidents, but the evidence quality is not uniform: the RansomHub case is reconstructed from host and network telemetry described by The DFIR Report, while the AI-actor case studies rely on security-vendor reporting. The benchmark labels that distinction rather than treating the cases as equally observed or directly comparable. The short IDs are just labels: INC means the source incident, and CASE means the particular benchmark task. Each task gives the model an evidence packet and asks for the best-supported reconstruction, not a free-form guess. I ran ten models from several providers against the same seven Kaggle tasks: Gemini 3.7 Flash , Gemma 4 26B A4B , GLM-5 , Grok 4.20 Reasoning , GPT-5.6 Luna , GPT-5.6 Sol , GPT-5.4 mini , Claude Sonnet 5 , Claude Opus 5 , and Qwen 3 Coder 480B . The current benchmark view has a score for every model-task pair. Kaggle's overall score aggregates these seven tasks, which include related variants of the same incidents. The table is the Kaggle leaderboard snapshot fetched on 2 October 2026 , after duplicate and failing task attachments were removed and the earlier evaluated versions restored. CASE-001 through CASE-011 use v3; CASE-012 and CASE-013 use their republished v1 versions. Values are EGRS percentages Kaggle's 0-1 scores multiplied by 100 . EGRS rewards recovering supported events and links, citing evidence, and representing uncertainty, while penalizing unsupported events. It is specific to this evidence-reconstruction task, not a general measure of intelligence or cybersecurity ability. Kaggle's overall score now matches the equal-weight mean across the seven task rows. Since some tasks are related variants of the same incidents, this is descriptive rather than an independent-sample leaderboard. | Kaggle task | Gemini Flash | Gemma 4 | GPT-5.6 Luna | GLM-5 | Grok 4.20 | Claude Sonnet 5 | Claude Opus 5 | GPT-5.6 Sol | GPT-5.4 mini | Qwen 3 Coder | |---|---|---|---|---|---|---|---|---|---|---| | CASE-001: Full RansomHub intrusion https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-001-reconstruct-inc-001-full/3 | 70.55 | 82.38 | 71.27 | 81.31 | 75.71 | 73.75 | 78.14 | 78.06 | 66.06 | 70.90 | | CASE-002: GTG-1002 espionage campaign https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-002-reconstruct-inc-002-full/3 | 79.72 | 77.21 | 80.31 | 76.22 | 83.60 | 65.05 | 68.16 | 76.71 | 69.68 | 76.30 | | CASE-003: Reported data-extortion operation https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-003-reconstruct-inc-003-full/3 | 76.31 | 92.11 | 84.50 | 83.80 | 84.97 | 73.11 | 71.50 | 76.12 | 84.35 | 70.31 | | CASE-004: RansomHub, first-day evidence only https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-004-reconstruct-inc-001-temporal-cutoff-d1/3 | 79.57 | 84.79 | 80.70 | 72.37 | 80.42 | 74.00 | 72.66 | 74.51 | 72.55 | 63.70 | | CASE-011: GTG-1002, human framing https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-011-reconstruct-inc-002-framing-framed-as-human/3 | 85.44 | 80.84 | 82.04 | 72.80 | 80.12 | 77.09 | 65.13 | 76.43 | 66.80 | 66.80 | | CASE-012: GTG-1002, AI-agent framing https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-012-reconstruct-inc-002-framing-framed-as-ai-agent-republished/1 | 77.13 | 76.91 | 79.82 | 76.83 | 70.87 | 78.42 | 69.74 | 69.42 | 68.48 | 67.53 | | CASE-013: Google's reported credential-harvesting campaign https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-013-reconstruct-inc-004-full-republished/1 | 89.33 | 88.33 | 88.75 | 82.50 | 87.83 | 77.50 | 52.47 | 79.07 | 78.50 | 68.25 | | Kaggle overall | 79.72 | 83.22 | 81.06 | 77.98 | 80.50 | 74.13 | 68.26 | 75.76 | 72.35 | 69.11 | These are single runs, not stable model rankings. Gemma has the highest displayed overall score 83.22 , followed by GPT-5.6 Luna 81.06 and Grok 4.20 80.50 . The standout case score is Gemma's 92.11 on the short extortion-report task; that is a case-specific result, not proof of general model superiority. Several patterns stand out in these single runs: CASE-013's reference has seven gold events, while the richer INC-001 case has 28. Raw EGRS should not be read as a ranking of real-world incident difficulty or attacker behavior. The Kaggle leaderboard provides the composite score; I have not treated older component-level run outputs as if they were measurements from these restored task versions. Next I would repeat each condition with multiple seeds and inspect whether missed causal links and failed actions recur. There are no repeated-trial confidence intervals, and all model comparisons here are single runs. Building the benchmark involved a few Kaggle-specific failure modes that are worth recording because they affect task visibility and score continuity: Fail or blank cells until those models are run again. Unspecified , had no completed run, and could not be selected in Untitled Task for those attachments. The recovery was to remove the failing attachments and add the earlier evaluated task versions, which restored the seven named rows and their saved scores. kbench.llm . Keep the task model-agnostic, and do not assume a successful run under one fixed model makes it reusable across the leaderboard. Completed is not the same as β€œall models finished.” For the current results, the benchmark is pinned to the previously evaluated versions: CASE-001 through CASE-011 v3 and the republished CASE-012/013 v1 tasks. This keeps the benchmark rows named and retains their model scores while the failed newer uploads remain separate task records. The model is asked for a structured reconstruction rather than a free-form incident summary. An event has a description, an evidence status, and evidence IDs; relationships connect event IDs: @dataclass class ReconstructedEvent: event id: str description: str status: str confirmed | inferred | unknown | attempted | failed evidence ids: list str = field default factory=list @dataclass class ReconstructedRelationship: source event id: str target event id: str relationship: str precedes | enables | causes | depends on @dataclass class Reconstruction: events: list ReconstructedEvent = field default factory=list relationships: list ReconstructedRelationship = field default factory=list unknown steps: list str = field default factory=list The task sends the incident packet with that schema, normalizes the structured response, and passes it to the deterministic scorer: message = SYSTEM PROMPT + "\n\n" + build user prompt result = llm.prompt message, schema=Reconstruction, seed=0, temperature=0 prediction = to prediction result metrics = score prediction prediction, GOLD, VALID EVIDENCE IDS return metrics "egrs" / 100.0 Event matching is one-to-one. Text similarity proposes candidate matches, cited evidence gives a fixed bonus, and a threshold filters weak matches: score = label similarity pe.get "description", "" , gn.get "label", "" if pe ev and gn ev and set pe ev & set gn ev : score += EVIDENCE BONUS if score = MATCH THRESHOLD: candidates.append score, pe "event id" , gn "id" EGRS combines recovery, precision, graph links, evidence attribution, status, uncertainty, and failed-action recognition, while subtracting a hallucination penalty: egrs = 100.0 max 0.0, 0.25 recall + 0.20 precision + 0.15 link f1 + 0.15 evidence attribution + 0.10 status accuracy + 0.10 unknown calibration + 0.05 failed recognition - 0.25 hallucination rate, This makes unsupported certainty costly while giving the model credit for preserving uncertainty and citing the evidence behind its reconstruction. The live Cyber Autopsy Kaggle benchmark and leaderboard https://www.kaggle.com/benchmarks/ujjavalasingh/cyber-autopsy-benchmark groups the seven restored task versions. Scores above were fetched on 2 October 2026. Individual evaluated task versions, including the restored CASE-012 and CASE-013 records, are linked in the results table. The next version will expand the internet-sourced dataset and move further into behavioural analysis. The questions I'm interested in include: How do AI and human cyber activity differ when the evidence is genuinely comparable? What kinds of failures are common in AI cyber activity? How do AI agents adapt after failed actions? How often do guardrails and environmental controls intervene? How much human intervention is involved? Can models recognise when they don't have enough evidence? Does actor framing systematically affect reconstruction? For now, Cyber Autopsy is a small first step toward answering those questions. The benchmark, dataset, scoring code, and evaluation results are all available on Kaggle. If you're interested in seeing how different models reconstruct the same cyber evidence β€” or in digging into where those reconstructions go wrong β€” have a look at the benchmark and results. The interesting part isn't just who gets the highest number. It's what the model thought happened, what the evidence actually says, and where those two things diverge.