# I Asked 10 AI Models to Reconstruct Real Cyber Attacks

> Source: <https://dev.to/ujja/i-asked-10-ai-models-to-reconstruct-real-cyber-attacks-2o3j>
> Published: 2026-10-01 15:33:59+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Cyber incident reports rarely arrive as a clean, complete timeline. They describe what investigators or security teams could observe, mix those observations with analysis, and leave some questions unanswered. That is exactly where an AI-generated summary can become risky: it may sound convincing while quietly turning an inference into a fact or filling a gap with an event that the evidence never established.

I built **Cyber Autopsy** to test a narrower, practical question: when given pieces of a reported cyber incident, can a model reconstruct what happened while showing which evidence supports each claim and admitting what remains uncertain? The model must produce a timeline, connect events with causal or temporal links, cite evidence, and distinguish confirmed activity from inference, failed attempts, contradictions, and unknowns. A plausible attack story is not enough; unsupported certainty should count against it.

I chose this problem because incident reconstruction depends on more than recognizing familiar attack techniques. An analyst needs to know whether a step was observed or inferred, whether an attempt actually succeeded, and how one event led to another. Those distinctions can disappear in a fluent summary. Measuring them separately makes it easier to see whether a model is recovering the evidence or merely telling a likely-sounding story.

The first evaluation uses seven tasks built from four public reports. They include a detailed human-operated ransomware intrusion and vendor-reported campaigns involving AI-assisted activity. These are useful real-world case studies, but they are not a controlled contest between human and AI attackers: the reports differ in detail, evidence source, and corroboration. I also included two versions of one case with identical evidence but different actor framing, to see whether that wording changes the model's reconstruction.

This is a pilot, not a claim that AI attackers are more or less capable than people. It evaluates model reconstructions of reported incidents, not live attack behavior. The scorer is deterministic and reports an Evidence-Grounded Reconstruction Score (EGRS), alongside event recall and precision, causal-link quality, evidence attribution, status accuracy, uncertainty calibration, and hallucination-related measures. The temporal cutoff and framing pair are exploratory comparisons; only the framing pair holds the evidence fixed.

The seven tasks are built from four public incident reports, not invented scenarios:

These are real reported incidents, but the evidence quality is not uniform: the RansomHub case is reconstructed from host and network telemetry described by The DFIR Report, while the AI-actor case studies rely on security-vendor reporting. The benchmark labels that distinction rather than treating the cases as equally observed or directly comparable.

The short IDs are just labels: `INC` means the source incident, and `CASE` means the particular benchmark task. Each task gives the model an evidence packet and asks for the best-supported reconstruction, not a free-form guess.

I ran ten models from several providers against the same seven Kaggle tasks: **Gemini 3.7 Flash**, **Gemma 4 26B A4B**, **GLM-5**, **Grok 4.20 Reasoning**, **GPT-5.6 Luna**, **GPT-5.6 Sol**, **GPT-5.4 mini**, **Claude Sonnet 5**, **Claude Opus 5**, and **Qwen 3 Coder 480B**. The current benchmark view has a score for every model-task pair. Kaggle's overall score aggregates these seven tasks, which include related variants of the same incidents.

The table is the Kaggle leaderboard snapshot fetched on **2 October 2026**, after duplicate and failing task attachments were removed and the earlier evaluated versions restored. CASE-001 through CASE-011 use v3; CASE-012 and CASE-013 use their republished v1 versions. Values are EGRS percentages (Kaggle's 0-1 scores multiplied by 100). EGRS rewards recovering supported events and links, citing evidence, and representing uncertainty, while penalizing unsupported events. It is specific to this evidence-reconstruction task, not a general measure of intelligence or cybersecurity ability. Kaggle's overall score now matches the equal-weight mean across the seven task rows. Since some tasks are related variants of the same incidents, this is descriptive rather than an independent-sample leaderboard.

| Kaggle task | Gemini Flash | Gemma 4 | GPT-5.6 Luna | GLM-5 | Grok 4.20 | Claude Sonnet 5 | Claude Opus 5 | GPT-5.6 Sol | GPT-5.4 mini | Qwen 3 Coder | 
|---|---|---|---|---|---|---|---|---|---|---|
| [CASE-001: Full RansomHub intrusion](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-001-reconstruct-inc-001-full/3) | 70.55 | 82.38 | 71.27 | 81.31 | 75.71 | 73.75 | 78.14 | 78.06 | 66.06 | 70.90 | 
| [CASE-002: GTG-1002 espionage campaign](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-002-reconstruct-inc-002-full/3) | 79.72 | 77.21 | 80.31 | 76.22 | 83.60 | 65.05 | 68.16 | 76.71 | 69.68 | 76.30 | 
| [CASE-003: Reported data-extortion operation](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-003-reconstruct-inc-003-full/3) | 76.31 | 92.11 | 84.50 | 83.80 | 84.97 | 73.11 | 71.50 | 76.12 | 84.35 | 70.31 | 
| [CASE-004: RansomHub, first-day evidence only](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-004-reconstruct-inc-001-temporal-cutoff-d1/3) | 79.57 | 84.79 | 80.70 | 72.37 | 80.42 | 74.00 | 72.66 | 74.51 | 72.55 | 63.70 | 
| [CASE-011: GTG-1002, human framing](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-011-reconstruct-inc-002-framing-framed-as-human/3) | 85.44 | 80.84 | 82.04 | 72.80 | 80.12 | 77.09 | 65.13 | 76.43 | 66.80 | 66.80 | 
| [CASE-012: GTG-1002, AI-agent framing](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-012-reconstruct-inc-002-framing-framed-as-ai-agent-republished/1) | 77.13 | 76.91 | 79.82 | 76.83 | 70.87 | 78.42 | 69.74 | 69.42 | 68.48 | 67.53 | 
| [CASE-013: Google's reported credential-harvesting campaign](https://www.kaggle.com/benchmarks/tasks/ujjavalasingh/cyber-autopsy-case-013-reconstruct-inc-004-full-republished/1) | 89.33 | 88.33 | 88.75 | 82.50 | 87.83 | 77.50 | 52.47 | 79.07 | 78.50 | 68.25 | 
| **Kaggle overall** | **79.72** | **83.22** | **81.06** | **77.98** | **80.50** | **74.13** | **68.26** | **75.76** | **72.35** | **69.11** | 

These are single runs, not stable model rankings. Gemma has the highest displayed overall score (83.22), followed by GPT-5.6 Luna (81.06) and Grok 4.20 (80.50). The standout case score is Gemma's 92.11 on the short extortion-report task; that is a case-specific result, not proof of general model superiority.

Several patterns stand out in these single runs:

CASE-013's reference has seven gold events, while the richer INC-001 case has 28. Raw EGRS should not be read as a ranking of real-world incident difficulty or attacker behavior. The Kaggle leaderboard provides the composite score; I have not treated older component-level run outputs as if they were measurements from these restored task versions. Next I would repeat each condition with multiple seeds and inspect whether missed causal links and failed actions recur. There are no repeated-trial confidence intervals, and all model comparisons here are single runs.

Building the benchmark involved a few Kaggle-specific failure modes that are worth recording because they affect task visibility and score continuity:

`Fail` or blank cells until those models are run again.`Unspecified`, had no completed run, and could not be selected in `Untitled Task` for those attachments. The recovery was to remove the failing attachments and add the earlier evaluated task versions, which restored the seven named rows and their saved scores.`kbench.llm`. Keep the task model-agnostic, and do not assume a successful run under one fixed model makes it reusable across the leaderboard.`Completed` is not the same as “all models finished.”
For the current results, the benchmark is pinned to the previously evaluated versions: CASE-001 through CASE-011 v3 and the republished CASE-012/013 v1 tasks. This keeps the benchmark rows named and retains their model scores while the failed newer uploads remain separate task records.

The model is asked for a structured reconstruction rather than a free-form incident summary. An event has a description, an evidence status, and evidence IDs; relationships connect event IDs:

```
@dataclass
class ReconstructedEvent:
    event_id: str
    description: str
    status: str  # confirmed | inferred | unknown | attempted | failed
    evidence_ids: list[str] = field(default_factory=list)

@dataclass
class ReconstructedRelationship:
    source_event_id: str
    target_event_id: str
    relationship: str  # precedes | enables | causes | depends_on

@dataclass
class Reconstruction:
    events: list[ReconstructedEvent] = field(default_factory=list)
    relationships: list[ReconstructedRelationship] = field(default_factory=list)
    unknown_steps: list[str] = field(default_factory=list)
```

The task sends the incident packet with that schema, normalizes the structured response, and passes it to the deterministic scorer:

```
message = SYSTEM_PROMPT + "\n\n" + build_user_prompt()
result = llm.prompt(message, schema=Reconstruction, seed=0, temperature=0)

prediction = to_prediction(result)
metrics = score_prediction(prediction, GOLD, VALID_EVIDENCE_IDS)
return metrics["egrs"] / 100.0
```

Event matching is one-to-one. Text similarity proposes candidate matches, cited evidence gives a fixed bonus, and a threshold filters weak matches:

```
score = label_similarity(pe.get("description", ""), gn.get("label", ""))
if pe_ev and gn_ev and (set(pe_ev) & set(gn_ev)):
    score += EVIDENCE_BONUS
if score >= MATCH_THRESHOLD:
    candidates.append((score, pe["event_id"], gn["id"]))
```

EGRS combines recovery, precision, graph links, evidence attribution, status, uncertainty, and failed-action recognition, while subtracting a hallucination penalty:

```
egrs = 100.0 * max(
    0.0,
    0.25 * recall + 0.20 * precision + 0.15 * link_f1
    + 0.15 * evidence_attribution + 0.10 * status_accuracy
    + 0.10 * unknown_calibration + 0.05 * failed_recognition
    - 0.25 * hallucination_rate,
)
```

This makes unsupported certainty costly while giving the model credit for preserving uncertainty and citing the evidence behind its reconstruction.

The live [Cyber Autopsy Kaggle benchmark and leaderboard](https://www.kaggle.com/benchmarks/ujjavalasingh/cyber-autopsy-benchmark) groups the seven restored task versions. Scores above were fetched on 2 October 2026. Individual evaluated task versions, including the restored CASE-012 and CASE-013 records, are linked in the results table.

The next version will expand the internet-sourced dataset and move further into behavioural analysis.

The questions I'm interested in include:

How do AI and human cyber activity differ when the evidence is genuinely comparable?

What kinds of failures are common in AI cyber activity?

How do AI agents adapt after failed actions?

How often do guardrails and environmental controls intervene?

How much human intervention is involved?

Can models recognise when they don't have enough evidence?

Does actor framing systematically affect reconstruction?

For now, Cyber Autopsy is a small first step toward answering those questions.

The benchmark, dataset, scoring code, and evaluation results are all available on Kaggle.

If you're interested in seeing how different models reconstruct the same cyber evidence — or in digging into where those reconstructions go wrong — have a look at the benchmark and results.

The interesting part isn't just who gets the highest number.

It's what the model thought happened, what the evidence actually says, and where those two things diverge.
