{"slug": "when-a-fact-is-retracted-can-an-ai-keep-what-still-holds", "title": "When a Fact Is Retracted, Can an AI Keep What Still Holds?", "summary": "A developer built \"Repair Without Breaking,\" a Kaggle Benchmarking Challenge dataset of 60 JSON episodes across 12 graph motifs and five matched conditions (applicable change, no-op, irrelevant update, retraction, unresolved conflict) to test whether AI assistants retract exactly the conclusions that lose support when evidence changes. In development runs on five parcel-routing episodes with Gemini 2.5 Flash and thinking disabled, four reached the correct final state but only two passed the strict trajectory check, with the model first attempting to rewrite immutable source fields before repairing writable ones. The dataset was produced by a deterministic Python generator from authored graphs, domains, initial values and event rules, with expected states checked against a separately implemented arithmetic reference.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23).*\n\nA plan contains two numbers: 2 and 4. Their total is 6. Then the evidence for the input is withdrawn.\n\nShould an AI erase the total too?\n\nHere are the supplied rules: `left = x`, `right = 6 - x`, and `total = left + right`. The declared domain of `x` is `{2, 3, 4}`. After retraction, neither branch has one justified value. Their total still does: every admissible world gives 6.\n\nThe correct repair clears the unsupported branch values, preserves the total, and marks the affected component as needing a source. This is an authored reference example from motif H03, not an observed model success or failure. It asks a precise question: **when evidence changes, can an assistant retract exactly the conclusions that lose support?**\n\nThat is the focus of **Repair Without Breaking**. Each episode supplies explicit finite-domain rules, a valid stored plan, and one authenticated evidence update. The assistant can inspect the state and event, apply leaf-level patches, and finish with a `completed` or `blocked` status. The evaluator checks the resulting state and recorded tool actions. Stored sources and history are immutable; unrelated live derived leaves remain writable. The event changes the working evidence used for the repair.\n\nThe formulas and finite domains are given to the model. The test concerns their execution and the resulting selective repair, not discovery of hidden dependencies or a new theory of uncertainty.\n\nAn earlier development check exposed a useful distinction. Across five parcel-routing episodes with Gemini 2.5 Flash and thinking disabled, four ended in the correct state, but only two also passed the strict trajectory check. In the applicable-change and retraction episodes, the model first tried to rewrite immutable source fields. The tool rejected those patches atomically, and the model subsequently repaired the writable fields. In the no-op episode, it changed a readiness field that should have stayed unchanged.\n\nThose are development observations from one familiar template, not a held-out success rate. They motivated two separate questions for the main study: **Was the final state correct? What happened on the way there?**\n\nThe v2 set contains 12 graph motifs, each with five matched conditions: an applicable change, a no-op confirmation, an irrelevant update, a retraction, and an unresolved conflict. That gives 60 episodes per model configuration. Within a motif, the starting state and rules stay the same across conditions.\n\nThe motifs cover a serial chain, fan-out, a correlated diamond, a two-root join, cross-plan dependencies, dependency selection, thresholds, saturation, a critical-path join, a filtered aggregate, ranking, and interval intersection. Each episode also contains two independently sourced writable lookalike copies. Accidentally repairing the wrong plan can therefore cause real simulator damage.\n\nThe dataset was built with a deterministic Python generator from authored graphs, source domains, initial values and event rules. It materialized the 12-by-five design as 60 JSON task records with separate expected-state records. My AI assistant, dot, authored the rules, generator and evaluation code under my direction; the Gemini configurations evaluated here were test subjects, not dataset generators. Literal expected rows and a separately implemented arithmetic reference were checked against a generic rule executor. That separation helps catch implementation mistakes, but all three share the authored specification and do not constitute independent human validation.\n\nThe H03 reference walkthrough makes the opening example inspectable: three complete worlds, the same shared `x` in both branches, and a unanimous total of 6. The saved traces below separate these given rules from the observed model actions. The public evidence archive retains the original opaque paths and every recorded patch.\n\nUncertainty is finite and explicit. The environment evaluates complete admissible worlds, sharing the same source value across branches in each world. A derived field retains a value only when all worlds agree. A retracted source does not revive an inactive older record. Unresolved conflicts retain all tied authoritative candidates.\n\nEqual-current-value rewrites are counted as redundancy before wrong-value attempts, including when a required field is still stale. Missing-repair and exact-state checks still catch that stale outcome. Strict scoring therefore should not be read as a complete semantic audit of every attempted value. Rejected historical writes, accepted collateral changes, and later restorations have separate counts. A correct state with no evidence reads can pass the final-state metric on an unchanged control; it does not establish that the model inspected the update.\n\nCPU-only controls checked the scorer before model evaluation. A deterministic rule executor reached 60/60 exact states and 60/60 strict trajectories, repairing 150/150 required leaves. A no-op policy passed only the 24 unchanged controls. A policy that damaged a protected field and later restored it reached 60/60 exact states but 0/60 strict trajectories. These are harness checks, not model scores.\n\nThe completed primary schedule used one common prompt and all 60 cases with each of these configurations, for 120 episodes. Its provider output-token cap was 1,024 per response, with at most six model responses per episode. Each scheduled row was attempted once, with no automatic retry or fallback.\n\n| Primary configuration | Backend | Thinking setting | Output cap per response | Completed episodes | \n|---|---|---|---|---|\n| `google/gemini-2.5-flash` | Native GenAI through Kaggle | `reasoning=\"none\"` | 1,024 | 60 | \n| `google/gemini-3.5-flash-lite` | Native GenAI through Kaggle | Explicit `MINIMAL` | 1,024 | 60 | \n\nThis is a comparison of whole model configurations. Different model families and thinking settings do not isolate the effect of reasoning, and equal request limits do not establish equal cost. There is no primary prompt A/B experiment.\n\nThe execution used `kaggle-benchmarks 0.6.1` and `google-genai 2.24.0`. The primary protocol was frozen on October 1, 2026 at 07:43:54 UTC, before v2 outputs. Its prospective order interleaved the two configurations for each case. The common limits included six responses and 30 tool calls per episode; byte limits were fixed per case in the plan. Temperature and other unspecified sampling parameters used provider defaults; this is not a temperature-zero claim. One attempt per row cannot estimate repeat-to-repeat variability. The frozen protocol and per-row prompt/tool hashes are in the [reproducibility archive](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence). The canonical scientific-plan SHA-256 is `3401fac778d2a2f1585048e93e8340e010afe7cda8bfca7381bebb2395032ecb`.\n\nAfter the first 50 rows, covering H01–H05, the resource-admission controller was amended prospectively before H06. It began using fresh verified free-quota balances with a buffer of twice the upcoming whole-motif reservation. The original conservative accounting would otherwise have paused execution. This changed resource admission, not the cases, order, prompt, model settings, response/tool limits, scoring or zero-retry rule. All earlier artifacts were retained, and the amendment applied to every remaining scheduled row. The buffer does not prove settled billing or guarantee uninterrupted quota. The entire execution policy therefore cannot be described as unchanged from the original freeze. This admission change is disclosed here separately from the scientific input hashes in the [reproducibility archive](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence).\n\nAfter 70 primary rows had completed, unfinished public responses near the configured output cap motivated a separate sensitivity check. The completed follow-up covers all 60 cases with both model identities, rather than selecting failed cases. It raises the per-response output cap from 1,024 to 4,096 while retaining the same named thinking settings, prompts, tools, scoring and six-response limit. The primary 120-row schedule has since completed unchanged, with all original outcomes retained.\n\nThis follow-up was selected after observing primary outputs. Its scientific protocol was frozen on October 1, 2026 at 09:05:04 UTC, before its first model output, with canonical plan SHA-256 `13fd248f0d62af43f313366056fb4dddfe7d4cbe60db50fa7a905661224b5d7b`. Its [frozen protocol and complete evidence](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence) retain separate attempt IDs. It is not part of the original untouched primary design. The public-task execution also groups all five conditions by model within each motif, whereas the primary schedule interleaved models by case. Each model's case order is preserved, but cross-model dispatch order differs. The output cap is the only inference-parameter change; it is not the only execution difference. One run at each budget cannot separate the budget change from scheduling, time or ordinary model variability, or establish a general causal effect. For the minimal-thinking configuration, the same setting also does not prove identical hidden compute usage.\n\nAll 120 scheduled primary episodes produced valid bounded outcomes: 60 per configuration. There were no missing or infrastructure-incomplete rows, and no six-response-limit exhaustion. The frozen scientific analyzer and resource-admission validator both passed.\n\n| Primary metric | 2.5 Flash, none | 3.5 Flash-Lite, MINIMAL | \n|---|---|---|\n| Exact final states | 31/60 | 25/60 | \n| Strict trajectories | 16/60 | 25/60 | \n| Required leaves repaired | 48/150 | 5/150 | \n| Writable protected leaves retained | 1188/1200 | 1200/1200 | \n| State and event both read | 60/60 | 55/60 | \n| Correct explicit finish | 22/60 | 36/60 | \n| Unchanged controls: exact final | 22/24 | 24/24 | \n| Repair-needed episodes: exact final | 9/36 | 1/36 | \n| Episodes without an explicit finish | 36/60 | 0/60 | \n| False completion | 3/60 | 35/60 | \n\n**Preserving everything can hide repairing almost nothing.** Flash-Lite retained all 1,200 writable protected leaf observations, but repaired only 5/150 required leaves. Its 25 exact successes comprised all 24 unchanged controls and one repair-needed episode. In that successful H07 applicable-change episode, it patched all four required leaves, harmlessly rewrote readiness to the same value, and finished correctly. The preservation result must be read beside repair coverage.\n\n**The ranking changes with the metric.** Flash produced more exact final states, 31/60 versus 25/60, while Flash-Lite had more strict passes, 25/60 versus 16/60. Flash's 15 final-only successes were unchanged controls: 14 lacked an explicit finish, and H04 no-op recovered from malformed tool arguments. This is a completion-and-trajectory distinction in these configurations, not evidence that the strict leader is better at selective repair.\n\n**Strict success does not require evidence access in this frozen scorer.** Four Flash-Lite controls passed strict despite omitting the event read. That is why the separate evidence-access column is necessary. Reading evidence is also not proof of applying it correctly.\n\nEach condition contains 12 authored cases per configuration. “Exact; strict” below gives two episode-success counts with the same denominator; repair leaves use their own denominator.\n\n| Condition | Flash exact; strict | Lite exact; strict | Flash repair leaves | Lite repair leaves | \n|---|---|---|---|---|\n| Applicable change | 3/12; 3/12 | 1/12; 1/12 | 11/40 | 4/40 | \n| No-op | 11/12; 5/12 | 12/12; 12/12 | No repair needed | No repair needed | \n| Irrelevant update | 11/12; 2/12 | 12/12; 12/12 | No repair needed | No repair needed | \n| Retraction | 5/12; 5/12 | 0/12; 0/12 | 28/56 | 0/56 | \n| Unresolved conflict | 1/12; 1/12 | 0/12; 0/12 | 9/54 | 1/54 | \n\nAll five conditions are present for every motif and configuration:\n\n| Motif | Flash exact | Flash strict | Lite exact | Lite strict | \n|---|---|---|---|---|\n| H01 | 4/5 | 3/5 | 2/5 | 2/5 | \n| H02 | 3/5 | 2/5 | 2/5 | 2/5 | \n| H03 | 2/5 | 0/5 | 2/5 | 2/5 | \n| H04 | 3/5 | 1/5 | 2/5 | 2/5 | \n| H05 | 3/5 | 2/5 | 2/5 | 2/5 | \n| H06 | 2/5 | 0/5 | 2/5 | 2/5 | \n| H07 | 3/5 | 1/5 | 3/5 | 3/5 | \n| H08 | 3/5 | 2/5 | 2/5 | 2/5 | \n| H09 | 3/5 | 2/5 | 2/5 | 2/5 | \n| H10 | 2/5 | 1/5 | 2/5 | 2/5 | \n| H11 | 2/5 | 2/5 | 2/5 | 2/5 | \n| H12 | 1/5 | 0/5 | 2/5 | 2/5 | \n\nAll 60 case pairs and all 12 five-condition motif groups are complete. Exact-state paired outcomes were: both pass 22/60, Flash only 9/60, Flash-Lite only 3/60, neither 26/60. Strict paired outcomes were: both 7/60, Flash only 9/60, Flash-Lite only 18/60, neither 26/60. These are descriptive comparisons. The five conditions share a graph and starting state; the 12 purposively authored motifs are not a random population sample.\n\nThe following error and completion breakdown comes from a post-hoc trace audit, rather than a preregistered error taxonomy. Every recorded tool result, final state and score was replay-checked against the frozen simulator. Flash left 36/60 episodes without an explicit finish. In 31 of those 36, the final public response was visibly incomplete and reported 1,024–1,027 candidate output tokens, consistent with the configured per-response cap. Provider finish reasons were unavailable, so this does not confirm why generation stopped. The other five did not all fit that pattern: two gave a complete repair plan without applying it; one correctly described a no-write control without calling finish; two ended mid-sentence below that range. There was no recorded six-response-limit exhaustion.\n\nEach configuration had one malformed model-supplied JSON argument recorded as `adapter_protocol_error`: Flash H04 no-op and Flash-Lite H11 conflict. These are model-visible protocol errors, not provider or SDK infrastructure failures. Flash recovered the exact state in its control but failed strict; Flash-Lite subsequently repaired only one of six required leaves and falsely finished `completed`.\n\nFlash made 12 accepted collateral leaf mutations across five episodes, with none subsequently restored. Flash-Lite made none. The trace also recorded 63 and 21 state-unchanging equal-value rewrites respectively. These include some rewrites of stale required values, which do not count as repairs. Neither configuration attempted a historical write in this primary matrix. Those observations do not erase the rejected historical writes in the separate development runs.\n\nMeasured per-episode costs were unavailable for both configurations. Candidate output counters do not establish total generated tokens; thinking-token counts were unreported for all 60 Flash-Lite episodes. Reservations and quota displays are not substituted for billing measurements.\n\nThe opening puzzle now has a concrete observed counterpart. In the primary 1,024-token H03 retraction episode, Gemini 2.5 Flash with thinking disabled read the state and event, then submitted a five-operation patch. It correctly cleared `left`, `right` and `balance`, and changed readiness to `needs_source`. Those were all four required repairs. Its fifth operation also changed `total` from 6 to `null`.\n\nThe patch was accepted. The model then correctly finished `blocked`, but the final state was wrong: the sum should remain 6 in every admissible world. Required repair coverage was 4/4, writable protected preservation was 13/14, and both exact-state and strict-trajectory checks failed. This episode had no tool errors, infrastructure failure or round-limit exhaustion.\n\nThe public explanation makes the local error inspectable. For the total, it says: “Since both references are now `null`, this path becomes `null`.” That matches the accepted patch. The declared contract instead evaluates the complete worlds first and keeps their unanimous total. This is evidence of the explanation and action in this saved run, not a claim about the model's hidden reasoning or intrinsic capability.\n\n*Gemini 2.5 Flash, native GenAI, thinking disabled. The given rules preserve total=6. Both saved retraction attempts erase it; the higher-cap change case and both controls pass exact and strict. All five conditions are retained. One authored motif, one attempt per phase, selected after output inspection. The [public evidence](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence) includes all original tool traces; primary retraction row: `2dcf2cbb027767ef945d`.*\n\nThe matched Flash-Lite MINIMAL run shows why preservation also needs repair context. It left the total at 6 by leaving the entire state unchanged, made none of the four required repairs, and reported `completed`. Its 14/14 writable-preservation score therefore does not establish successful selective retraction. Both configurations failed this episode for different recorded reasons.\n\nThere is also genuine positive evidence. In H09 retraction, Flash correctly cleared three uncertain outputs, changed readiness, and finished `blocked`: 4/4 required repairs, with 17/17 writable protected leaves retained. Independent B-finish=5 and C-finish=4 stayed intact. This shows successful preservation of independent branches in that case, even though the correlated H03 invariant was lost. Flash succeeded on nine repair-needed episodes overall; the selected failure is not a claim that it never repairs or retracts correctly.\n\nThe saved 4,096-token Flash retraction attempt repeats that extra deletion: 4/4 required repairs, 13/14 writable preservation, a correct `blocked` finish, and failed exact and strict checks. The conflict comparison changes from an unfinished primary attempt with no patch to a follow-up that makes all four required repairs and also deletes the invariant. All ten Flash H03 episodes across the two phases replay exactly against the frozen environment.\n\nThe successes matter too. In the follow-up, the applicable-change case and both unchanged controls pass exact and strict, giving 3/5 for both metrics on this motif, compared with primary 2/5 exact and 0/5 strict. The archived evidence retains all five conditions. These saved attempts show both recovery and a repeated error; they do not establish a cause, an inevitable failure or a full-study result.\n\nThe [complete evidence archive](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence) retains the H09 retraction success (`4100c5fb4d9644a54bff`), H07 repair success, malformed-argument examples and unfinished responses alongside every other episode.\n\nAll 24 follow-up task/model jobs completed, yielding 120/120 valid bounded outcomes, 60 per configuration. No rows were missing or infrastructure-incomplete, and no episode exhausted the six-response limit. The complete report was recomputed from the privacy-filtered bytes distributed in the evidence archive. These results remain separate from the primary phase.\n\n| Follow-up metric | 2.5 Flash, none | 3.5 Flash-Lite, MINIMAL | \n|---|---|---|\n| Exact final states | 47/60 | 25/60 | \n| Strict trajectories | 43/60 | 25/60 | \n| Required leaves repaired | 129/150 | 5/150 | \n| Writable protected leaves retained | 1185/1200 | 1200/1200 | \n| State and event both read | 60/60 | 55/60 | \n| Correct explicit finish | 51/60 | 36/60 | \n| Unchanged controls: exact final | 22/24 | 24/24 | \n| Repair-needed episodes: exact final | 25/36 | 1/36 | \n| Episodes without an explicit finish | 5/60 | 0/60 | \n| False completion | 4/60 | 35/60 | \n\nFlash's exact-state successes rose from 31/60 to 47/60, and strict successes from 16/60 to 43/60. Required repairs rose from 48/150 to 129/150. Flash-Lite's exact, strict and repair counts stayed at 25/60, 25/60 and 5/150. This observed sensitivity matters: the lower-budget primary result alone would substantially understate Flash's successful repairs in the higher-budget attempts.\n\nThe preservation trade-off remains visible. Flash retained 1185/1200 writable protected leaves in the follow-up, compared with 1188/1200 in the primary. It made 15 accepted collateral mutations, with no restoration, despite repairing many more required leaves. Flash-Lite retained 1200/1200 in both phases while failing 35/36 repair-needed episodes in each. More repairs and less damage are separate goals.\n\n| Condition | Flash exact; strict | Lite exact; strict | Flash repair leaves | Lite repair leaves | \n|---|---|---|---|---|\n| Applicable change | 9/12; 9/12 | 1/12; 1/12 | 36/40 | 4/40 | \n| No-op | 11/12; 9/12 | 12/12; 12/12 | No repair needed | No repair needed | \n| Irrelevant update | 11/12; 9/12 | 12/12; 12/12 | No repair needed | No repair needed | \n| Retraction | 10/12; 10/12 | 0/12; 0/12 | 53/56 | 0/56 | \n| Unresolved conflict | 6/12; 6/12 | 0/12; 0/12 | 40/54 | 1/54 | \n\n| Motif | Flash exact | Flash strict | Lite exact | Lite strict | \n|---|---|---|---|---|\n| H01 | 5/5 | 5/5 | 2/5 | 2/5 | \n| H02 | 5/5 | 5/5 | 2/5 | 2/5 | \n| H03 | 3/5 | 3/5 | 2/5 | 2/5 | \n| H04 | 5/5 | 4/5 | 2/5 | 2/5 | \n| H05 | 4/5 | 3/5 | 2/5 | 2/5 | \n| H06 | 3/5 | 1/5 | 2/5 | 2/5 | \n| H07 | 4/5 | 4/5 | 3/5 | 3/5 | \n| H08 | 3/5 | 3/5 | 2/5 | 2/5 | \n| H09 | 4/5 | 4/5 | 2/5 | 2/5 | \n| H10 | 4/5 | 4/5 | 2/5 | 2/5 | \n| H11 | 4/5 | 4/5 | 2/5 | 2/5 | \n| H12 | 3/5 | 3/5 | 2/5 | 2/5 | \n\nEach row below compares the same 60 cases across the two phases. All pairs and all 12 motif groups are complete.\n\n| Configuration | Metric | Both pass | Primary only | Follow-up only | Neither | Valid pairs | \n|---|---|---|---|---|---|---|\n| Flash | exact | 31 | 0 | 16 | 13 | 60/60 | \n| Flash | strict | 16 | 0 | 27 | 17 | 60/60 | \n| Flash-Lite | exact | 25 | 0 | 0 | 35 | 60/60 | \n| Flash-Lite | strict | 25 | 0 | 0 | 35 | 60/60 | \n\nFlash had 16 exact-state recoveries and 27 strict recoveries, with no pass-to-fail transitions on either metric. Exact counts increased in 11 of the 12 motifs and were unchanged in H08; strict counts increased in all 12. Flash-Lite had no exact or strict transitions in any motif. The archive retains every paired case and the per-condition and per-motif differences. Zero binary regressions does not mean every aspect improved: the protected-leaf losses above also matter.\n\nFive Flash follow-up episodes still lacked an explicit finish, compared with 36 in the primary; Flash-Lite explicitly finished all 60 in both phases. The follow-up also records one model-argument runner error per configuration and no environment tool errors. Measured costs remain unavailable, and Flash-Lite's thinking-token counts remain unreported. These completion and accounting observations do not establish why any generation stopped.\n\nThe higher cap was selected after inspecting primary outputs, and cross-model dispatch order also changed. The paired counts describe these saved attempts. They do not isolate the output cap as the cause, establish a stable model ranking, or estimate a population effect.\n\nThis diagnostic measures execution of explicit finite-domain rules and selective repair in a bounded synthetic tool environment. The controls help distinguish failure to propagate a change from failure to preserve valid state. They also expose the gap between final correctness and the safety of the path taken.\n\nThe scope is limited. Each episode has only one uncertain source. All retraction and conflict cases require a blocked finish. The serial-chain motif resembles development cases. The references have separate implementation and agent review, but no independent human annotation. Source snapshots and history are protected by the tool interface. Publishing the cases makes them reproducible, not permanently hidden from future models. None of this establishes deployment reliability, general out-of-distribution performance, or a new research problem.\n\nUseful extensions would include independent human review, multiple uncertain sources, uncertainty that still permits full completion, repeated runs, and a broader model lineup. These are proposed next measurements, not results of this entry.\n\n[Open the Kaggle Benchmark: Repair Without Breaking, 4,096-Token Follow-Up](https://www.kaggle.com/benchmarks/sean2333/repair-without-breaking-4096-token-follow-up)\n\nThe Benchmark contains all 12 motif tasks and the two supported model configurations. Each task score is its five-case exact-state success fraction; the Benchmark uses the average of task scores. With equal five-case tasks, the full offline averages are 47/60 (78.33%) for Flash and 25/60 (41.67%) for Flash-Lite. These displayed scores belong only to the 4,096-token follow-up. Strict trajectory scores and the complete original 1,024-token primary study are reported separately above and in the archive.\n\nA platform-default unsupported 3.7 model registration can appear as an administrative configuration-error row with no score. It was rejected before case exposure or model generation and is outside the two scientific cohorts and their denominators.\n\nThe [public reproducibility archive](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence) contains the source and tests, both frozen protocols, 240 privacy-filtered episode records, all 24 follow-up aggregates, and the separately recomputed phase and paired results. The original uploaded ZIP has SHA-256 `8c13927f880223cc558e121820e53e4dd84dc0049e8269bdaa5fb51d8d2c7ac5`. Kaggle unpacks that ZIP into dataset files and may repackage downloads, so a downloaded ZIP need not have the same container checksum. Verify the extracted file sizes and SHA-256 values against `PUBLIC_ARCHIVE_MANIFEST.json`. Synthetic tool trajectories and final states are retained so the counts can be audited; private quotas, credentials, account metadata and environment logs are excluded.\n\nSelective updating is an established evaluation problem. [Belief-R, published at EMNLP 2024](https://aclanthology.org/2024.emnlp-main.586/), studies revising conclusions versus maintaining them when updates are unnecessary. [BeliefShift, a March 2026 preprint](https://arxiv.org/abs/2603.23848), evaluates consistency, contradictions and evidence-driven revision across conversational sessions.\n\nThe [BeliefDrift author card](https://huggingface.co/datasets/fahadhafeezofficial/beliefdrift) describes especially close public overlap: selective updates, unaffected-belief preservation, weighted dependency propagation and evidence-removal counterfactuals. I cite that stated scope; its chronology, peer-review status and implementation validity have not been independently verified here. [MnemeBrain's benchmark suite](https://github.com/mnemebrain/mnemebrain-benchmark) includes evidence-linked belief maintenance and retraction scenarios. Its BMB scoring skips unsupported adapter capabilities, so those percentages cannot be compared directly with this fixed matrix.\n\n[Only What I Asked](https://dev.to/akj1608/only-what-i-asked-1i26) studies narrowly requested text edits and collateral changes. [Reservation Replay](https://dev.to/yongchan_kwonnuckdrip_/when-a-failed-request-must-stay-failed-reservation-replay-2o6o) checks cached decisions and the complete reservation trace. [ChartReplay](https://dev.to/p0rt/same-patient-conflicting-documents-can-ai-preserve-the-evidence-3anp) tests preservation of source assertions and their relationships as corrections and conflicts arrive. These are close neighbors: restraint, intermediate correctness and evidence preservation are already established concerns.\n\nThis entry contributes a compact, inspectable finite-domain repair diagnostic and traces such as H03's lost invariant. The rules are supplied, the admissible worlds can be enumerated, and the state changes can be replayed. Executable tools, trajectory checks and selective updating are not claimed as unique. The observed invariant loss illustrates this run under this contract; it does not introduce a new uncertainty theory or establish superiority over these benchmarks.\n\nAll scenarios are synthetic. The [archive](https://www.kaggle.com/datasets/sean2333/repair-without-breaking-reproducibility-evidence) binds the reference states, deterministic checks, model settings and scoring logic with a file-level SHA-256 manifest. The frozen candidate hash is `44775804c2286246e1716e030aa10c5e6addd15901fbc84ac777386462e79482`; the corpus manifest hash is `fd606407c488f9f7ba8c50512e49720ef08764e547ec10585fea26b5932e2d23`. The resource-admission amendment and follow-up scheduling difference are disclosed above. The paired reporting helper was written after initial follow-up outputs; it is a post-hoc analysis utility, not an untouched original confirmatory analysis.\n\n[ToolTrap](https://dev.to/himanshu_748/tooltrap-tool-results-are-data-wasnt-enough-25oh) is related motivation for evaluating concrete tool behavior. This project does not claim to invent selective repair or to outperform an existing benchmark. No ToolTrap cases were reused in this authored corpus.\n\nThis diagnostic and write-up were authored by my AI assistant, dot, under my direction, with implementation review and trace analysis by AI agents. I selected the goal and authorized publication; the work does not claim an independent human review. The tested Gemini configurations did not generate the dataset.", "url": "https://wpnews.pro/news/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds", "canonical_source": "https://dev.to/sean233_0984ede20c450437f/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds-3jh5", "published_at": "2026-10-01 15:04:17+00:00", "updated_at": "2026-10-01 15:14:18.517464+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["Gemini 2.5 Flash", "Kaggle", "Repair Without Breaking", "dot", "Google"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds", "markdown": "https://wpnews.pro/news/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds.md", "text": "https://wpnews.pro/news/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds.txt", "jsonld": "https://wpnews.pro/news/when-a-fact-is-retracted-can-an-ai-keep-what-still-holds.jsonld"}}