{"slug": "ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought", "title": "AI Agent Decision Logs: Record Evidence Without Storing Chain of Thought", "summary": "A developer has published a guide for building AI agent decision logs: compact, structured records that capture the task, bounded evidence, policy and release versions, chosen action, and independently verified outcome for each consequential agent action. The approach deliberately avoids storing chain-of-thought or full prompts, instead using stable references to access-controlled data and versioned policy identifiers so incidents can be reconstructed without duplicating sensitive records.", "body_md": "An agent can update a customer record, trigger a refund, or route a support case while every dashboard stays green. Then someone asks the only question that matters: **why was this action allowed, and what actually changed?** A trace may show tokens and latency. An application log may show an HTTP 200. Neither is a dependable answer.\n\nThe missing piece is an **AI agent decision log**: a small, structured record for each meaningful choice. It links the task, the bounded evidence, the policy and release in force, the chosen action, and the independently verified outcome. It is not a transcript, and it should not be a hidden-reasoning archive.\n\nThis guide shows how to build that record so an engineer can debug an incident, an operator can review a risky action, and a customer-data team does not inherit a second copy of every prompt.\n\nKeep both. They answer different questions.\n\n| Artifact | Best for | Usually misses | \n|---|---|---|\n| Trace | latency, retries, spans, model calls | business intent and policy verdict | \n| Tool log | request/response diagnostics | whether the response changed the target state | \n| Prompt release manifest | which behavior bundle was deployed | why this request took this branch | \n| Decision log | why the system was allowed to act and what was verified | low-level performance detail | \n\nSuppose a support agent changes a renewal status. A trace can show `update_subscription` completed in 180 ms. A useful decision log can say that the request came from an authenticated account owner, the account was eligible under `renewal-policy-v4`, the action stayed below the risk threshold, the expected version was 27, and the billing API confirmed version 28 with the requested status.\n\nThat is enough to investigate and reproduce the decision without pretending that a model's private intermediate text is a reliable audit artifact.\n\nLogging every token is expensive, risky, and noisy. It also makes review harder. Define a **decision boundary**: log when the agent crosses from interpretation into a consequential branch or side effect.\n\nGood boundaries include:\n\nDo not make a decision log for each retrieval chunk, retry, or sentence of a draft. Put those in traces and link their stable IDs when they matter. A compact log should be readable when an incident is active.\n\nBefore adding fields, ask: *Could a teammate who was absent reconstruct the reason, authority, and result of this action from this one record?*\n\nIf the answer is no, add the missing reference. If the answer needs the full prompt, full customer record, and every generated token, redesign the workflow so the record can cite safe, scoped evidence instead.\n\nThe schema needs stable identities, not a giant JSON blob. Here is a TypeScript shape for a material action:\n\n```\ntype DecisionLog = {\n  id: string;\n  occurredAt: string;\n  tenantId: string;\n  runId: string;\n  decision: \"route\" | \"allow\" | \"hold\" | \"block\" | \"complete\";\n  task: { type: string; requestRef: string };\n  release: { manifestId: string; model: string; policyVersion: string };\n  evidence: Array<{ kind: string; ref: string; sha256?: string }>;\n  action?: { tool: string; targetRef: string; idempotencyKey: string };\n  policy: { verdict: \"allow\" | \"hold\" | \"block\"; reasons: string[] };\n  outcome: { state: \"verified\" | \"rejected\" | \"unknown\"; resultRef?: string };\n  redactionVersion: string;\n};\n```\n\nThree choices matter here.\n\nFirst, `requestRef`, `targetRef`, and `resultRef` point to access-controlled data; they do not copy it. A support transcript or invoice belongs in its existing system with its existing retention rules.\n\nSecond, store versioned identifiers for the release, model, and policy. An answer may be acceptable under one policy version and blocked under the next. Without IDs, you cannot explain a historical decision honestly.\n\nThird, store the *verdict and its public reasons*, such as `account_owner_verified` or `refund_below_limit`. Do not store a fabricated summary of hidden reasoning. Design the policy engine to emit explicit reason codes that are safe to expose to operators.\n\nThe safest shape is a two-phase record. Write a planned decision before an external side effect; append the verified result after it.\n\n``` js\nasync function changeRenewal(input: RenewalInput) {\n  const decision = await decide(input); // typed facts + deterministic policy\n  const record = await decisions.insert({\n    ...decision,\n    outcome: { state: \"unknown\" },\n  });\n\n  if (decision.policy.verdict !== \"allow\") return record;\n\n  const result = await billing.updateRenewal({\n    subscriptionId: input.subscriptionId,\n    expectedVersion: input.version,\n    idempotencyKey: record.action!.idempotencyKey,\n  });\n\n  return decisions.appendOutcome(record.id, {\n    state: result.version === input.version + 1 ? \"verified\" : \"unknown\",\n    resultRef: `billing-event:${result.eventId}`,\n  });\n}\n```\n\nWhy not write the record only after success? Because a timeout, worker crash, or duplicate retry is itself important. The planned record gives you a durable idempotency key. The outcome distinguishes **attempted**, **verified**, and **unknown** instead of silently turning uncertainty into success.\n\nFor high-risk actions, send `hold` records to an approval queue. The reviewer should see the task summary, safe evidence links, policy reason codes, expected change, and expiry. Their approval becomes another append-only event, not an edit to history.\n\nA decision log can become a data leak if it stores raw prompts, retrieval text, tool arguments, or secrets. Use a privacy budget as deliberately as a token budget.\n\n| Store directly | Store by reference or hash | Never store in the decision log | \n|---|---|---|\n| policy version, verdict, time, action type | request ID, document ID, sanitized evidence hash | API keys, access tokens, raw passwords | \n| model and release IDs | encrypted artifact location | full hidden reasoning | \n| idempotency key, result status | redacted tool-response reference | unrestricted customer transcript | \n\nRun a redactor before persistence, and version it. A later investigation needs to know which rule set removed or transformed a field. Do not solve this with a vague instruction like “avoid PII.” Make it a schema rule: only allow-listed fields can enter the decision table.\n\nAlso separate read permissions. An on-call engineer may need policy codes and status, while a privacy officer may be the only person allowed to open the underlying request reference. The decision log should remain useful even when its linked artifacts are unavailable to the reader.\n\nReplay means testing whether the same typed facts and policy would produce the same verdict—not firing the live tool again.\n\nCreate a replay fixture with:\n\n``` js\nit(\"holds a renewal change when ownership evidence is missing\", async () => {\n  const result = await evaluate({\n    accountOwnerVerified: false,\n    requestedChange: \"renew\",\n    riskAmount: 0,\n  }, policy(\"renewal-policy-v4\"));\n\n  expect(result.verdict).toBe(\"hold\");\n  expect(result.reasons).toContain(\"account_owner_not_verified\");\n});\n```\n\nThis is more stable than asking a model to narrate its prior thinking. Keep LLM interpretation upstream, constrained by a typed output schema. Make the consequential verdict deterministic where possible. If the model must classify, capture the allowed label, evidence references, confidence as a diagnostic signal, and the policy that converted the label into an action.\n\nGood records turn vague questions into bounded queries:\n\n`hold` spike?”\nBuild a small operational view around those questions. Show the decision timeline, state transitions, policy reason codes, release ID, and safe links. Do not start with a polished chat summary; start with data that makes a summary checkable.\n\nAn especially useful alert is a rising count of `unknown` outcomes. It catches the uncomfortable middle state: a tool call may have reached the destination, but your service cannot yet prove what happened. That deserves reconciliation, not an automatic retry.\n\nIt feels thorough until a customer asks for deletion, a secret appears in context, or an operator needs one fact quickly. Store a constrained summary and references; keep original data in the system responsible for it.\n\nConfidence is not permission. A high score does not replace ownership checks, spend limits, or a reviewer. Policy should decide whether a classification may become an action.\n\nAn accepted request can still be rejected later, applied twice, or applied to the wrong version. Prefer a returned version, downstream event ID, read-after-write check, or reconciliation job.\n\nCorrections are valuable; overwriting history is not. Append a correction event that names the original record, the actor, and the reason. That preserves the investigation trail.\n\nStart with one high-value tool action, such as changing a ticket state or creating an outbound message. Define its input facts, risk tier, policy reasons, idempotency key, and verification method. Put the log schema under test before adding a dashboard.\n\nNext, add `hold` and `unknown` states. Teams often log only success and failure, but these two states capture the real operational ambiguity. Finally, link the record to your existing trace and release manifest. You get an explanation layer without replacing observability, deployments, or access control.\n\nThe result is less glamorous than an autonomous demo, but more useful in production: an agent can act quickly and still leave behind a clear, privacy-aware explanation of what it was allowed to do and what the system proved afterward.\n\nTreat a decision log as an API contract. If a release removes the policy version, changes a reason code, or starts serializing a raw email address, the build should fail before that change reaches production.\n\nStart with three small tests:\n\n`unknown` may become `verified` or `rejected`, but a terminal outcome cannot be silently overwritten.\n\n``` js\nit(\"does not persist raw customer text\", () => {\n  const record = toDecisionLog({\n    task: { summary: \"Email ava@example.com about invoice 8841\" },\n    policy: { verdict: \"hold\", reasons: [\"human_review_required\"] },\n  });\n\n  expect(JSON.stringify(record)).not.toContain(\"ava@example.com\");\n  expect(record.task.requestRef).toMatch(/^request:/);\n});\n```\n\nUse fixtures from incidents and near misses after redacting them. A fixture that proves a duplicated retry preserves its idempotency key is more valuable than a generic happy-path example. When a policy changes intentionally, review the fixture diff as carefully as an application permission change.\n\nOne final rule helps with multi-agent systems: assign a `parentDecisionId` when one agent delegates to another. The child still logs its own policy and result. This lets an investigator follow the tree from a user request to a tool action without merging all agents into one unsearchable transcript.\n\nInclude a stable decision ID, time, task reference, release and policy versions, safe evidence references, policy verdict and reason codes, the planned action, idempotency key, and a verified or unknown outcome. Avoid raw prompts and hidden reasoning.\n\nNo. Traces diagnose execution performance and call paths. Decision logs explain a consequential branch or action in business and policy terms. Link the two rather than forcing either one to do both jobs.\n\nNo. Store explicit, reviewable policy reasons and safe evidence references instead. They are more reliable for operations and avoid creating an unnecessary store of sensitive, unbounded model text.\n\nUse an idempotency key plus a domain-level confirmation: a returned version, durable downstream event, read-after-write check, or reconciliation process. Record `unknown` until that confirmation exists.\n\nMatch retention to the underlying business process, contractual needs, and privacy policy. Keep the smallest useful record, retain linked sensitive artifacts under their own controls, and test deletion and access rules.\n\nYes. Start with an append-only database table or event stream, typed schema validation, policy reason codes, and links to existing traces and records. Add specialized tooling only after the basic contract is working.", "url": "https://wpnews.pro/news/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought", "canonical_source": "https://dev.to/jackm-singularity/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought-3gp0", "published_at": "2026-10-10 05:50:28+00:00", "updated_at": "2026-10-10 06:00:48.853567+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "mlops", "developer-tools", "ai-infrastructure"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought", "markdown": "https://wpnews.pro/news/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought.md", "text": "https://wpnews.pro/news/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought.txt", "jsonld": "https://wpnews.pro/news/ai-agent-decision-logs-record-evidence-without-storing-chain-of-thought.jsonld"}}