{"slug": "the-llm-rewrote-the-email-after-i-approved-it", "title": "The LLM rewrote the email after I approved it", "summary": "Crusbro agent-OS, an agent operating system with 600,000+ hours of runtime, addresses AI agent failures by implementing gates, freeze snapshots, and replay mechanisms to prevent unauthorized outbound sends and writes. The system, designed for business teams, ties evidence to business order IDs and supports canary rollbacks, aiming to make AI controllable and auditable for enterprise use.", "body_md": "# crusbro agent-OS\n\nAgent Operating System · HARNESS · 600,000+ hours\n\nAgent failures are rarely “the model isn’t smart enough.” They are boundary failures: outbound sends, writes, session drift, version drift. crusbro agent-OS hardens harness insight into a runtime—so models finish defined goals under gates, freeze and replay.\n\n## What it is\n\ncrusbro agent-OS is a mature harness for driving models to complete defined work: goals, tools, memory, failure handling, human gates and evidence in one foundation. The model reasons; agent-OS keeps that reasoning **on an acceptable business track**.\n\nWhat to finish · How to accept\n\n**→**\n\nHarness · Gate · Freeze · Replay\n\n**→**\n\nSystem actions · Documents · Devices · Reports\n\n## What is not substitutable\n\nAnyone can rewrite another agent loop. What is hard to replace: production insight into harness failure modes—and defaults shaped by runtime data.\n\n| Attribute | Why generic frameworks fall short |\n|---|---|\n| Harness engineering insight | Addresses tool hallucination, accidental outbound, session confusion, missing rules and version drift with gate + freeze + replay—not a demo loop. |\n| 600,000 hours of runtime | Failures and dispositions feed defaults: retries, idempotency, gate thresholds and risk tags come from real operations. |\n| Freeze before outbound / write | Immutable snapshot before high-risk actions; execute only after approval. Models cannot bypass the physical gate. |\n| Order-level evidence packs | Sessions, tool calls, approvals and outbound tied to business order IDs. Acceptance unit is business fact, not scattered logs. |\n| Config revision bound to Run | On failure, return to the exact rules, prompts and tool permissions then in force—with canary and rollback. |\n| One acceptance language across lines | Automation, drawing, vision and robot share orchestration and audit. Org capability transfers; no new platform per agent. |\n\n## Role in build and orchestration\n\nTeams write goals, tools, gates and acceptance. agent-OS turns proven harness behavior into configurable defaults.\n\n| Stage | What the team does | What agent-OS hardens |\n|---|---|---|\n| Define | Goals, boundaries, acceptance | Task state machine, pass/fail, milestone gates |\n| Connect | Mail, forms, MES/QMS, device actions | Skills / Tools constraints, idempotency keys, risk tags |\n| Gate | Who approves, when humans must step in | HITL, freeze snapshots, high-risk blocks |\n| Pilot | Exceptions, rule fixes, sample return | Replay localization, metrics, failure feedback |\n| Release | Change flow, model or roles | Revision bound to Run, canary, audit, rollback |\n\n## For business teams\n\n- Replay by business order: context, tool sequence and versions are visible\n- Outbound and writes freeze before approval—no silent model sends\n- Gray and high-risk steps carry human gates; rules inherit across scenes\n- New scenes reuse proven failure handling instead of rediscovering edges\n- Connect drawing / vision / X-ray models through one tool surface\n\n## For managers\n\n- Deliverables fit ops manuals and audit trails: versions, gates, freeze snapshots, rollback paths\n- Turn “uncontrollable AI” into signable acceptance items\n- Observe success rate, human-intervention rate and outbound blocks\n- One foundation compounds across product lines—no chimney rebuilds\n- Control and customer acceptance share one ownership chain\n\n## Typical scenarios\n\n| Scenario | Where the non-substitutable edge shows |\n|---|---|\n| Cross-system workflows | Mail, forms, approvals, writeback; freeze + idempotency; compensate and replay by order |\n| Station QC / film reading | Schedule vision or X-ray models; gray-zone review and dual-sign; evidence to QMS/MES |\n| Drawing to process / BOM | Chain drawing models and rules; escalate low confidence; field–geometry lineage |\n| Robot task loops | Step perception–decision–execution; force gates and human takeover on hazardous acts |\n| Multi-role delivery | Manager / worker / auditor; Gate accept/reject drives the remediation loop |\n\n## Capability snapshot\n\n| Capability | What it does |\n|---|---|\n| Harness loop | Goal → plan → tools → observe → finish or escalate |\n| Gate & freeze | Snapshot first, approve, then execute high-risk actions |\n| Order-level replay | Tool sequences, HITL decisions and outbound indexed by order ID |\n| Multi-agent | Roles, delegation, parallel work and merge |\n| Tiered memory | Session vs long-term; searchable and expirable |\n| Sandbox | Allow-lists for files, network and credentials |\n| Skills / Tools | Shared I/O and permissions against systems and domain models |\n| Governance | Versioning, audit, observability, cost control and rollback |", "url": "https://wpnews.pro/news/the-llm-rewrote-the-email-after-i-approved-it", "canonical_source": "https://www.crusbro.com/en/product-architecture.html", "published_at": "2026-08-21 07:11:02+00:00", "updated_at": "2026-08-21 07:43:51.533099+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-infrastructure", "ai-tools"], "entities": ["crusbro agent-OS"], "alternates": {"html": "https://wpnews.pro/news/the-llm-rewrote-the-email-after-i-approved-it", "markdown": "https://wpnews.pro/news/the-llm-rewrote-the-email-after-i-approved-it.md", "text": "https://wpnews.pro/news/the-llm-rewrote-the-email-after-i-approved-it.txt", "jsonld": "https://wpnews.pro/news/the-llm-rewrote-the-email-after-i-approved-it.jsonld"}}