{"slug": "ai-agents-need-more-than-logs-why-i-built-novafabric", "title": "AI Agents Need More Than Logs: Why I Built NovaFabric", "summary": "A developer built NovaFabric, an observability system that captures AI agent executions into portable, tamper-evident \"Run Capsules\" containing traces, model calls, tool calls, environment locks, and replay metadata. The project combines existing cryptographic standards rather than new cryptography, and in a controlled experiment its structural diff correctly localized all 140 injected mutations across seven mutation classes. The developer notes the system is tamper-evident, not tamper-proof, and defines four replay modes since AI executions cannot be reproduced bit-for-bit.", "body_md": "AI agents are becoming more powerful.\n\nThey can:\n\nBut there is a simple question that is still surprisingly hard to answer:\n\n**After an AI agent finishes its work, can we prove what it actually did?**\n\nUsually, the answer is: **not completely**.\n\nThat is the problem behind **NovaFabric**.\n\nToday we have many good observability tools. They can show things like:\n\n```\nAgent started\n↓\nLLM called\n↓\nTool called\n↓\nAPI called\n↓\nFile modified\n↓\nAgent finished\n```\n\nThis is useful for debugging. But imagine that six months later something goes wrong. An auditor, security engineer, researcher, or developer may ask:\n\nA traditional trace does not necessarily answer all of these questions. And more importantly, a normal database record can potentially be changed later.\n\nThat is where the idea of **execution evidence** becomes useful.\n\nNovaFabric takes a different approach. Instead of thinking:\n\n```\nAI Agent → Logs\n```\n\nit thinks:\n\n```\nAI Agent\n   ↓\nExecution\n   ↓\nPortable Evidence\n```\n\nNovaFabric captures a run into something called a **Run Capsule**: a structured folder describing one execution.\n\n```\ncapsule/\n├── capsule.yaml\n├── trace.jsonl\n├── model-calls.jsonl\n├── tool-calls.jsonl\n├── env.lock\n├── redaction-proof.json\n├── replay.yaml\n└── ...\n```\n\nThe idea is simple:\n\nThe execution should leave behind something you can keep, inspect, compare, replay, sign, and share.\n\nThe research paper models the capsule using **15 entity types**, including model calls, tool calls, file events, network calls, human approvals, environment information, lineage, redaction records, replay information, and tamper-evidence.\n\n```\nCapture → Seal → Replay → Diff → Audit\n```\n\nFirst, capture what happened during the execution:\n\n```\nnova capture python my_agent.py\n```\n\nNovaFabric can wrap a command without requiring the application itself to be rewritten. It aims to capture model calls, tool calls, files, network activity, environment, inputs and outputs. The result is a portable **Run Capsule**.\n\nCapturing information is not enough. If somebody modifies the record afterward, we want that modification to be detectable.\n\nNovaFabric combines established technologies:\n\nThe important point: NovaFabric is **not inventing new cryptography**. It combines existing standards into one evidence system for AI-agent executions. That integration, rather than a new signing algorithm, is the paper's main contribution.\n\n```\nBefore sealing:  \"This file says the agent did X.\"\n\nAfter sealing:   \"This file says the agent did X,\n                  and we can detect if somebody changes it later.\"\n```\n\nThere is an important limitation here. **Tamper-evident does not mean tamper-proof.** NovaFabric can detect later modification of properly sealed evidence. It cannot prove that a trusted signing party told the truth when the evidence was originally created. The paper explicitly defines this trust boundary.\n\nAI systems have another problem: reproducibility. Run the same prompt tomorrow and you may not get the same answer. So NovaFabric does not pretend every AI execution can be reproduced bit-for-bit. Instead, it defines four replay modes:\n\nThe paper is deliberately careful here: only the mocked mode is deterministic by construction with respect to model inputs, while forensic mode executes nothing.\n\nRun B produces a bad result. What changed? Maybe the model, the prompt, a dependency, a tool response, the environment, or the output. NovaFabric provides a structural diff between capsules, comparing the structure of runs instead of giant text logs.\n\nIn the paper's controlled experiment, structural diff correctly localized all **140 injected mutations** across seven tested mutation classes. That could make questions like this easier to investigate:\n\n\"It worked yesterday. Why did the agent fail today?\"\n\nFinally, the evidence can be packaged into an **Evidence Bundle**. Another person should not have to trust your database; you can hand them a portable artifact containing the evidence and the cryptographic information needed to verify it.\n\n```\nAgent execution → Run Capsule → Cryptographic seal → Evidence Bundle → Auditor / Researcher / Security Team\n```\n\nThe paper specifies an offline verification path using standard technologies rather than requiring the verifier to trust a NovaFabric server. However, it also states that this independent-verifier interoperability path had **not yet been exercised by an independently implemented verifier** at publication time. That distinction matters.\n\nI like OpenTelemetry, and NovaFabric uses OpenTelemetry concepts. But the problems are different:\n\n| Observability | Execution evidence | \n|---|---|\n| What is happening? | What happened? | \n| Find errors | Reconstruct the execution | \n| Metrics and traces | Portable run artifact | \n| Operational debugging | Audit and forensics | \n| Usually mutable storage | Tamper-evident evidence | \n| Live system | Historical execution | \n\nNovaFabric is **not intended to replace Prometheus, OpenTelemetry, Langfuse, LangSmith, or similar systems**. It focuses on another question:\n\nCan a past execution be replayed, compared, and proven?\n\nCapturing everything creates another problem: an agent may see API keys, tokens, passwords, or credentials, and you do not want them inside a bundle sent to another company or auditor. NovaFabric includes secret redaction and a redaction attestation.\n\nBut a redaction attestation means \"the scanner says it detected and removed this secret.\" It does **not** prove that an unknown secret does not exist somewhere else in the capsule. The paper makes this explicit: completeness depends on the scanner, its rules, and the locations it checks. Being clear about these boundaries is especially important for security tools.\n\n```\nModel calls captured:                 100% in tested scenarios\nMCP tool calls captured:              100% in tested scenarios\nTested tampering classes detected:    3 / 3\nStructural mutations localized:       140 / 140\nRedaction (14 credential types):      14 / 14 removed\nProvenance graph, 10M-edge blast-radius p99:   45.5 ms\nProvenance graph, 100M-edge single-client p99: 167.9 ms\n```\n\nBut some results exposed weaknesses too. Mocked replay avoided live model calls in **10/10** tested scenarios, but only **2/10** tool-using workloads completed, because tool responses were not yet being substituted in the evaluated version. That was one of the most useful outcomes of the research: the evaluation did not only produce good benchmark numbers, it found real defects.\n\n**\"The pipeline ran successfully\" is not the same as \"the evidence is correct.\"**\n\nSeveral problems were discovered because the tests checked what *should have been recorded*, rather than simply checking whether the program exited successfully. At one point some evidence streams were missing even though capture appeared to work. In another experiment, an ingest service returned successful HTTP responses even though its configured metadata database had no tables.\n\n```\n\"Did it run?\" is a weak test.\n\"Can we prove the expected things actually happened?\" is a much stronger one.\n```\n\nNovaFabric is open source and self-hosted. The project has moved beyond the exact version evaluated in the paper, so the paper should be read as a **research snapshot**, not a complete description of today's repository.\n\n```\npip install novafabric\n\nnova capture python my_agent.py\nnova validate <capsule>\nnova replay <capsule> --mode forensic\nnova diff <capsule-a> <capsule-b>\n```\n\nThe project is still **beta**. Some local functionality is more mature, while server, cluster-scale, and some at-scale components remain experimental. I do not want to describe experimental infrastructure as production-proven.\n\nAs AI agents move from chat interfaces into systems that actually **do things** - changing production infrastructure, approving financial operations, modifying software, moving important data, operating HPC systems, coordinating other agents - observability alone will not be enough. Eventually someone will ask: *What exactly happened? Can you prove it? Can you reproduce it?*\n\nThat is the space NovaFabric is exploring. Not another agent framework, not another tracing dashboard, but infrastructure for turning an execution into something we can **capture → seal → replay → compare → audit**.\n\nNovaFabric is Apache-2.0 and open source:\n\nIf you work on AI agents, MLOps, HPC, security, observability, provenance, or reproducibility, I would love your feedback. Especially this question:\n\n**When your AI agent makes an important decision today, what evidence will you still have six months from now?**", "url": "https://wpnews.pro/news/ai-agents-need-more-than-logs-why-i-built-novafabric", "canonical_source": "https://dev.to/mskazemi/ai-agents-need-more-than-logs-why-i-built-novafabric-1emd", "published_at": "2026-10-07 10:10:58+00:00", "updated_at": "2026-10-07 10:17:04.996922+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "mlops", "developer-tools", "ai-infrastructure"], "entities": ["NovaFabric"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agents-need-more-than-logs-why-i-built-novafabric", "markdown": "https://wpnews.pro/news/ai-agents-need-more-than-logs-why-i-built-novafabric.md", "text": "https://wpnews.pro/news/ai-agents-need-more-than-logs-why-i-built-novafabric.txt", "jsonld": "https://wpnews.pro/news/ai-agents-need-more-than-logs-why-i-built-novafabric.jsonld"}}