AI Agents Need a System of Record, Not Just a Dashboard Multisynapse introduced an Agent System of Record that catalogs every AI agent version, records each model call, tool call and policy decision at the source, and produces tamper-evident audit trails mapped to the NIST AI Risk Management Framework and ISO/IEC 42001. The product requires a signed-in human to approve, promote or enable automatic promotion of an agent version, and can require that a grading judge come from a different model vendor than the agent it evaluates. Multisynapse said the system works across frameworks and model vendors so teams can switch models without losing history, policies or grades. Your AI Agents Need a System of Record, Not Just a Dashboard Your agents have stopped drafting and started doing. They log in to vendor portals, approve invoices, file tickets and hand work to other agents. Most teams still run them the way they run a web service: ship it, watch a dashboard, read the logs when something breaks. That works until the first time someone asks a question a dashboard can't answer. Which version of this agent paid that invoice? Who approved putting it live? What was it allowed to touch that day, and can you prove it? Watching is not governing Observability tells you what happened. Governance decides what is allowed to happen, and keeps the proof. For software that only answers questions, watching was enough. For software that acts on your behalf, you need both, and you need them to agree. An answer in minutes, not an investigation Picture this: on Monday morning, finance tells you an agent paid a vendor twice over the weekend. With a dashboard, you start reading logs. With Multisynapse, you open the run and the answers are already there. - Which agent, which version. Every agent lives in one catalog, and every run names the exact version it ran. An edit stays a draft until someone promotes it. - What it was allowed to do. Every tool call was checked against your policies when it was made, under the identity of whoever started the run. The decision sits on the record next to the call. - What it actually did, and what it cost. Every model call, tool call and policy decision was recorded once, at the source. - Whether it was good enough to ship. That version was graded before it went live and sampled after, against the standard you set. If the new version is the problem, rolling back is one decision, not a redeploy. That is what we mean by an Agent System of Record: one place that knows every agent, what it was allowed to do and what it did. It works whatever framework or model vendor runs the agent, so you can switch models without losing the history, the policies or the grades. We wrote about why the version has to follow the agent all the way from evaluation to execution in You Didn't Deploy the Agent You Evaluated https://www.anuclei.com/blog/you-didnt-deploy-the-agent-you-evaluated . Graded before it goes live A new version of an agent earns its way into production. It runs against datasets built from your own real runs, with the outside world replayed rather than touched, so an eval never pays a real invoice. Then it has to clear a gate you define: quality, error rate, cost. We are strict about who does the grading: - The judge can't grade its own homework. A promotion can require the judge to come from a different model vendor than the agent it is grading. - The judge has to be good at judging. A promotion can also require a judge that has been measured against human labels and met a quality bar. If it drifts, it stops qualifying. - A person makes the call. Approving a version, promoting it, or switching on automatic promotion takes a signed-in human. An API key, or an agent holding one, can do none of these. When you want to grade a lot of runs at once, Multisynapse can use your model provider's batch interface for the judge calls, at the provider's batch price, without changing what the judge is asked. Evidence you can hand an auditor When a regulator, a customer or your own risk team asks what an agent did, the answer should be a record, not a reconstruction. - Tamper-evident history. Approvals, permission changes, promotions and policy decisions go into an audit trail where a changed or removed entry shows. - Signed run provenance. A run can carry a signed manifest of what produced it: the version, the context it was given, the tools it was offered. - Framework mapping. Multisynapse maps that evidence to the NIST AI Risk Management Framework and ISO/IEC 42001. It shows your alignment posture control by control, separates what is checked automatically from what a person attests, and exports a bundle for review. We say alignment posture on purpose. These frameworks describe good practice. Software can't certify you against them, but it can show the evidence in one place. Built for how teams actually work Governance that only works for a team of three isn't governance. Multisynapse is multi-tenant from the ground up, and the newest release adds the controls larger organizations ask for first: - Need-to-know projects. An organization can limit members to the projects they are added to, with no surprise access to everything else. Project groups give a team one grant that covers a whole set of projects, including ones added later. - Separated duties. Authoring an agent, stewarding its evals and approving its promotion can sit with three different people, each holding exactly that slice of authority. - Removal means removal. Taking someone out of an organization ends every role they held in it. - An assistant that shares, carefully. The built-in assistant answers questions about your agents, runs and costs, and prepares changes for you to confirm. Its conversations are private by default. You can share one with a teammate, and when they add to it, the assistant acts with their permissions, not yours. It can also read from the tools your admins choose to offer it, and only from tools that declare themselves read-only. Last month's release, which took certification from a single agent to whole conversations, is in Certify the Conversation, Not Just the Agent https://www.anuclei.com/blog/certify-the-conversation-not-just-the-agent . Works with what you already send You don't have to rebuild your agents to get any of this. Multisynapse speaks OpenTelemetry, the open standard for traces and logs, including the emerging conventions for generative AI. Point your existing exporter at it and your agents' traces and application logs land side by side, tied to the same runs. Your data is protected on the way in: - Your rules, before it leaves. The SDKs let you drop or edit a record before it is ever sent. - Your rules, at the door. Each organization can set its own redaction rules, applied as data arrives, on top of built-in detection of personal data. - Spend that adds up. Cost is computed from what each call actually used, at the price in force. You can break usage down by project and export it to the spreadsheet your finance team already uses. Pick any point on a cost or error chart and you can go straight to the runs behind it. At a glance Every run is checked and recorded, its results are graded, and the next version goes live only when a person decides it should. How we build We measure before we build. When a feature would cost more than it returns at the scale our customers run today, we don't ship it to tick a box. We set up the check that tells us when it's time. When we find a gap in our own product, we fix it at the cause and add a test that keeps it fixed. If your agents act on your behalf, in your systems, with your money or your customers' data, you need more than a view of what they did. You need a record of what they were allowed to do, who said so, and how you know it worked. That's what we built. To see it on your own agents, email info@anuclei.com mailto:info@anuclei.com and we'll set up a walkthrough.