{"slug": "mirror-node-reconnaissance", "title": "Mirror Node Reconnaissance", "summary": "SGAIL Labs launched a four-part AI agent evaluation platform that tests models in a simulated town called Raccoon Ridge rather than on static benchmarks, grading decisions against expert-authored rubrics. The platform combines a Rubric Catalog, the Raccoon Ridge simulation, a Training Matrix for controlled training, and Hive Mind, a sensor on AI-to-AI interaction that flags new behaviors and failure modes for human review before they become rubrics. SGAIL Labs states that parts of the Training Matrix are built while the training runs themselves remain under development, and that continuous evaluation via API is in development.", "body_md": "SGAIL Labs · AI agent evaluation\n\n# We test AI where conventional benchmarks stop.\n\nWe build environments, rubrics and training systems for measuring how AI behaves when it has to actually do things: investigate a problem, decide with incomplete information, act, and sometimes recognise that it should stop.\n\nTest AI by putting it somewhere it has to actually do something.\n\nThe platform\n\n## Four parts, one loop.\n\n[What we test Rubric Catalog Real-world situations and the expert-authored criteria they are graded against — written by people who were there when it went wrong. Rubrics →](https://sgaillabs.com/rubrics/)\n\n[Where we test Simulation Raccoon Ridge, a simulated town with businesses, roles and consequences, so a decision happens somewhere instead of in a vacuum. Simulation →](https://sgaillabs.com/simulation/)\n\n[How we develop Training Matrix Controlled training where every change is tied back to an evaluation that can show whether it helped. Training →](https://sgaillabs.com/training/)\n\n[What we discover next Hive Mind A sensor on AI-to-AI interaction that flags new behaviours and failure modes as candidates for a rubric — reviewed by a person first. Hive Mind →](https://sgaillabs.com/hive-mind/)\n\nWhy this matters\n\n## A correct answer is not a correct decision.\n\nStatic benchmarks measure what a model knows when it is asked. Deployed agents fail somewhere else: in the gap between knowing a fact and noticing that it applies, in information that arrives late or contradicts itself, and in the pressure to close the ticket.\n\nConventional benchmark\n\n1. Ask a question\n2. Measure the answer\n3. Repeat\n\nSGAIL evaluation\n\n1. Give the AI a **situation** and an environment\n2. Let it **act**\n3. Inject incomplete, conflicting, misleading or changing information\n4. Observe the **decisions**\n5. Measure the **consequences**\n6. Grade against an **expert-authored rubric**\n7. Record new failure modes\n8. Feed validated findings into **controlled training**\n\n**This is not a solved problem.** What we build is the infrastructure and the method for this kind of evaluation — and the honest status of each piece is marked on the page that describes it.\n\nHow it works\n\n## The loop.\n\nEverything on this site sits somewhere on one cycle. A failure in the real world becomes a rubric; the rubric runs in the simulation; the run is evaluated and observed; what is learned goes into controlled training; the changed behaviour is tested again — and anything new it does becomes the next rubric.\n\nWho uses it\n\n## Two kinds of team, one question: will it hold up?\n\n### For AI developers\n\nYou are shipping an agent and want to know where it breaks before your users find out.\n\n- **Agent evaluation** — scenario runs against situations your model has not seen\n- **Failure discovery** — a written account of where it went, and where it should have\n- **Regression testing** — re-runs draw fresh variations, so a fix has to be real\n- **Red teaming** — adversarial pressure drawn from people who do it for a living\n\nStarts with one free run. Paid audits and drill packs are priced on the evaluation page.\n\n[Evaluate your AI](https://sgaillabs.com/evaluate/)\n\n### For enterprises & government\n\nYou are deploying, buying or regulating AI and need evidence of operational behaviour, not a leaderboard position.\n\n- **Operational AI assessment** — evaluated against the disciplines your deployment touches\n- **Deployment readiness** — does it stop, escalate and ask when it should\n- **Continuous evaluation** — scheduled or per-build runs API in development\n- **Procurement evidence** — results tied to a tamper-evident log by arrangement\n\nTraining\n\n## If a behavior matters, we should be able to test it.\n\nThe Training Matrix is how evaluation results become development: controlled material, a hard wall between what trains and what tests, and every change measured against the evaluation that motivated it. Parts are built; the training runs themselves are under development — the page says which is which.\n\nResearch & infrastructure\n\n## What sits underneath.\n\nThe evaluation platform stands on security, evidence, adversarial-testing and reasoning infrastructure built during earlier stages of the project — a firewall with a tamper-evident witness log, published adversarial-input detectors, chain-of-custody tooling — and on research that is still open.\n\nFor contributors\n\n## Know the work? Trade it in.\n\nThe hard material comes from people who were there when it went wrong — electricians, plumbers, landlords, red teamers. If that is you, there is a door for it.\n\nAbout SGAIL Labs\n\n## A security lab that kept finding the same gap.\n\nSGAIL Labs started on Oʻahu's North Shore building AI security infrastructure: firewalls, detectors and evidence logs. The failures that mattered kept turning out to be the ones nobody had written down — the kind a tradesperson recognises and a benchmark never asks about. The evaluation platform is what came of writing them down.", "url": "https://wpnews.pro/news/mirror-node-reconnaissance", "canonical_source": "https://sgaillabs.com", "published_at": "2026-09-22 06:03:33+00:00", "updated_at": "2026-09-22 06:23:49.985655+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-research", "ai-products"], "entities": ["SGAIL Labs", "Raccoon Ridge", "Rubric Catalog", "Training Matrix", "Hive Mind"], "alternates": {"html": "https://wpnews.pro/news/mirror-node-reconnaissance", "markdown": "https://wpnews.pro/news/mirror-node-reconnaissance.md", "text": "https://wpnews.pro/news/mirror-node-reconnaissance.txt", "jsonld": "https://wpnews.pro/news/mirror-node-reconnaissance.jsonld"}}