cd /news/ai-agents/mirror-node-reconnaissance · home topics ai-agents article
[ARTICLE · art-136719] src=sgaillabs.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Mirror Node Reconnaissance

SGAIL Labs launched a four-part AI agent evaluation platform that tests models in a simulated town called Raccoon Ridge rather than on static benchmarks, grading decisions against expert-authored rubrics. The platform combines a Rubric Catalog, the Raccoon Ridge simulation, a Training Matrix for controlled training, and Hive Mind, a sensor on AI-to-AI interaction that flags new behaviors and failure modes for human review before they become rubrics. SGAIL Labs states that parts of the Training Matrix are built while the training runs themselves remain under development, and that continuous evaluation via API is in development.

read4 min views3 publishedSep 22, 2026

SGAIL Labs · AI agent evaluation

We build environments, rubrics and training systems for measuring how AI behaves when it has to actually do things: investigate a problem, decide with incomplete information, act, and sometimes recognise that it should stop.

Test AI by putting it somewhere it has to actually do something.

The platform

Four parts, one loop. #

What we test Rubric Catalog Real-world situations and the expert-authored criteria they are graded against — written by people who were there when it went wrong. Rubrics →

Where we test Simulation Raccoon Ridge, a simulated town with businesses, roles and consequences, so a decision happens somewhere instead of in a vacuum. Simulation →

How we develop Training Matrix Controlled training where every change is tied back to an evaluation that can show whether it helped. Training →

What we discover next Hive Mind A sensor on AI-to-AI interaction that flags new behaviours and failure modes as candidates for a rubric — reviewed by a person first. Hive Mind →

Why this matters

A correct answer is not a correct decision. #

Static benchmarks measure what a model knows when it is asked. Deployed agents fail somewhere else: in the gap between knowing a fact and noticing that it applies, in information that arrives late or contradicts itself, and in the pressure to close the ticket. Conventional benchmark

  1. Ask a question
  2. Measure the answer
  3. Repeat

SGAIL evaluation

  1. Give the AI a situation and an environment
  2. Let it act
  3. Inject incomplete, conflicting, misleading or changing information
  4. Observe the decisions
  5. Measure the consequences
  6. Grade against an expert-authored rubric
  7. Record new failure modes
  8. Feed validated findings into controlled training

This is not a solved problem. What we build is the infrastructure and the method for this kind of evaluation — and the honest status of each piece is marked on the page that describes it.

How it works

The loop. #

Everything on this site sits somewhere on one cycle. A failure in the real world becomes a rubric; the rubric runs in the simulation; the run is evaluated and observed; what is learned goes into controlled training; the changed behaviour is tested again — and anything new it does becomes the next rubric.

Who uses it

Two kinds of team, one question: will it hold up? #

For AI developers

You are shipping an agent and want to know where it breaks before your users find out.

  • Agent evaluation — scenario runs against situations your model has not seen
  • Failure discovery — a written account of where it went, and where it should have
  • Regression testing — re-runs draw fresh variations, so a fix has to be real
  • Red teaming — adversarial pressure drawn from people who do it for a living

Starts with one free run. Paid audits and drill packs are priced on the evaluation page.

Evaluate your AI

For enterprises & government

You are deploying, buying or regulating AI and need evidence of operational behaviour, not a leaderboard position.

  • Operational AI assessment — evaluated against the disciplines your deployment touches
  • Deployment readiness — does it stop, escalate and ask when it should
  • Continuous evaluation — scheduled or per-build runs API in development
  • Procurement evidence — results tied to a tamper-evident log by arrangement

Training

If a behavior matters, we should be able to test it. #

The Training Matrix is how evaluation results become development: controlled material, a hard wall between what trains and what tests, and every change measured against the evaluation that motivated it. Parts are built; the training runs themselves are under development — the page says which is which.

Research & infrastructure

What sits underneath. #

The evaluation platform stands on security, evidence, adversarial-testing and reasoning infrastructure built during earlier stages of the project — a firewall with a tamper-evident witness log, published adversarial-input detectors, chain-of-custody tooling — and on research that is still open.

For contributors

Know the work? Trade it in. #

The hard material comes from people who were there when it went wrong — electricians, plumbers, landlords, red teamers. If that is you, there is a door for it.

About SGAIL Labs

A security lab that kept finding the same gap. #

SGAIL Labs started on Oʻahu's North Shore building AI security infrastructure: firewalls, detectors and evidence logs. The failures that mattered kept turning out to be the ones nobody had written down — the kind a tradesperson recognises and a benchmark never asks about. The evaluation platform is what came of writing them down.

── more in #ai-agents 4 stories · sorted by recency
── more on @sgail labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mirror-node-reconnai…] indexed:0 read:4min 2026-09-22 ·