cd /news/ai-agents/novelty-gate-sentinel-confidence-is-… · home › topics › ai-agents › article
[ARTICLE · art-149112] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Novelty Gate Sentinel— confidence is not novelty.

A developer built Sentinel, an autonomous code agent that watches a JavaScript codebase and separates confidence from novelty via a "Novelty Gate" that fingerprints each file's content, risk, role and rules so unchanged files reuse a prior SKIP decision instead of re-reasoning. Live analytics from the production instance on Render, read 2026-10-11, showed 188,000 reuses against 42 fresh evaluations across 5 tracked files — roughly 4,476 reuses per fresh evaluation — with an estimated 314,893,830 tokens not re-read, a figure the developer cautions is a rough size estimate rather than actual billed tokens.

by read21 min views1 publishedOct 11, 2026

How an autonomous code agent learned to say "I've seen this exact file before, nothing moved, my answer is still no" — and what that looks like after three months in production.

Live numbers in this article come from the production Sentinel instance on Render, read from the read-only GET /analytics endpoint on 2026-10-11 08:24 UTC (instance 4cd95249), then re-read at 08:26:55 UTC with identical results. The captured raw output is reproduced at the end of the article, in Captured output. Nothing here is a projection or a synthetic benchmark. Where a number is an estimate, I say so and show the formula.

Sentinel is an autonomous agent that watches a JavaScript codebase. Every couple of minutes it wakes up, looks at a batch of files, and decides one of three things for each file:

Most of the time the answer is SKIP. That is good. A security-focused agent that changes code on every pass is not an agent, it is an incident.

The trouble was how it got to SKIP. Every pass it re-read the file, rebuilt its picture of the file's risk and role, re-ran its decision policy, and — surprise — arrived at the same SKIP as two minutes earlier. Same file, same risk, same role, same rules, same answer. Over and over.

If you have ever re-opened the fridge hoping something new appeared since thirty seconds ago, you already understand the bug.

This is the core idea of the article, so here it is in one line:

Confidence answers "how reliable is my judgement?". Novelty answers "has anything changed since I last judged?". They are different questions.

It is tempting to treat repeated agreement as growing certainty: "I said SKIP ten times in a row, so I must be really sure." That is a trap. Ten identical inputs give you the same evidence ten times, not ten pieces of evidence. If confidence drifted upward just because a judgement was repeated, the agent would slowly talk itself into trusting a module it has learned nothing new about.

So Sentinel keeps the two apart:

If it is the same situation, and the last answer was a plain SKIP, Sentinel simply reuses that answer. No re-reasoning, no new memories, no LLM. That check is the Novelty Gate.

Not just "the file's bytes are identical". The gate takes a fingerprint of everything that should force a fresh look:

Change any one of these and the fingerprint changes, so the file gets a full, fresh evaluation.

The gate is deliberately timid:

SENTINEL_CONTEXT.md, line 104). So "unchanged" means "unchanged Here is the novelty block exactly as the live instance returned it:

{
  "trackedFiles": 5,
  "freshEvaluations": 42,
  "reuses": 188000,
  "reuseRate": 1,
  "estimatedTokensSaved": 314893830
}

In plain words:

Field Value What it means
trackedFiles 5 Files the gate currently remembers a decision for. Small, because Sentinel's scanner looks at one directory ( libs/shared , 7 files) and processes at most 5 per pass.
freshEvaluations 42 Times one of those files went through the full decision pipeline and ended in a SKIP that was remembered.
reuses 188,000 Times a remembered SKIP was reused instead of re-reasoning.
reuseRate 1 (really 0.99978) reuses / (reuses + freshEvaluations) , rounded to two decimals — so itdisplays as 1. It is not literally 100 %.
estimatedTokensSaved 314,893,830 An estimate of how much input Sentinel did not have to re-read. Read the next section before quoting this anywhere.

That is roughly 4,476 reuses for every fresh evaluation. For a five-file working set that mostly doesn't change, that is what you would expect: once a stable file has been judged, the gate keeps saying "same as before" until something moves.

That number is real in the sense that the endpoint returns it. It is not "315 million OpenAI tokens we would otherwise have paid for". Three reasons:

ceil(chars / 4) + 400 "tokens" — a rough size of the file as an LLM prompt plus a fixed overhead. Nobody counted actual tokens. The honest summary: the Novelty Gate turned an agent that re-thought the same five files on every pass (every two minutes by design, and for a stretch in July every ~5 seconds, see Part 2) into one that re-thinks them only when something relevant changes. The token estimate is a size-of-avoided-work indicator, not a bill.

If you only came for the idea, you can stop here. The rest is code.

The gate lives in libs/core/novelty-gate.js and is called from Executor.processFile (libs/core/executor.js), right after the decision envelope is built and before the compact IR, reference assessment, decision policy and any LLM call.

flowchart TD
    A["mainLoop tick"] --> B{"Wake Gate: actionable event?"}
    B -- "no" --> Z["no-op, nothing scanned"]
    B -- "yes" --> C["RepositoryScanner: libs/shared, batch of 5"]
    C --> D["processFile: build envelope"]
    D --> E["NoveltyGate.fingerprint"]
    E --> F{"scheduled review due?"}
    F -- "yes" --> H["full pipeline"]
    F -- "no" --> G{"NoveltyGate.evaluate"}
    G -- "novel" --> H
    G -- "unchanged and prev SKIP" --> R{"autonomy token reconsiders?"}
    R -- "yes" --> H
    R -- "no" --> S["reuse previous SKIP, count tokens, return reused:true"]
    H --> P["IR, policy, maybe LLM, gates"]
    P --> Q{"outcome"}
    Q -- "SKIP" --> M["NoveltyGate.remember"]
    Q -- "EVOLVE / REJECT / DEFERRED" --> N["not remembered"]

Note the two escape hatches before reuse: a due scheduled review (added after a real bug where a cached SKIP swallowed a review, pinned in tests/unit/executor-scheduled-review.test.js) and an autonomy token ( reconsider("novelty-reuse")).

libs/core/novelty-gate.js, lines 35–69:

function fingerprint(input = {}) {
    const {
        content,
        risk = null,
        role = null,
        confidence = null,
        status = null,
        rollbackStreak = 0,
        penalty = 0,
        evidenceVersion = null
    } = input;

    const contentHash = crypto
        .createHash("sha256")
        .update(String(content ?? ""))
        .digest("hex")
        .slice(0, 16);

    const conf = Number.isFinite(Number(confidence))
        ? Math.round(Number(confidence) * 100) / 100
        : "na";

    return [
        contentHash,
        risk ?? "na",
        role ?? "na",
        conf,
        status ?? "na",
        Number(rollbackStreak) || 0,
        Number(penalty) || 0,
        evidenceVersion ?? "na"
    ].join("|");
}

A few deliberate choices:

0.701 and 0.704 produce the same fingerprint (there is a test for exactly this). Floating-point jitter is not novelty. A move from 0.70 to 0.71 evidenceVersion is the world outside the file.OpportunityEngine.evidenceVersion(memory, { dependencies: readDependencyManifest() }). Without it, a frozen module would reuse its old SKIP forever, even after the governance rules or dependencies moved — and the Opportunity Engine could never reach exactly the modules it exists to re-examine. Small aside: the fingerprint description in console.log. When your cache key is human-readable, "why did this miss?" takes seconds instead of an afternoon. What is not in it, by design for now: branch/MR state, human confirmation, production runtime signals. Those are external observers, and adding them is listed as a future extension in SENTINEL_CONTEXT.md:104.

// Only a prior SKIP is deterministic enough to reuse.
function reusable(decision) {
    return !!decision &&
        decision.action === "SKIP" &&
        typeof decision.reason === "string";
}

function evaluate(memory, filePath, fp) {
    const prev = store(memory)[filePath];
    if (!prev) return { novel: true, reason: "first-seen", previous: null };
    if (prev.fingerprint !== fp) return { novel: true, reason: "changed", previous: prev };
    if (!reusable(prev.decision)) return { novel: true, reason: "prev-not-reusable", previous: prev };
    return { novel: false, reason: "unchanged", previous: prev };
}

remember() is only called on the three SKIP return paths in the executor — lifecycle skip, local "LLM not needed" skip, and intent-gate skip. Its bookkeeping per file:

s[filePath] = {
    fingerprint: fp,
    decision: { action: decision.action, reason: decision.reason ?? null },
    firstTs: prev?.firstTs ?? now,
    lastTs: now,
    evaluations: (prev?.evaluations ?? 0) + 1,
    reuses: sameFp ? (prev.reuses ?? 0) : 0,
    tokensSaved: sameFp ? (prev.tokensSaved ?? 0) : 0
};

Read the last two lines twice, because they matter for interpreting the production numbers: when the fingerprint changes, reuses and tokensSaved reset to zero, but evaluations keeps counting. So the live reuses: 188000 is "reuses since each file's current fingerprint was established", not an all-time total. If anything, it under-counts lifetime reuse.

On the hot path, the executor does this when the fingerprint is unchanged (libs/core/executor.js, ~lines 308–329):

const saved = DecisionAnalytics.estimateTokens(
    typeof content === "string" ? content.length : 0
);
NoveltyGate.reuse(memory, filePath, saved);
// ... ObservationMode.observe(...) labels the SKIP, never changes it
MemoryManager.clearWorking(memory);
return {
    action: novelty.previous.decision.action,
    reason: novelty.previous.decision.reason,
    reused: true,
    envelope
};

No decision-stream entry, no experience summary, no analytics row. That is the point: a reused decision is not a new memory.

I had been describing the estimate as "chars / 4". The actual function is in libs/knowledge/decision-analytics.js, lines 18–21:

function estimateTokens(contentLength = 0) {
    const body = Math.ceil(Math.max(0, Number(contentLength) || 0) / 4);
    return body + 400;
}

So it is ceil(chars / 4) + 400 — the body plus a fixed prompt overhead. It is the same estimator DecisionAnalytics uses for every locally-handled decision, so novelty numbers and analytics numbers are at least in the same (rough) currency.

Can estimatedTokensSaved = 314,893,830 actually come from reuses = 188,000 under that formula?

crypto-vault.js, 18,743 chars → 5,086): 956,168,000. ✔ below it. master ranged from 1,370 to 18,743 chars over the period (translator-gateway.js from 5,235 to 7,241). An average of ~5,100 chars across a 5-file batch is plausible for that mix. What I could not do: an exact per-file reconciliation. The per-file memory.decision_memory records (with their own reuses/ tokensSaved) are not exposed by any read-only endpoint, and file sizes changed mid-period. So the claim is "consistent with the formula and within bounds", not "verified to the token".

Being evidence-first means writing these down instead of rounding them away. My first draft had a section called "things I can't explain", and its main item was this:

The loop runs every 2 minutes with a batch of 5. From the gate's merge (2026-07-18) to the Wake Gate (2026-09-02) is ~46 days → at most ~165,600 reuses, plus ~5,520 since (1,104 wakes × 5). Ceiling ≈ 171,000. Reported: 188,000.

188,000 is above that, which should be impossible. So I went digging. The explanation is that the ceiling was wrong, not the counter.

The loop did not run every 2 minutes. Sentinel's Observation Mode keeps a log of six-hour "watch windows". Each window records how many times a file was observed, and the novelty-reuse path itself calls ObservationMode.observe() (executor.js, inside the reuse branch). Here are all the retained windows for libs/shared/translator-gateway.js, straight from GET /observations:

Window (UTC) Observations Seconds per observation
07-24 23:43 → 07-25 05:43 3,322 6.5
07-25 23:43 → 07-26 05:43 4,613 4.7
07-27 05:43 → 07-27 11:43 4,011 5.4
07-28 11:44 → 07-28 17:44 3,421 6.3
07-29 17:44 → 07-29 20:05 (incident) 1,589 5.3

(20 windows retained, 2026-07-24 → 07-29; all between 4.7 and 6.5 s.) A 2-minute timer would give 180 observations per window. The real number was 18–26× higher: one pass roughly every 5 seconds.

Who called the loop that often? There are exactly two callers of mainLoop() in app.js:

app.post('/webhook',(_req,res)=>{
    res.status(200).send('OK');
    mainLoop();
});

setInterval(mainLoop, CONFIG.SYSTEM.INTERVAL); // 2 min

The guard sits at the very top of mainLoop(), before the first await (app.js, master):

async function mainLoop(){
    if(isBusy) return;          // first statement
    isBusy=true;                // set synchronously, no await in between
    let memory=null;
    let lockAcquired=false;
    try{
        memory=await MemoryManager.load();
        lockAcquired=RuntimeLock.acquire(memory, INSTANCE_ID);
        if(!lockAcquired){ return; }
        // ... wake gate, batch, executor ...
    }catch(err){ /* log */ }
    finally{
        if(lockAcquired && memory){ RuntimeLock.release(memory); }
        isBusy=false;           // only reset here
    }
}

The check and the set run in the same synchronous step, and Node runs one step at a time, so two passes can never overlap inside one process. RuntimeLock does something else: it keeps a different instance out (owner = INSTANCE_ID, 240 s TTL). Render ran a single instance anyway. The guard was in the same position in the July code (app.js at 3d8a0b3, 2026-07-18).

So POST /webhook is the only code path that can produce passes faster than every 2 minutes. A webhook call is not queued: it calls mainLoop() without await, and if a pass is already running, the call returns immediately and nothing happens. A pass every ~5 seconds therefore means webhook calls were arriving at least that often, each one landing after the previous pass had finished.

A second, independent check. Observation Mode has labelled 1,678,193 skips since it was merged on 2026-07-18, 85 days ago. At one pass every 2 minutes with 5 files, that would take ~466 days of non-stop running. At one pass every ~5 seconds, it takes ~19 days.

What this means for 188,000. At ~5 s per pass and 5 files per batch, the gate could log ~86,000 reuses a day, so 188,000 is roughly two days' worth. Also, a file's reuses counter resets to 0 each time its fingerprint changes. The live number is therefore "reuses since each file's last fingerprint change", not a lifetime total. That makes it smaller than the lifetime count, not larger.

What I still can't prove. Who sent those webhook calls. The code explains how the cadence was possible, and the observation log shows it happened. Render logs from July are past retention, and I had no access to the GitLab project's webhook settings. So the sender (most likely a GitLab project hook firing on Sentinel's own activity, but that's a guess) stays unconfirmed. Today it no longer matters much for cost: since 2026-09-02 every pass goes through the Wake Gate first, so a webhook-triggered pass with nothing new is a cheap no-op. The live counter shows 14,625 ticks, 92.45 % of them no-ops.

Two smaller notes:

/analytics reads and one /dashboard read, 08:24–08:48 UTC) and it didn't move. That's expected: since the Wake Gate, a no-op tick does not open files, so it does not reuse anything. A round number is a coincidence, not a sign of a cap; novelty-gate.js has no cap on reuses. reuseRate rounds to 1.stats() rounds to two decimals; the exact value is 0.999777. None of this changes the qualitative result, but it changes the honest reading. The big counter partly measures a loop that ran far more often than designed. Some of the "saved" work is work that only existed because the loop was over-triggered. The Novelty Gate made that over-triggering cheap; the Wake Gate later made it pointless.

const STORE_LIMIT = 500;

function enforceLimit(s, limit = STORE_LIMIT) {
    const keys = Object.keys(s);
    if (keys.length <= limit) return;
    keys
        .map(k => [k, Number(s[k]?.lastTs) || 0])
        .sort((a, b) => a[1] - b[1])
        .slice(0, keys.length - limit)
        .forEach(([k]) => { delete s[k]; });
}

At most 500 files are remembered; the least-recently touched ones are evicted first. Today the instance tracks 5, so this limit is a seatbelt, not a feature in use. Eviction is safe by construction: a forgotten file is just "first-seen" next time and gets a fresh evaluation.

function stats(memory) {
    const s = (memory && typeof memory.decision_memory === "object")
        ? memory.decision_memory
        : {};
    const files = Object.values(s);

    const freshEvaluations = files.reduce((n, r) => n + (Number(r.evaluations) || 0), 0);
    const reuses = files.reduce((n, r) => n + (Number(r.reuses) || 0), 0);
    const tokensSaved = files.reduce((n, r) => n + (Number(r.tokensSaved) || 0), 0);
    const denom = freshEvaluations + reuses;

    return {
        trackedFiles: files.length,
        freshEvaluations,
        reuses,
        reuseRate: denom > 0 ? Math.round((reuses / denom) * 100) / 100 : null,
        estimatedTokensSaved: tokensSaved
    };
}

It never mutates memory. Both GET /analytics (app.js, novelty: NoveltyGate.stats(snapshot)) and GET /dashboard (libs/knowledge/health-dashboard.js, novelty: NoveltyGate.stats(safe)) call it on a freshly loaded snapshot — which is why the two endpoints returned byte-identical blocks.

sequenceDiagram
    participant Client
    participant App as "app.js"
    participant MM as "MemoryManager"
    participant NG as "NoveltyGate"
    Client->>App: "GET /analytics"
    App->>MM: "load()"
    MM-->>App: "snapshot (persisted state)"
    App->>NG: "stats(snapshot)"
    NG-->>App: "trackedFiles, freshEvaluations, reuses, reuseRate, estimatedTokensSaved"
    App-->>Client: "{ ..., novelty: {...}, ... }"

On by default. Off with config.NOVELTY_GATE.ENABLED === false or NOVELTY_GATE_ENABLED="false". When off, every file goes through the full pipeline exactly as before — the gate adds no other behaviour.

Unit tests live in tests/unit/novelty-gate.test.js (12 tests); executor-level tests live in tests/unit/novelty-executor.test.js (2 tests). All 14 pass on Node 24, and the full suite is 639/639:

Test What it pins
fingerprint: stable for identical inputs, changes when a signal changes Same input → same key; changing content, risk, role, confidence, status, rollbackStreak or penalty each changes it.
fingerprint: evidenceVersion alone (deps/governance) flips it Same file, new evidenceVersion → new key; missing version ≠ any version.
fingerprint: tiny confidence jitter within rounding does NOT flip it 0.701 and 0.704 → same fingerprint.
evaluate: first-seen file is always novel No memory → first-seen .
evaluate: unchanged file with a prior SKIP is reused The happy path.
evaluate: a changed fingerprint forces a fresh evaluation One-character edit → changed .
evaluate: a non-SKIP prior decision is never reused A remembered EVOLVE → prev-not-reusable .
remember: resets reuse counters when the fingerprint changes reuses /tokensSaved back to 0,evaluations keeps counting.
reuse: accumulates counters and token savings Two reuses of 50 → 100.
stats: aggregates evaluations, reuses, rate and tokens 2 files, 3 reuses → rate 0.6, 100 tokens.
stats: empty memory returns safe zero/null values reuseRate: null , notNaN .
enforceLimit: decision_memory is bounded, oldest entries drop first 505 inserts → 500 kept, f0–f4 evicted.

The key one, because it encodes the whole philosophy:

test('evaluate: a non-SKIP prior decision is never reused', () => {
    const memory = {};
    const fp = NoveltyGate.fingerprint(base);
    // EVOLVE / REJECT are not deterministic -> must re-reason even if unchanged.
    NoveltyGate.remember(memory, 'a.js', fp, { action: 'EVOLVE', reason: 'committed' });
    const r = NoveltyGate.evaluate(memory, 'a.js', fp);
    assert.equal(r.novel, true);
    assert.equal(r.reason, 'prev-not-reusable');
});

And the stats arithmetic the live numbers rely on:

test('stats: aggregates evaluations, reuses, rate and tokens', () => {
    // a.js and b.js each remembered once, then 3 reuses (40 + 40 + 20 tokens)
    const s = NoveltyGate.stats(memory);
    assert.equal(s.trackedFiles, 2);
    assert.equal(s.freshEvaluations, 2);
    assert.equal(s.reuses, 3);
    assert.equal(s.reuseRate, 0.6); // 3 / (2 + 3)
    assert.equal(s.estimatedTokensSaved, 100);
});

Integration coverage lives elsewhere: tests/unit/executor-scheduled-review.test.js pins that a due scheduled review is not swallowed by a cached novelty SKIP, and that a not-due file uses the novelty cache exactly as before.

The first draft of this article admitted two gaps: nothing asserted that evidenceVersion alone flips the fingerprint, and nothing tied estimatedTokensSaved to estimateTokens() through the real executor. Both are now pinned.

Test ( novelty-executor.test.js ) What it pins
executor: each reuse adds exactly estimateTokens(content.length) and stats() reports it Real Executor.processFile() , 1 fresh pass + 3 reuses →tokensSaved = 3 × (ceil(chars/4) + 400) ,reuseRate 0.75 . That's the exact arithmetic behind the production 314,893,830.
executor: a governance change (evidenceVersion) forces a fresh evaluation and resets reuse counters Same bytes, one extra governance rule → fresh evaluation, reuses /tokensSaved back to 0, cache resumes afterwards.
test('executor: a governance change (evidenceVersion) forces a fresh evaluation and resets reuse counters', async () => {
    const memory = certifiedMemory();
    await run(memory);
    assert.equal((await run(memory)).reused, true);

    // Same bytes, but the rules moved (a real identity edit keeps version/mission).
    memory.identity = { ...memory.identity, rules: [...memory.identity.rules, 'new governance rule'] };

    const afterRuleChange = await run(memory);
    assert.equal(afterRuleChange.reused, undefined, 'identical file is re-evaluated after a rule change');
    const rec = memory.decision_memory[FILE];
    assert.equal(rec.evaluations, 2);
    assert.equal(rec.reuses, 0);
    assert.equal(rec.tokensSaved, 0);
    assert.equal((await run(memory)).reused, true, 'cache works again under the new rules');
});

A confession from writing that test: my first version did memory.identity = { rules: ['new rule'] }, and the gate reused anyway. The test was wrong, not the gate. The memory layer treats an identity object without a current version as stale and replaces it with the default DNA (ensureBrainShape() in memory-manager.js). That happens before the fingerprint is computed, so my "rule change" disappeared before the gate could see it. A real governance edit keeps version, and then the gate behaves as designed. If you write governance tooling for Sentinel, keep the version, or your edit will be quietly reset.

Gaps still open: the fingerprint still sees only per-file signals plus evidenceVersion. Branch, MR, human-review and runtime signals are future extensions (SENTINEL_CONTEXT.md). No test covers webhook-driven cadence, because that is a property of the deployment, not of the gate.

Run them:

node --test tests/unit/novelty-gate.test.js tests/unit/novelty-executor.test.js

The Novelty Gate is about 150 lines of boring code (comments included), and that's the compliment. It doesn't make Sentinel smarter; it stops it from mistaking repetition for insight. In production it remembered 5 files, judged them fresh 42 times, and declined to re-judge them 188,000 times — while the much bigger saving today happens one step earlier, in the Wake Gate, where 92 % of ticks never open a file at all.

Confidence is how sure you are. Novelty is whether there's anything new to be sure about. Keep them in separate variables.

The repository behind Sentinel is private, so I can't link to the evidence files. Below is the captured output, excerpted field-for-field from the exact responses. The production instance was 4cd95249, running commit 7df2969, at https://autodoc-sentinel.onrender.com. The endpoints are public and read-only, but the numbers keep moving, so a read today will differ.

GET /analytics — 2026-10-11 08:24 UTC (fields excerpted from the full response):

{
  "instanceId": "4cd95249",
  "novelty": {
    "trackedFiles": 5,
    "freshEvaluations": 42,
    "reuses": 188000,
    "reuseRate": 1,
    "estimatedTokensSaved": 314893830
  },
  "wake": {
    "ticks": 14613,
    "wakes": 1104,
    "noops": 13509,
    "noopRate": 0.9245
  },
  "summary": {
    "total": 141,
    "llmConsulted": 40,
    "llmSkipped": 101,
    "estimatedTokensSpent": 204059,
    "estimatedTokensSaved": 17136
  }
}

GET /analytics — 2026-10-11 08:26:55 UTC (fields excerpted from the full response):

{
  "instanceId": "4cd95249",
  "novelty": {
    "trackedFiles": 5,
    "freshEvaluations": 42,
    "reuses": 188000,
    "reuseRate": 1,
    "estimatedTokensSaved": 314893830
  },
  "wake": {
    "ticks": 14614,
    "wakes": 1104,
    "noops": 13510,
    "noopRate": 0.9245
  },
  "summary": {
    "total": 141,
    "llmConsulted": 40,
    "llmSkipped": 101,
    "estimatedTokensSpent": 204059,
    "estimatedTokensSaved": 17136
  }
}

GET /analytics — 2026-10-11 08:48:52 UTC (fields excerpted from the full response):

{
  "instanceId": "4cd95249",
  "novelty": {
    "trackedFiles": 5,
    "freshEvaluations": 42,
    "reuses": 188000,
    "reuseRate": 1,
    "estimatedTokensSaved": 314893830
  },
  "wake": {
    "ticks": 14625,
    "wakes": 1104,
    "noops": 13521,
    "noopRate": 0.9245
  },
  "summary": {
    "total": 141,
    "llmConsulted": 40,
    "llmSkipped": 101,
    "estimatedTokensSpent": 204059,
    "estimatedTokensSaved": 17136
  }
}

GET /dashboard — 2026-10-11 08:24 UTC, novelty block:

{
  "instanceId": "4cd95249",
  "novelty": {
    "trackedFiles": 5,
    "freshEvaluations": 42,
    "reuses": 188000,
    "reuseRate": 1,
    "estimatedTokensSaved": 314893830
  }
}

GET /observations — 2026-10-11 08:46 UTC. Stats and all 20 retained resolved windows. Timestamps are converted from epoch ms to UTC, all other fields are as returned:

{
  "instanceId": "4cd95249",
  "stats": {"watching": 0, "dueForReeval": 0, "labelledSkips": 1678193, "byCategory": {"NO_VALUE": 1506470, "INSUFFICIENT_DATA": 0, "OBSERVE": 171723}, "predictions": {"resolved": 46, "correct": 45, "wrong": 1, "accuracy": 0.98}},
  "resolved": [
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-24T23:43Z", "to": "2026-07-25T05:43Z", "observations": 3322, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-25T05:43Z", "to": "2026-07-25T11:43Z", "observations": 3586, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-25T11:43Z", "to": "2026-07-25T17:43Z", "observations": 4152, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-25T17:43Z", "to": "2026-07-25T23:43Z", "observations": 4265, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-25T23:43Z", "to": "2026-07-26T05:43Z", "observations": 4613, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-26T05:43Z", "to": "2026-07-26T11:43Z", "observations": 4471, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-26T11:43Z", "to": "2026-07-26T17:43Z", "observations": 4422, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-26T17:43Z", "to": "2026-07-26T23:43Z", "observations": 4350, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-26T23:43Z", "to": "2026-07-27T05:43Z", "observations": 4261, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-27T05:43Z", "to": "2026-07-27T11:43Z", "observations": 4011, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-27T11:43Z", "to": "2026-07-27T17:43Z", "observations": 3789, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-27T17:43Z", "to": "2026-07-27T23:43Z", "observations": 3952, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-27T23:43Z", "to": "2026-07-28T05:44Z", "observations": 4063, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-28T05:44Z", "to": "2026-07-28T11:44Z", "observations": 3795, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-28T11:44Z", "to": "2026-07-28T17:44Z", "observations": 3421, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-28T17:44Z", "to": "2026-07-28T23:44Z", "observations": 3855, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-28T23:44Z", "to": "2026-07-29T05:44Z", "observations": 4133, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-29T05:44Z", "to": "2026-07-29T11:44Z", "observations": 4160, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-29T11:44Z", "to": "2026-07-29T17:44Z", "observations": 3801, "reason": "window-elapsed"},
    {"path": "libs/shared/translator-gateway.js", "from": "2026-07-29T17:44Z", "to": "2026-07-29T20:05Z", "observations": 1589, "reason": "incident-detected"}
  ]
}

Repo maintainers can rerun the same checks with one command. It runs the 14 Novelty Gate tests and prints these figures from the committed raw files (docs/articles/evidence-2026-10-11/):

node --test tests/unit/novelty-gate.test.js tests/unit/novelty-executor.test.js && node -e 'const r=f=>require("./docs/articles/evidence-2026-10-11/"+f);for(const f of ["analytics-0824Z.json","analytics-0826Z.json","analytics-0848Z.json"]){const d=r(f);console.log(f,d.instanceId,JSON.stringify(d.novelty),"ticks",d.wake.ticks,"noops",d.wake.noops,"wakes",d.wake.wakes,"decisions",d.summary.total,"spent",d.summary.estimatedTokensSpent)}const o=r("observations-0846Z.json");console.log("labelledSkips",o.stats.labelledSkips);for(const w of o.resolved)console.log(w.path,w.observations,((w.resolvedTs-w.firstTs)/1000/w.observations).toFixed(1)+" s/obs");console.log("dashboard",JSON.stringify(r("dashboard-0824Z.json").novelty))'

npm test runs the whole suite: 639 tests at the time of writing.

── more in #ai-agents 4 stories · sorted by recency
── more on @sentinel 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/novelty-gate-sentine…] indexed:0 read:21min 2026-10-11 · —