{"slug": "novelty-gate-sentinel-confidence-is-not-novelty", "title": "Novelty Gate Sentinel— confidence is not novelty.", "summary": "A developer built Sentinel, an autonomous code agent that watches a JavaScript codebase and separates confidence from novelty via a \"Novelty Gate\" that fingerprints each file's content, risk, role and rules so unchanged files reuse a prior SKIP decision instead of re-reasoning. Live analytics from the production instance on Render, read 2026-10-11, showed 188,000 reuses against 42 fresh evaluations across 5 tracked files — roughly 4,476 reuses per fresh evaluation — with an estimated 314,893,830 tokens not re-read, a figure the developer cautions is a rough size estimate rather than actual billed tokens.", "body_md": "*How an autonomous code agent learned to say \"I've seen this exact file before, nothing moved, my answer is still no\" — and what that looks like after three months in production.*\n\n**Live numbers in this article** come from the production Sentinel instance on Render, read from the read-only `GET /analytics` endpoint on **2026-10-11 08:24 UTC** (instance `4cd95249`), then re-read at 08:26:55 UTC with identical results. The captured raw output is reproduced at the end of the article, in *Captured output*. Nothing here is a projection or a synthetic benchmark. Where a number is an *estimate*, I say so and show the formula.\n\nSentinel is an autonomous agent that watches a JavaScript codebase. Every couple of minutes it wakes up, looks at a batch of files, and decides one of three things for each file:\n\nMost of the time the answer is SKIP. That is good. A security-focused agent that changes code on every pass is not an agent, it is an incident.\n\nThe trouble was *how* it got to SKIP. Every pass it re-read the file, rebuilt its picture of the file's risk and role, re-ran its decision policy, and — surprise — arrived at the same SKIP as two minutes earlier. Same file, same risk, same role, same rules, same answer. Over and over.\n\nIf you have ever re-opened the fridge hoping something new appeared since thirty seconds ago, you already understand the bug.\n\nThis is the core idea of the article, so here it is in one line:\n\n**Confidence answers \"how reliable is my judgement?\". Novelty answers \"has anything changed since I last judged?\". They are different questions.**\n\nIt is tempting to treat repeated agreement as growing certainty: \"I said SKIP ten times in a row, so I must be really sure.\" That is a trap. Ten identical inputs give you the same evidence ten times, not ten pieces of evidence. If confidence drifted upward just because a judgement was repeated, the agent would slowly talk itself into trusting a module it has learned nothing new about.\n\nSo Sentinel keeps the two apart:\n\nIf it is the same situation, and the last answer was a plain SKIP, Sentinel simply reuses that answer. No re-reasoning, no new memories, no LLM. That check is the **Novelty Gate**.\n\nNot just \"the file's bytes are identical\". The gate takes a fingerprint of everything that should force a fresh look:\n\nChange any one of these and the fingerprint changes, so the file gets a full, fresh evaluation.\n\nThe gate is deliberately timid:\n\n`SENTINEL_CONTEXT.md`, line 104). So \"unchanged\" means \"unchanged Here is the `novelty` block exactly as the live instance returned it:\n\n```\n{\n  \"trackedFiles\": 5,\n  \"freshEvaluations\": 42,\n  \"reuses\": 188000,\n  \"reuseRate\": 1,\n  \"estimatedTokensSaved\": 314893830\n}\n```\n\nIn plain words:\n\n| Field | Value | What it means | \n|---|---|---|\n| `trackedFiles` | **5** | Files the gate currently remembers a decision for. Small, because Sentinel's scanner looks at one directory ( `libs/shared` , 7 files) and processes at most 5 per pass. | \n| `freshEvaluations` | **42** | Times one of those files went through the full decision pipeline and ended in a SKIP that was remembered. | \n| `reuses` | **188,000** | Times a remembered SKIP was reused instead of re-reasoning. | \n| `reuseRate` | **1** (really 0.99978) | `reuses / (reuses + freshEvaluations)` , rounded to two decimals — so it*displays* as 1. It is not literally 100 %. | \n| `estimatedTokensSaved` | **314,893,830** | An *estimate* of how much input Sentinel did not have to re-read. Read the next section before quoting this anywhere. | \n\nThat is roughly **4,476 reuses for every fresh evaluation**. For a five-file working set that mostly doesn't change, that is what you would expect: once a stable file has been judged, the gate keeps saying \"same as before\" until something moves.\n\nThat number is real in the sense that the endpoint returns it. It is **not** \"315 million OpenAI tokens we would otherwise have paid for\". Three reasons:\n\n`ceil(chars / 4) + 400` \"tokens\" — a rough size of the file as an LLM prompt plus a fixed overhead. Nobody counted actual tokens.\nThe honest summary: **the Novelty Gate turned an agent that re-thought the same five files on every pass (every two minutes by design, and for a stretch in July every ~5 seconds, see Part 2) into one that re-thinks them only when something relevant changes.** The token estimate is a size-of-avoided-work indicator, not a bill.\n\nIf you only came for the idea, you can stop here. The rest is code.\n\nThe gate lives in `libs/core/novelty-gate.js` and is called from `Executor.processFile` (`libs/core/executor.js`), right after the decision envelope is built and **before** the compact IR, reference assessment, decision policy and any LLM call.\n\n``` php\nflowchart TD\n    A[\"mainLoop tick\"] --> B{\"Wake Gate: actionable event?\"}\n    B -- \"no\" --> Z[\"no-op, nothing scanned\"]\n    B -- \"yes\" --> C[\"RepositoryScanner: libs/shared, batch of 5\"]\n    C --> D[\"processFile: build envelope\"]\n    D --> E[\"NoveltyGate.fingerprint\"]\n    E --> F{\"scheduled review due?\"}\n    F -- \"yes\" --> H[\"full pipeline\"]\n    F -- \"no\" --> G{\"NoveltyGate.evaluate\"}\n    G -- \"novel\" --> H\n    G -- \"unchanged and prev SKIP\" --> R{\"autonomy token reconsiders?\"}\n    R -- \"yes\" --> H\n    R -- \"no\" --> S[\"reuse previous SKIP, count tokens, return reused:true\"]\n    H --> P[\"IR, policy, maybe LLM, gates\"]\n    P --> Q{\"outcome\"}\n    Q -- \"SKIP\" --> M[\"NoveltyGate.remember\"]\n    Q -- \"EVOLVE / REJECT / DEFERRED\" --> N[\"not remembered\"]\n```\n\nNote the two escape hatches before reuse: a due scheduled review (added after a real bug where a cached SKIP swallowed a review, pinned in `tests/unit/executor-scheduled-review.test.js`) and an autonomy token (` reconsider(\"novelty-reuse\")`).\n\n`libs/core/novelty-gate.js`, lines 35–69:\n\n```\nfunction fingerprint(input = {}) {\n    const {\n        content,\n        risk = null,\n        role = null,\n        confidence = null,\n        status = null,\n        rollbackStreak = 0,\n        penalty = 0,\n        evidenceVersion = null\n    } = input;\n\n    const contentHash = crypto\n        .createHash(\"sha256\")\n        .update(String(content ?? \"\"))\n        .digest(\"hex\")\n        .slice(0, 16);\n\n    const conf = Number.isFinite(Number(confidence))\n        ? Math.round(Number(confidence) * 100) / 100\n        : \"na\";\n\n    return [\n        contentHash,\n        risk ?? \"na\",\n        role ?? \"na\",\n        conf,\n        status ?? \"na\",\n        Number(rollbackStreak) || 0,\n        Number(penalty) || 0,\n        evidenceVersion ?? \"na\"\n    ].join(\"|\");\n}\n```\n\nA few deliberate choices:\n\n`0.701` and `0.704` produce the same fingerprint (there is a test for exactly this). Floating-point jitter is not novelty. A move from `0.70` to `0.71` `evidenceVersion` is the world outside the file.`OpportunityEngine.evidenceVersion(memory, { dependencies: readDependencyManifest() })`. Without it, a frozen module would reuse its old SKIP forever, even after the governance rules or dependencies moved — and the Opportunity Engine could never reach exactly the modules it exists to re-examine. Small aside: the fingerprint description in `console.log`. When your cache key is human-readable, \"why did this miss?\" takes seconds instead of an afternoon.\nWhat is *not* in it, by design for now: branch/MR state, human confirmation, production runtime signals. Those are external observers, and adding them is listed as a future extension in `SENTINEL_CONTEXT.md:104`.\n\n```\n// Only a prior SKIP is deterministic enough to reuse.\nfunction reusable(decision) {\n    return !!decision &&\n        decision.action === \"SKIP\" &&\n        typeof decision.reason === \"string\";\n}\n\nfunction evaluate(memory, filePath, fp) {\n    const prev = store(memory)[filePath];\n    if (!prev) return { novel: true, reason: \"first-seen\", previous: null };\n    if (prev.fingerprint !== fp) return { novel: true, reason: \"changed\", previous: prev };\n    if (!reusable(prev.decision)) return { novel: true, reason: \"prev-not-reusable\", previous: prev };\n    return { novel: false, reason: \"unchanged\", previous: prev };\n}\n```\n\n`remember()` is only called on the three SKIP return paths in the executor — lifecycle skip, local \"LLM not needed\" skip, and intent-gate skip. Its bookkeeping per file:\n\n```\ns[filePath] = {\n    fingerprint: fp,\n    decision: { action: decision.action, reason: decision.reason ?? null },\n    firstTs: prev?.firstTs ?? now,\n    lastTs: now,\n    evaluations: (prev?.evaluations ?? 0) + 1,\n    reuses: sameFp ? (prev.reuses ?? 0) : 0,\n    tokensSaved: sameFp ? (prev.tokensSaved ?? 0) : 0\n};\n```\n\nRead the last two lines twice, because they matter for interpreting the production numbers: **when the fingerprint changes, `reuses` and `tokensSaved` reset to zero, but `evaluations` keeps counting.** So the live `reuses: 188000` is \"reuses since each file's *current* fingerprint was established\", not an all-time total. If anything, it under-counts lifetime reuse.\n\nOn the hot path, the executor does this when the fingerprint is unchanged (`libs/core/executor.js`, ~lines 308–329):\n\n``` js\nconst saved = DecisionAnalytics.estimateTokens(\n    typeof content === \"string\" ? content.length : 0\n);\nNoveltyGate.reuse(memory, filePath, saved);\n// ... ObservationMode.observe(...) labels the SKIP, never changes it\nMemoryManager.clearWorking(memory);\nreturn {\n    action: novelty.previous.decision.action,\n    reason: novelty.previous.decision.reason,\n    reused: true,\n    envelope\n};\n```\n\nNo decision-stream entry, no experience summary, no analytics row. That is the point: a reused decision is not a new memory.\n\nI had been describing the estimate as \"chars / 4\". The actual function is in `libs/knowledge/decision-analytics.js`, lines 18–21:\n\n``` js\nfunction estimateTokens(contentLength = 0) {\n    const body = Math.ceil(Math.max(0, Number(contentLength) || 0) / 4);\n    return body + 400;\n}\n```\n\nSo it is **`ceil(chars / 4) + 400`** — the body plus a fixed prompt overhead. It is the same estimator `DecisionAnalytics` uses for every locally-handled decision, so novelty numbers and analytics numbers are at least in the same (rough) currency.\n\nCan `estimatedTokensSaved = 314,893,830` actually come from `reuses = 188,000` under that formula?\n\n`crypto-vault.js`, 18,743 chars → 5,086): 956,168,000. ✔ below it.` master` ranged from 1,370 to 18,743 chars over the period (`translator-gateway.js` from 5,235 to 7,241). An average of ~5,100 chars across a 5-file batch is plausible for that mix.\nWhat I **could not** do: an exact per-file reconciliation. The per-file `memory.decision_memory` records (with their own `reuses`/` tokensSaved`) are not exposed by any read-only endpoint, and file sizes changed mid-period. So the claim is \"consistent with the formula and within bounds\", not \"verified to the token\".\n\nBeing evidence-first means writing these down instead of rounding them away. My first draft had a section called \"things I can't explain\", and its main item was this:\n\nThe loop runs every 2 minutes with a batch of 5. From the gate's merge (2026-07-18) to the Wake Gate (2026-09-02) is ~46 days → at most ~165,600 reuses, plus ~5,520 since (1,104 wakes × 5). Ceiling ≈ 171,000. Reported: 188,000.\n\n188,000 is above that, which should be impossible. So I went digging. The explanation is that the ceiling was wrong, not the counter.\n\n**The loop did not run every 2 minutes.** Sentinel's Observation Mode keeps a log of six-hour \"watch windows\". Each window records how many times a file was observed, and the novelty-reuse path itself calls `ObservationMode.observe()` (`executor.js`, inside the reuse branch). Here are all the retained windows for `libs/shared/translator-gateway.js`, straight from `GET /observations`:\n\n| Window (UTC) | Observations | Seconds per observation | \n|---|---|---|\n| 07-24 23:43 → 07-25 05:43 | 3,322 | 6.5 | \n| 07-25 23:43 → 07-26 05:43 | 4,613 | 4.7 | \n| 07-27 05:43 → 07-27 11:43 | 4,011 | 5.4 | \n| 07-28 11:44 → 07-28 17:44 | 3,421 | 6.3 | \n| 07-29 17:44 → 07-29 20:05 (incident) | 1,589 | 5.3 | \n\n(20 windows retained, 2026-07-24 → 07-29; all between 4.7 and 6.5 s.) A 2-minute timer would give 180 observations per window. The real number was **18–26× higher**: one pass roughly every 5 seconds.\n\n**Who called the loop that often?** There are exactly two callers of `mainLoop()` in `app.js`:\n\n``` js\napp.post('/webhook',(_req,res)=>{\n    res.status(200).send('OK');\n    mainLoop();\n});\n\nsetInterval(mainLoop, CONFIG.SYSTEM.INTERVAL); // 2 min\n```\n\nThe guard sits at the very top of `mainLoop()`, before the first `await` (`app.js`, master):\n\n```\nasync function mainLoop(){\n    if(isBusy) return;          // first statement\n    isBusy=true;                // set synchronously, no await in between\n    let memory=null;\n    let lockAcquired=false;\n    try{\n        memory=await MemoryManager.load();\n        lockAcquired=RuntimeLock.acquire(memory, INSTANCE_ID);\n        if(!lockAcquired){ return; }\n        // ... wake gate, batch, executor ...\n    }catch(err){ /* log */ }\n    finally{\n        if(lockAcquired && memory){ RuntimeLock.release(memory); }\n        isBusy=false;           // only reset here\n    }\n}\n```\n\nThe check and the set run in the same synchronous step, and Node runs one step at a time, so two passes can never overlap inside one process. `RuntimeLock` does something else: it keeps a *different* instance out (owner = `INSTANCE_ID`, 240 s TTL). Render ran a single instance anyway. The guard was in the same position in the July code (`app.js` at `3d8a0b3`, 2026-07-18).\n\nSo `POST /webhook` is the only code path that can produce passes faster than every 2 minutes. A webhook call is not queued: it calls `mainLoop()` without `await`, and if a pass is already running, the call returns immediately and nothing happens. A pass every ~5 seconds therefore means webhook calls were arriving at least that often, each one landing after the previous pass had finished.\n\n**A second, independent check.** Observation Mode has labelled **1,678,193** skips since it was merged on 2026-07-18, 85 days ago. At one pass every 2 minutes with 5 files, that would take ~466 days of non-stop running. At one pass every ~5 seconds, it takes ~19 days.\n\n**What this means for 188,000.** At ~5 s per pass and 5 files per batch, the gate could log ~86,000 reuses a day, so 188,000 is roughly two days' worth. Also, a file's `reuses` counter resets to 0 each time its fingerprint changes. The live number is therefore \"reuses since each file's last fingerprint change\", not a lifetime total. That makes it smaller than the lifetime count, not larger.\n\n**What I still can't prove.** *Who* sent those webhook calls. The code explains how the cadence was possible, and the observation log shows it happened. Render logs from July are past retention, and I had no access to the GitLab project's webhook settings. So the sender (most likely a GitLab project hook firing on Sentinel's own activity, but that's a guess) stays unconfirmed. Today it no longer matters much for cost: since 2026-09-02 every pass goes through the Wake Gate first, so a webhook-triggered pass with nothing new is a cheap no-op. The live counter shows 14,625 ticks, 92.45 % of them no-ops.\n\nTwo smaller notes:\n\n`/analytics` reads and one `/dashboard` read, 08:24–08:48 UTC) and it didn't move. That's expected: since the Wake Gate, a no-op tick does not open files, so it does not reuse anything. A round number is a coincidence, not a sign of a cap; `novelty-gate.js` has no cap on `reuses`.` reuseRate` rounds to 1.`stats()` rounds to two decimals; the exact value is 0.999777.\nNone of this changes the qualitative result, but it changes the honest reading. The big counter partly measures **a loop that ran far more often than designed**. Some of the \"saved\" work is work that only existed because the loop was over-triggered. The Novelty Gate made that over-triggering cheap; the Wake Gate later made it pointless.\n\n``` js\nconst STORE_LIMIT = 500;\n\nfunction enforceLimit(s, limit = STORE_LIMIT) {\n    const keys = Object.keys(s);\n    if (keys.length <= limit) return;\n    keys\n        .map(k => [k, Number(s[k]?.lastTs) || 0])\n        .sort((a, b) => a[1] - b[1])\n        .slice(0, keys.length - limit)\n        .forEach(([k]) => { delete s[k]; });\n}\n```\n\nAt most 500 files are remembered; the least-recently touched ones are evicted first. Today the instance tracks 5, so this limit is a seatbelt, not a feature in use. Eviction is safe by construction: a forgotten file is just \"first-seen\" next time and gets a fresh evaluation.\n\n``` js\nfunction stats(memory) {\n    const s = (memory && typeof memory.decision_memory === \"object\")\n        ? memory.decision_memory\n        : {};\n    const files = Object.values(s);\n\n    const freshEvaluations = files.reduce((n, r) => n + (Number(r.evaluations) || 0), 0);\n    const reuses = files.reduce((n, r) => n + (Number(r.reuses) || 0), 0);\n    const tokensSaved = files.reduce((n, r) => n + (Number(r.tokensSaved) || 0), 0);\n    const denom = freshEvaluations + reuses;\n\n    return {\n        trackedFiles: files.length,\n        freshEvaluations,\n        reuses,\n        reuseRate: denom > 0 ? Math.round((reuses / denom) * 100) / 100 : null,\n        estimatedTokensSaved: tokensSaved\n    };\n}\n```\n\nIt never mutates memory. Both `GET /analytics` (`app.js`, `novelty: NoveltyGate.stats(snapshot)`) and `GET /dashboard` (`libs/knowledge/health-dashboard.js`, `novelty: NoveltyGate.stats(safe)`) call it on a freshly loaded snapshot — which is why the two endpoints returned byte-identical blocks.\n\n```\nsequenceDiagram\n    participant Client\n    participant App as \"app.js\"\n    participant MM as \"MemoryManager\"\n    participant NG as \"NoveltyGate\"\n    Client->>App: \"GET /analytics\"\n    App->>MM: \"load()\"\n    MM-->>App: \"snapshot (persisted state)\"\n    App->>NG: \"stats(snapshot)\"\n    NG-->>App: \"trackedFiles, freshEvaluations, reuses, reuseRate, estimatedTokensSaved\"\n    App-->>Client: \"{ ..., novelty: {...}, ... }\"\n```\n\nOn by default. Off with `config.NOVELTY_GATE.ENABLED === false` or `NOVELTY_GATE_ENABLED=\"false\"`. When off, every file goes through the full pipeline exactly as before — the gate adds no other behaviour.\n\nUnit tests live in `tests/unit/novelty-gate.test.js` (12 tests); executor-level tests live in `tests/unit/novelty-executor.test.js` (2 tests). All 14 pass on Node 24, and the full suite is 639/639:\n\n| Test | What it pins | \n|---|---|\n| `fingerprint: stable for identical inputs, changes when a signal changes` | Same input → same key; changing content, risk, role, confidence, status, rollbackStreak or penalty each changes it. | \n| `fingerprint: evidenceVersion alone (deps/governance) flips it` | Same file, new `evidenceVersion` → new key; missing version ≠ any version. | \n| `fingerprint: tiny confidence jitter within rounding does NOT flip it` | 0.701 and 0.704 → same fingerprint. | \n| `evaluate: first-seen file is always novel` | No memory → `first-seen` . | \n| `evaluate: unchanged file with a prior SKIP is reused` | The happy path. | \n| `evaluate: a changed fingerprint forces a fresh evaluation` | One-character edit → `changed` . | \n| `evaluate: a non-SKIP prior decision is never reused` | A remembered EVOLVE → `prev-not-reusable` . | \n| `remember: resets reuse counters when the fingerprint changes` | `reuses` /`tokensSaved` back to 0,`evaluations` keeps counting. | \n| `reuse: accumulates counters and token savings` | Two reuses of 50 → 100. | \n| `stats: aggregates evaluations, reuses, rate and tokens` | 2 files, 3 reuses → rate 0.6, 100 tokens. | \n| `stats: empty memory returns safe zero/null values` | `reuseRate: null` , not`NaN` . | \n| `enforceLimit: decision_memory is bounded, oldest entries drop first` | 505 inserts → 500 kept, f0–f4 evicted. | \n\nThe key one, because it encodes the whole philosophy:\n\n``` js\ntest('evaluate: a non-SKIP prior decision is never reused', () => {\n    const memory = {};\n    const fp = NoveltyGate.fingerprint(base);\n    // EVOLVE / REJECT are not deterministic -> must re-reason even if unchanged.\n    NoveltyGate.remember(memory, 'a.js', fp, { action: 'EVOLVE', reason: 'committed' });\n    const r = NoveltyGate.evaluate(memory, 'a.js', fp);\n    assert.equal(r.novel, true);\n    assert.equal(r.reason, 'prev-not-reusable');\n});\n```\n\nAnd the stats arithmetic the live numbers rely on:\n\n``` js\ntest('stats: aggregates evaluations, reuses, rate and tokens', () => {\n    // a.js and b.js each remembered once, then 3 reuses (40 + 40 + 20 tokens)\n    const s = NoveltyGate.stats(memory);\n    assert.equal(s.trackedFiles, 2);\n    assert.equal(s.freshEvaluations, 2);\n    assert.equal(s.reuses, 3);\n    assert.equal(s.reuseRate, 0.6); // 3 / (2 + 3)\n    assert.equal(s.estimatedTokensSaved, 100);\n});\n```\n\nIntegration coverage lives elsewhere: `tests/unit/executor-scheduled-review.test.js` pins that *a due scheduled review is not swallowed by a cached novelty SKIP*, and that a not-due file uses the novelty cache exactly as before.\n\nThe first draft of this article admitted two gaps: nothing asserted that `evidenceVersion` alone flips the fingerprint, and nothing tied `estimatedTokensSaved` to `estimateTokens()` through the real executor. Both are now pinned.\n\n| Test ( `novelty-executor.test.js` ) | What it pins | \n|---|---|\n| `executor: each reuse adds exactly estimateTokens(content.length) and stats() reports it` | Real `Executor.processFile()` , 1 fresh pass + 3 reuses →`tokensSaved = 3 × (ceil(chars/4) + 400)` ,`reuseRate 0.75` . That's the exact arithmetic behind the production 314,893,830. | \n| `executor: a governance change (evidenceVersion) forces a fresh evaluation and resets reuse counters` | Same bytes, one extra governance rule → fresh evaluation, `reuses` /`tokensSaved` back to 0, cache resumes afterwards. | \n\n```\ntest('executor: a governance change (evidenceVersion) forces a fresh evaluation and resets reuse counters', async () => {\n    const memory = certifiedMemory();\n    await run(memory);\n    assert.equal((await run(memory)).reused, true);\n\n    // Same bytes, but the rules moved (a real identity edit keeps version/mission).\n    memory.identity = { ...memory.identity, rules: [...memory.identity.rules, 'new governance rule'] };\n\n    const afterRuleChange = await run(memory);\n    assert.equal(afterRuleChange.reused, undefined, 'identical file is re-evaluated after a rule change');\n    const rec = memory.decision_memory[FILE];\n    assert.equal(rec.evaluations, 2);\n    assert.equal(rec.reuses, 0);\n    assert.equal(rec.tokensSaved, 0);\n    assert.equal((await run(memory)).reused, true, 'cache works again under the new rules');\n});\n```\n\nA confession from writing that test: my first version did `memory.identity = { rules: ['new rule'] }`, and the gate *reused* anyway. The test was wrong, not the gate. The memory layer treats an identity object without a current `version` as stale and replaces it with the default DNA (`ensureBrainShape()` in `memory-manager.js`). That happens before the fingerprint is computed, so my \"rule change\" disappeared before the gate could see it. A real governance edit keeps `version`, and then the gate behaves as designed. If you write governance tooling for Sentinel, keep the version, or your edit will be quietly reset.\n\n**Gaps still open:** the fingerprint still sees only per-file signals plus `evidenceVersion`. Branch, MR, human-review and runtime signals are future extensions (`SENTINEL_CONTEXT.md`). No test covers webhook-driven cadence, because that is a property of the deployment, not of the gate.\n\nRun them:\n\n```\nnode --test tests/unit/novelty-gate.test.js tests/unit/novelty-executor.test.js\n```\n\nThe Novelty Gate is about 150 lines of boring code (comments included), and that's the compliment. It doesn't make Sentinel smarter; it stops it from mistaking repetition for insight. In production it remembered 5 files, judged them fresh 42 times, and declined to re-judge them 188,000 times — while the much bigger saving today happens one step earlier, in the Wake Gate, where 92 % of ticks never open a file at all.\n\nConfidence is how sure you are. Novelty is whether there's anything new to be sure about. Keep them in separate variables.\n\nThe repository behind Sentinel is private, so I can't link to the evidence files. Below is the captured output, excerpted field-for-field from the exact responses. The production instance was `4cd95249`, running commit `7df2969`, at `https://autodoc-sentinel.onrender.com`. The endpoints are public and read-only, but the numbers keep moving, so a read today will differ.\n\n**GET /analytics — 2026-10-11 08:24 UTC** (fields excerpted from the full response):\n\n```\n{\n  \"instanceId\": \"4cd95249\",\n  \"novelty\": {\n    \"trackedFiles\": 5,\n    \"freshEvaluations\": 42,\n    \"reuses\": 188000,\n    \"reuseRate\": 1,\n    \"estimatedTokensSaved\": 314893830\n  },\n  \"wake\": {\n    \"ticks\": 14613,\n    \"wakes\": 1104,\n    \"noops\": 13509,\n    \"noopRate\": 0.9245\n  },\n  \"summary\": {\n    \"total\": 141,\n    \"llmConsulted\": 40,\n    \"llmSkipped\": 101,\n    \"estimatedTokensSpent\": 204059,\n    \"estimatedTokensSaved\": 17136\n  }\n}\n```\n\n**GET /analytics — 2026-10-11 08:26:55 UTC** (fields excerpted from the full response):\n\n```\n{\n  \"instanceId\": \"4cd95249\",\n  \"novelty\": {\n    \"trackedFiles\": 5,\n    \"freshEvaluations\": 42,\n    \"reuses\": 188000,\n    \"reuseRate\": 1,\n    \"estimatedTokensSaved\": 314893830\n  },\n  \"wake\": {\n    \"ticks\": 14614,\n    \"wakes\": 1104,\n    \"noops\": 13510,\n    \"noopRate\": 0.9245\n  },\n  \"summary\": {\n    \"total\": 141,\n    \"llmConsulted\": 40,\n    \"llmSkipped\": 101,\n    \"estimatedTokensSpent\": 204059,\n    \"estimatedTokensSaved\": 17136\n  }\n}\n```\n\n**GET /analytics — 2026-10-11 08:48:52 UTC** (fields excerpted from the full response):\n\n```\n{\n  \"instanceId\": \"4cd95249\",\n  \"novelty\": {\n    \"trackedFiles\": 5,\n    \"freshEvaluations\": 42,\n    \"reuses\": 188000,\n    \"reuseRate\": 1,\n    \"estimatedTokensSaved\": 314893830\n  },\n  \"wake\": {\n    \"ticks\": 14625,\n    \"wakes\": 1104,\n    \"noops\": 13521,\n    \"noopRate\": 0.9245\n  },\n  \"summary\": {\n    \"total\": 141,\n    \"llmConsulted\": 40,\n    \"llmSkipped\": 101,\n    \"estimatedTokensSpent\": 204059,\n    \"estimatedTokensSaved\": 17136\n  }\n}\n```\n\n**GET /dashboard — 2026-10-11 08:24 UTC**, `novelty` block:\n\n```\n{\n  \"instanceId\": \"4cd95249\",\n  \"novelty\": {\n    \"trackedFiles\": 5,\n    \"freshEvaluations\": 42,\n    \"reuses\": 188000,\n    \"reuseRate\": 1,\n    \"estimatedTokensSaved\": 314893830\n  }\n}\n```\n\n**GET /observations — 2026-10-11 08:46 UTC**. Stats and all 20 retained resolved windows. Timestamps are converted from epoch ms to UTC, all other fields are as returned:\n\n```\n{\n  \"instanceId\": \"4cd95249\",\n  \"stats\": {\"watching\": 0, \"dueForReeval\": 0, \"labelledSkips\": 1678193, \"byCategory\": {\"NO_VALUE\": 1506470, \"INSUFFICIENT_DATA\": 0, \"OBSERVE\": 171723}, \"predictions\": {\"resolved\": 46, \"correct\": 45, \"wrong\": 1, \"accuracy\": 0.98}},\n  \"resolved\": [\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-24T23:43Z\", \"to\": \"2026-07-25T05:43Z\", \"observations\": 3322, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-25T05:43Z\", \"to\": \"2026-07-25T11:43Z\", \"observations\": 3586, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-25T11:43Z\", \"to\": \"2026-07-25T17:43Z\", \"observations\": 4152, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-25T17:43Z\", \"to\": \"2026-07-25T23:43Z\", \"observations\": 4265, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-25T23:43Z\", \"to\": \"2026-07-26T05:43Z\", \"observations\": 4613, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-26T05:43Z\", \"to\": \"2026-07-26T11:43Z\", \"observations\": 4471, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-26T11:43Z\", \"to\": \"2026-07-26T17:43Z\", \"observations\": 4422, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-26T17:43Z\", \"to\": \"2026-07-26T23:43Z\", \"observations\": 4350, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-26T23:43Z\", \"to\": \"2026-07-27T05:43Z\", \"observations\": 4261, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-27T05:43Z\", \"to\": \"2026-07-27T11:43Z\", \"observations\": 4011, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-27T11:43Z\", \"to\": \"2026-07-27T17:43Z\", \"observations\": 3789, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-27T17:43Z\", \"to\": \"2026-07-27T23:43Z\", \"observations\": 3952, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-27T23:43Z\", \"to\": \"2026-07-28T05:44Z\", \"observations\": 4063, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-28T05:44Z\", \"to\": \"2026-07-28T11:44Z\", \"observations\": 3795, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-28T11:44Z\", \"to\": \"2026-07-28T17:44Z\", \"observations\": 3421, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-28T17:44Z\", \"to\": \"2026-07-28T23:44Z\", \"observations\": 3855, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-28T23:44Z\", \"to\": \"2026-07-29T05:44Z\", \"observations\": 4133, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-29T05:44Z\", \"to\": \"2026-07-29T11:44Z\", \"observations\": 4160, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-29T11:44Z\", \"to\": \"2026-07-29T17:44Z\", \"observations\": 3801, \"reason\": \"window-elapsed\"},\n    {\"path\": \"libs/shared/translator-gateway.js\", \"from\": \"2026-07-29T17:44Z\", \"to\": \"2026-07-29T20:05Z\", \"observations\": 1589, \"reason\": \"incident-detected\"}\n  ]\n}\n```\n\nRepo maintainers can rerun the same checks with one command. It runs the 14 Novelty Gate tests and prints these figures from the committed raw files (`docs/articles/evidence-2026-10-11/`):\n\n```\nnode --test tests/unit/novelty-gate.test.js tests/unit/novelty-executor.test.js && node -e 'const r=f=>require(\"./docs/articles/evidence-2026-10-11/\"+f);for(const f of [\"analytics-0824Z.json\",\"analytics-0826Z.json\",\"analytics-0848Z.json\"]){const d=r(f);console.log(f,d.instanceId,JSON.stringify(d.novelty),\"ticks\",d.wake.ticks,\"noops\",d.wake.noops,\"wakes\",d.wake.wakes,\"decisions\",d.summary.total,\"spent\",d.summary.estimatedTokensSpent)}const o=r(\"observations-0846Z.json\");console.log(\"labelledSkips\",o.stats.labelledSkips);for(const w of o.resolved)console.log(w.path,w.observations,((w.resolvedTs-w.firstTs)/1000/w.observations).toFixed(1)+\" s/obs\");console.log(\"dashboard\",JSON.stringify(r(\"dashboard-0824Z.json\").novelty))'\n```\n\n`npm test` runs the whole suite: 639 tests at the time of writing.", "url": "https://wpnews.pro/news/novelty-gate-sentinel-confidence-is-not-novelty", "canonical_source": "https://dev.to/jackymencz/novelty-gate-sentinel-confidence-is-not-novelty-1e94", "published_at": "2026-10-11 10:08:48+00:00", "updated_at": "2026-10-11 10:21:39.629869+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["Sentinel", "Render"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/novelty-gate-sentinel-confidence-is-not-novelty", "markdown": "https://wpnews.pro/news/novelty-gate-sentinel-confidence-is-not-novelty.md", "text": "https://wpnews.pro/news/novelty-gate-sentinel-confidence-is-not-novelty.txt", "jsonld": "https://wpnews.pro/news/novelty-gate-sentinel-confidence-is-not-novelty.jsonld"}}