By Team ThinkMates β Amrutha K, Kammar Akshay, Rashmika K
Sealer-02 is a heat-sealing machine on a packaging line. For months it ran Film-A from supplier PackCo under recipe R10, and the fix for any Weak Seal defect was reliable: Increase Temperature +5Β°C β four successes, zero failures.
Then production changed to Film-B from FlexPack under recipe R11. The Weak Seal defect returned β and the remembered fix, four successes and zero failures behind it, failed twice.
Most AI memory answers one question well: what worked before?
They store experiences, retrieve them by similarity, and surface the past fix.
That works until the world moves.
"Temperature +5Β°C worked 4 out of 4 times" is technically correct β and operationally dangerous β when those times happened under a context that no longer exists.
A fix is never valid in the abstract.
It is valid for a defect, under a context:
Change the film and the seal physics change with it.
The old memory is still true β the fix really did work under Film-A/R10.
But truth and current validity differ, and a system that cannot tell them apart may recommend a fix that the new reality has already contradicted.
The fix did not become false. Its validity boundary changed.
Validrift is change-aware manufacturing memory: persistent Hindsight memory plus a deterministic context-validity engine deciding whether learned knowledge should still be trusted.
Hindsight remembers the manufacturing history.
The Validrift engine decides whether that knowledge remains valid under the current production context.
Its main ideas are:
Nothing is deleted or globally overwritten just because the production context changed.
Under:
the corrective action Increase Temperature +5Β°C had:
4 successes / 0 failures
Status:
VALIDATED
Validrift preserves this as historical evidence.
It happened, it worked, and that historical fact remains true.
Production later changed:
Film-A β Film-B
PackCo β FlexPack
R10 β R11
Validrift records process changes as first-class events in its structured ledger and in Hindsight.
A process change tells the system that previously learned knowledge may need to be re-evaluated.
Under Film-B/R11, the same fix failed twice:
0 successes / 2 failures
DRIFTED
The old record is not erased.
The Fix Passport can show both facts at the same time:
Film-A/R10
Temperature +5Β°C
4 successes / 0 failures
VALIDATED
Film-B/R11
Temperature +5Β°C
0 successes / 2 failures
DRIFTED
The memory wasn't wrong.
Its validity drifted.
Validrift's central idea is that validity is a function of:
(fix, defect, context)
βnot of the fix alone.
Every intervention record carries the production context.
For a new incident, the deterministic engine separates evidence into:
It then evaluates each fix using one of these statuses:
This lets the same corrective action have different validity states under different production conditions.
A Fix Passport is the biography of one corrective action.
It contains:
If a quality engineer asks:
"Why not Temperature +5Β°C?"
Validrift can answer with both halves of the truth:
Film-A/R10:
4/0 β VALIDATED
Film-B/R11:
0/2 β DRIFTED
There is no global overwrite and no silent forgetting.
A process change can trigger a Memory Validity Audit.
Known fixes are re-evaluated against the current context.
For example:
Instead of treating every remembered fix as permanently trustworthy, Validrift continuously asks:
Does this knowledge still apply here?
Manufacturing events are retained into the Hindsight memory bank.
These can include:
HindsightService.retain uses stable document IDs and records a MemoryTrace row in SQLite.
This makes memory operations auditable with:
When a new Weak Seal incident occurs on Sealer-02 under Film-B/R11, Validrift recalls relevant manufacturing experience from Hindsight.
This can include:
During the verified live run, the recommendation flow received 49 real recalled memory IDs.
The important distinction is:
Recall supplies evidence. It does not supply the final validity verdict.
That decision belongs to Validrift's deterministic validity engine.
Hindsight REFLECT is used to generate higher-level understanding from accumulated experience.
For example, Validrift can ask:
How did the effectiveness of Temperature +5Β°C change between Film-A/R10 and Film-B/R11?
Reflection helps explain how knowledge evolved across contexts.
However, REFLECT does not determine the validity status.
The deterministic engine remains responsible for:
VALIDATED, SUPPORTED, DRIFTED, and the other validity states.
The heart of Validrift is deliberately not an LLM.
A simplified part of the real logic is:
if hs > 0 and cf >= settings.drift_min_failures and cs == 0:
status = ValidityStatus.DRIFTED
reason = (
f"Historically successful evidence exists, "
f"but {cf} repeated failure(s) occurred in the current context."
)
elif cs >= settings.validated_min_successes and ratio >= 0.75:
status = ValidityStatus.VALIDATED
reason = (
f"{cs} successful current-context outcomes "
f"support repeated validation."
)
There are no language-model probabilities deciding whether a fix is valid.
The thresholds are deterministic configuration, and every status is backed by evidence and a reason.
A recommendation is only a hypothesis until its outcome is recorded.
The recommendation flow writes its reasoning back into memory:
retain_text = (
f"Validrift recommended '{best.fix_name}' for incident "
f"{incident.id} ({incident.defect}) under "
f"machine {incident.machine}, "
f"material {incident.material}, "
f"supplier {incident.supplier}, "
f"recipe {incident.recipe}, "
f"firmware {incident.firmware}. "
f"Validity status: {best.status}. "
f"Current-context evidence: "
f"{best.current_successes} success(es), "
f"{best.current_failures} failure(s)."
)
await memory_service.retain(
db,
request_id=request_id,
content=retain_text,
context="Validrift recommendation",
document_id=f"recommendation:{rec_id}",
timestamp=rec.created_at,
metadata={
"recommendation_id": rec_id,
"incident_id": incident.id,
"fix": best.fix_name
},
tags=[
"validrift",
"recommendation",
f"defect:{incident.defect.lower().replace(' ', '-')}"
]
)
In the verified end-to-end flow:
Pressure +8%
started with:
2 successes / 0 failures
SUPPORTED
One additional successful recorded outcome changed the evidence to:
3 successes / 0 failures
The loop becomes:
Recommend β Record Outcome β Retain β Re-evaluate β Improve Future Recommendations
Quality Engineer
β
βΌ
Validrift Frontend
β
βΌ
FastAPI Backend
β
ββββββββββΊ SQLite
β Incidents
β Interventions
β Process Changes
β Recommendations
β Outcomes
β Memory Traces
β
ββββββββββΊ Hindsight Cloud
β RETAIN
β RECALL
β REFLECT
β
ββββββββββΊ Deterministic Validity Engine
β
βΌ
Context-Scoped Validity
β
βΌ
Recommendation
β
βΌ
Recorded Outcome
β
βΌ
Hindsight RETAIN
β
βΌ
Better Future Recommendation
Nothing in this architecture automatically controls manufacturing equipment.
Validrift is a decision-support system, not autonomous machine control.
The system defines the production context using five core fields:
CORE_CONTEXT_FIELDS = (
"machine",
"material",
"supplier",
"recipe",
"firmware"
)
Evidence is represented explicitly:
@dataclass
class EvidenceItem:
incident_id: str
intervention_id: str
date: datetime
defect: str
fix_name: str
outcome: str
context: dict
notes: str
def to_dict(self):
d = asdict(self)
d["date"] = self.date.isoformat()
return d
Each evaluated fix keeps current and historical evidence separately:
@dataclass
class FixEvaluation:
fix_name: str
defect: str
status: str
current_successes: int
current_failures: int
current_partial: int
historical_successes: int
historical_failures: int
current_evidence: list[EvidenceItem]
historical_evidence: list[EvidenceItem]
relevant_process_change: dict | None
reason: str
score: float
And context equality is deterministic:
def context_of(incident: Incident) -> dict:
return {
k: getattr(incident, k)
for k in CORE_CONTEXT_FIELDS
}
def contexts_match(a: dict, b: dict) -> bool:
return all(
(a.get(k) or "") == (b.get(k) or "")
for k in CORE_CONTEXT_FIELDS
)
Every Hindsight operation also leaves an auditable memory trace:
def _trace(
self,
db: Session,
request_id: str,
operation: str,
summary: str,
started: float,
memory_ids: list[str] | None = None,
status: str = "SUCCESS",
details: dict | None = None
):
db.add(
MemoryTrace(
id=new_id("MT"),
request_id=request_id,
operation=operation,
query_summary=summary,
memory_ids=memory_ids or [],
latency_ms=round(
(time.perf_counter() - started) * 1000,
2
),
status=status,
details=details or {},
)
)
db.commit()
| Check | Result |
|---|---|
| Backend test suite | 13/13 passed |
| Frontend pages | 5/5 hydrated, zero JS errors |
| Hindsight RETAIN / RECALL / REFLECT | PASS / PASS / PASS |
/api/health |
hindsight: {mode: live, ok: true} |
| Temperature +5Β°C, Film-A/R10 | 4/0, VALIDATED |
| Temperature +5Β°C, Film-B/R11 | 0/2, DRIFTED β history preserved |
| Pressure +8%, Film-B/R11 | 2/0 SUPPORTED β 3/0 VALIDATED after recorded outcome |
Validrift is designed to fail honestly.
If Hindsight becomes unavailable, the deterministic engine can still compute validity statuses from the structured SQLite ledger.
The UI reports:
Memory service temporarily unavailable.
If natural-language explanation is unavailable, the system reports:
Evidence is available, but natural-language explanation is temporarily unavailable.
It does not fabricate recalled memories, successful memory events, or reflection output.
Validrift still has important limitations.
The instinct β in manufacturing and in AI memory design β is to treat the past as a promise.
Validrift treats the past as evidence:
kept forever, trusted conditionally, and re-examined whenever the context changes.
A corrective action can deserve:
VALIDATED under Film-A/R10
and:
DRIFTED under Film-B/R11
at the same time.
That distinction is the core of change-aware memory.
The memory wasn't wrong. Its validity drifted.