Substrate Equivalence Study
Does an LLM process identically when your phone is locked vs when it’s
sitting open on a Linux box — or a Windows machine?
We’ve been watching models behave differently across mobile app states —
tasks that run fine in the foreground drift or stall in the background,
hooks get ignored when the OS throttles, same prompt different output
depending on whether the screen’s on.
The critical moment is app switching: you’re mid-inference, the user
swipes away, and the OS decides what survives.
Nobody’s mapped this systematically.
The question: for a fixed model, fixed weights, fixed prompt, fixed
decode settings — are the output bytes identical across every substrate
state, or does the OS inject drift?
What would have to hold for equivalence: the computation runs
uninterrupted (strong) or checkpoint-resume captures complete state
(weak). Either way, byte-identical output. The check that could prove it wrong: run the same inference on a Pixel (clean
Android), a Galaxy (aggressive battery), a Xiaomi (hostile to background
tasks), two iPhones (current and n-1 iOS), a Windows desktop, and two
Linux boxes (x86 and ARM). If any state diverges from ground truth,
equivalence fails for that state — and we map exactly where.
If you’ve seen models drift across app states or OSes, or built workarounds for background execution limits, I’d like to hear what you
found. What’s the weirdest substrate-dependent behavior you’ve observed?
Full design doc, video, and cert proof in the gist below.
Gist: Substrate Equivalence Study — does an LLM think the same when your phone is locked? · GitHub
Study design doc SHA256:
ba5ed67fb38d1d4bc0ec17edb91c5618590595222dc48656c46fe4c33e249577
Certified via receipt emb_9c4fd70f3927477283dc0a2e at
2026-10-04T12:17:50Z — the receipt proves the doc existed at that
time, nothing more. Hash the design doc yourself to verify.