cd /news/large-language-models/guarantees-on-dynamical-system-disti… · home topics large-language-models article
[ARTICLE · art-84166] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Guarantees on Dynamical System Distinguishability for LLM Token Generation

A new arXiv preprint (2607.28667v1) proves that classifying large language models (LLMs) by modeling token embeddings as trajectories of a black-box dynamical system (DS) achieves misclassification probability that decays exponentially in sequence length L, governed by a dynamical discriminability quantity δ². The authors show that total variation distance between stationary marginal distributions of two stochastic linear DSs can be arbitrarily small even when dynamics differ, establishing a fundamental accuracy floor for classifiers ignoring token dynamics. They also characterize cross-embedding generalization via an approximate intertwining condition, linking transferable discriminability to the smallest singular value of the intertwining map.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28667v1 Announce Type: new Abstract: Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the token sequence, and when it transfers across embedding models remains lacking. We address these questions by formalizing the classification task as a binary hypothesis test between two stochastic linear DSs. We show that the total variation distance between the stationary marginal distributions of the two DSs can be arbitrarily small even when the dynamics differ substantially, which provides a fundamental accuracy floor for any classifier that ignores token dynamics. We then show that the misclassification probability of DS-based classification decays exponentially in the sequence length $L$, with the decay governed by a dynamical discriminability quantity $\delta^2$ that captures the spectral distance between the two DSs. We also characterize cross-embedding generalization by introducing an approximate intertwining condition between embedding models and establishing a lower bound on the transferable discriminability in terms of the intertwining map's smallest singular value. Together, these results explain the empirical performance of DS-based classification and motivate further investigation into using DS theory to analyze AI systems, in contrast to the more common approach of using AI to model dynamical systems.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/guarantees-on-dynami…] indexed:0 read:1min 2026-08-03 ·