{"slug": "from-token-to-trajectory", "title": "From Token to Trajectory", "summary": "A small experiment using Pythia-160M traces the token \"light\" through six contexts across its two senses — weight and color — to measure where contextual separation emerges in hidden states and how it changes near the output. A separate experiment compares JEV's predefined-choice interface with label generation by Claude and GPT on the same banking-intent tasks, examining accuracy, output validity, response time, and available cost data, though the differing underlying models mean the comparison describes the systems as tested rather than isolating output format. The work argues that using a model in an application requires connecting internal representations with the output the application needs.", "body_md": "*The suitcase felt surprisingly light.**The color looked surprisingly light.*\n\n**The same token can begin with the same embedding and develop different representations.** Calling a word “a vector” hides this transformation: inside a Transformer, the initial embedding is only the starting point of a computation shaped by context. Here, *light* refers to weight in one sentence and color in the other. By holding its token ID and position fixed while changing the preceding words, we can follow its hidden state through successive layers. **Where does contextual separation emerge — and what happens to it near the output?**\n\nFollowing that transformation requires deciding what counts as a difference. Do we care about the vectors’ directions, their magnitudes, or the distances between them? These measurements capture different properties of the same representations and can reveal different patterns across depth. Understanding how context reshapes a token therefore requires examining both the model’s computation and the measurements we use to describe it.\n\nUsing Pythia-160M, I trace *light* through six contexts spanning these two uses. I compare representations within and between the context groups, then follow the final hidden state through the language-model head to next-token probabilities. The experiment is deliberately small: it makes contextual variation measurable while leaving open how much of that variation can be attributed to word sense.\n\nFollowing this path raises a second engineering question: **what output does the application need?** A banking support system may need one intent label from a predefined set. In that setting, both the correctness of the decision and the validity of its format matter.\n\nA separate experiment compares JEV’s predefined-choice interface with label generation by Claude and GPT on the same banking-intent tasks. I examine accuracy, output validity, response time, and available cost data. Because the underlying models differ, this comparison describes the systems as tested; it does not isolate the effect of output format alone.\n\nUnderstanding a model well enough to use it in an application requires connecting what happens inside the model with what the application needs from its output. These experiments examine two parts of that process: choosing measurements that support a careful interpretation of internal representations, and testing whether a system’s outputs meet the requirements of a concrete task.\n\nBefore following *light* through the model, we need to distinguish the numerical objects along its path. In a causal language model, the broad computational sequence is:\n\nA tokenizer first converts text into token IDs. Each ID retrieves a learned vector from an embedding table. At this lookup stage, the same token ID receives the same vector regardless of the surrounding sentence. Context enters through the computations that follow, together with the model’s mechanism for representing position.\n\n**There is no single “vector space.”** The objects along this path serve different purposes, even when some share the same dimensionality.\n\nThe distinction matters because a measurement is meaningful only in relation to the object being measured. Comparing two token hidden states inside a language model asks a different question from comparing a query with a document embedding.\n\nHere, the object of interest is the **hidden-state trajectory at a fixed token position**.\n\nLet *h* denote the representation at the target position after each Transformer block. Following the target through a model with (L) blocks gives a sequence:\n\nThis is a trajectory through computational depth: each state results from another transformation of the preceding state.\n\nAttention allows the target position to incorporate information from other available positions. In a causal Transformer, it can attend to itself and preceding tokens, but not future tokens. The MLP applies a learned nonlinear transformation at each position, while residual connections add the computed updates to the running representation. Their arrangement varies by architecture; Pythia uses parallel attention and MLP branches within each block.\n\nFor *light*, the preceding words therefore matter. A prefix about a suitcase and a prefix about a color provide different information for attention to combine. Even when the target token ID and position are held constant, its later hidden states can differ. Those differences reflect the entire available context, however, so they cannot automatically be attributed to lexical sense alone.\n\nEarlier research provides two complementary perspectives. In *A Primer in BERTology: What We Know About How BERT Works*, Rogers, Kovaleva, and Rumshisky (2020) review evidence about the information encoded in BERT and its distribution across the model. Their survey motivates examining representations at multiple layers rather than treating the final layer as the only informative one.\n\nEthayarajh (2019), in *How Contextual are Contextualized Word Representations?*, examines the geometry of BERT, ELMo, and GPT-2. One measure is **self-similarity**: the average cosine similarity between representations of the same word across different contexts. Interpreting this measure requires accounting for anisotropy — the tendency of representations to favor certain directions. If randomly sampled representations already have high cosine similarity, a high similarity between two occurrences of a word says less than it initially appears to.\n\nAfter adjusting for this background similarity, Ethayarajh found that representations were generally more context-specific in higher layers. This motivates a layer-wise investigation, but does not establish the pattern we should expect from Pythia or from a small set of contrasting uses of one word.\n\nThe question is therefore empirical: **When, and in what geometric form, does contextual differentiation emerge across Transformer layers?**\n\nTo make contextual change measurable, I followed one token through Pythia-160M in six short contexts. Three use *light* in relation to weight; three use it in relation to color.\n\nAll six inputs contain exactly five tokens. The target, Ġlight (token ID 1708), appears once in each input, at position 4 under zero-based indexing. No special tokens are added, and each input ends at the target. The token’s identity, position, and available context length are therefore held constant.\n\nWith the model in evaluation mode, I extracted the target-position vector from each entry in the returned hidden-state sequence. This produced 13 recorded states, indexed 0–12, with 768 dimensions per state. In the implementation checked here (Transformers 5.17.0), state 12 is the output after the final Transformer block and final LayerNorm. The transition from state 11 to state 12 therefore includes both operations. The extracted tensors were FP16 and were converted to FP32 before calculating the geometric metrics.\n\nThese controls narrow the comparison, but they do not isolate lexical sense. The prefixes contain different nouns and, in the first pair, different verbs. The experiment measures geometry associated with the weight and color context groups; differences in vocabulary and sentence construction remain part of that comparison.\n\nTwo representations can differ in direction, length, or both. I therefore measured cosine similarity, Euclidean distance, and vector norm, then repeated the distance analysis after normalizing each vector to unit length.\n\nThe mathematical setting is simply a finite-dimensional real inner-product space — a Hilbert space in which angles, norms, and distances are well defined. The practical question is which of these properties captures the distinction we want to examine.\n\nFor each recorded state, I compared all six unordered within-group pairs with all nine between-group pairs. Cosine separation was defined as\n\nFor Euclidean distance, I reversed the subtraction:\n\nUnder both definitions, a positive value indicates that representations are, on average, more similar within the context groups than between them. These are descriptive contrasts: the pairs share the same six inputs and are not independent observations.\n\nThe normalized-distance analysis removes each vector’s length before comparison. It is closely related to cosine similarity:\n\nConsequently, it provides a complementary view of directional geometry rather than independent evidence of contextual separation. Unit-length normalization also does not remove anisotropy; no random-context baseline subtraction is applied in this analysis.\n\n**Choosing a similarity metric is already an engineering assumption about which geometric property matters.**\n\nThe main figure compares cosine separation, raw Euclidean separation, representation norms, and unit-normalized Euclidean separation across the 13 recorded states.\n\nAt state 0, the target representations are identical across all six inputs: cosine similarity is 1 and pairwise distance is 0. This is the shared embedding starting point. At states 1–2, all three separation measures are negative. The representations have changed with context, but they do not yet exhibit the intended within-group advantage.\n\nAll three contrasts become positive at state 3. Cosine separation reaches its largest observed value at state 6, at 0.014845, and remains between approximately 0.010 and 0.015 through state 11. Unit-normalized Euclidean separation likewise remains positive over this interval. In these examples, a sustained group distinction is visible in intermediate states.\n\nThe final recorded state presents a different pattern.\n\nAt state 12, mean within-group and between-group cosine similarities reach 0.996840 and 0.996080. Their difference remains positive but is much smaller than at state 11. Meanwhile, raw Euclidean separation increases, alongside a roughly ninefold increase in mean vector norm in both groups. Once vector lengths are normalized, the distance contrast decreases instead.\n\nThere is no mathematical contradiction. Euclidean distance combines scale and direction:\n\nLarge vector norms can magnify absolute distances even when directions are close. A larger raw distance gap therefore does not, by itself, establish stronger directional separation. Equally, increasing vector length alone cannot explain the change in cosine similarity, which is invariant to positive rescaling.\n\n**Distance is not a single thing in representation space.**\n\nHere, absolute and direction-based measures describe different aspects of the final state. The measurements describe the combined effect of the final Transformer block and final LayerNorm; they do not isolate each operation’s contribution to the geometric changes.\n\nThe observed separation does not increase steadily with depth. It emerges, remains positive across intermediate states, and becomes less pronounced under direction-based measures at the final recorded state. That pattern is more informative than a single similarity score taken from the model’s output.\n\nIts scope is nevertheless narrow. **Sense-related geometric separation is not semantic understanding.** Six selected contexts cannot establish a general layer profile, and the design cannot fully separate lexical sense from other differences in the prefixes. Nor do these measurements show that the observed geometry causes a particular prediction.\n\nWhat the experiment makes visible is a context-dependent trajectory whose interpretation changes with the measurement. The next step is to follow that trajectory to its computational endpoint: how does the final hidden state become a distribution over possible next tokens?\n\nThe trajectory so far ends with a vector. For a language model, however, that vector is an input to another computation: predicting what comes next.\n\nAt the final position of each prefix — the token *light* — the representation supplied to the language-model head is projected into vocabulary logits. Softmax then converts those scores into a next-token probability distribution:\n\nHere,\n\ndenotes the representation passed to the LM head, including the model’s final normalization. Each vocabulary token has an associated output weight vector; its logit depends on how the hidden state projects onto that vector.\n\nThis projection connects the geometry inside the model to its language output. Cosine similarity and Euclidean distance summarize relationships between hidden states, but neither directly determines how similar their next-token distributions will be. The output head responds to particular directions in representation space. A difference that looks small under a global distance metric may still change the relative scores of candidate tokens.\n\nFor the six *light* contexts, the next comparison is therefore at the vocabulary level: the five highest-probability continuations for each prefix, using probabilities computed over the full vocabulary.\n\nThe candidate sets overlap substantially across the six contexts. A comma, a period, and “ and” appear among the top five in every case, while their ranks and probabilities vary. For example, the comma receives 29.07% in the backpack context and 17.33% in the package context. These outputs show variation in the model’s continuation probabilities, but the displayed candidates do not provide a clear division between the weight and color groups.\n\nThese are predictions about the token **after** *light*, not probabilities assigned to its two senses. The model is continuing a prefix, rather than explicitly choosing between “not heavy” and “not dark.” Its predictions can reflect syntax, familiar phrases, punctuation, and other properties of the preceding text. The five displayed probabilities also need not sum to one: the remaining probability mass belongs to the rest of the vocabulary.\n\nThis completes the computational path introduced at the beginning: text becomes token IDs, token embeddings become contextual hidden states, and the final representation becomes a distribution over possible continuations. Selecting a token from that distribution would begin the next step of generation.\n\nFollowing this path also clarifies what the experiment can establish. The measurements describe how six context-dependent representations differ across depth. They do not establish that geometric separation is equivalent to meaning, or that a larger separation indicates better understanding.\n\nLikewise, observing hidden-state differences alongside output differences does not identify which internal features caused a particular prediction. The final representation participates directly in computing the logits, but our summary metrics do not isolate the features responsible. A targeted causal test would require an intervention — for example, modifying a selected activation while holding the input fixed, then measuring how the output changes.\n\n**The trajectory makes the computation visible; it does not make geometry a complete explanation of meaning.**\n\nThat distinction leads to the second experiment. Next-token prediction is a natural endpoint for a language model, but an application may need a category, a score, or a routing decision. Evaluating such outputs requires a different level of analysis: whether the system makes the right decision and returns it in a usable form.\n\nThe two experiments therefore address complementary engineering questions. The first examines representations inside an inspectable model. The second compares the behavior of systems at their interfaces, without assuming that their internal computations are equivalent.\n\n**Why generate language when the application does not need language?**\n\nThe first experiment followed a representation through a language model to next-token probabilities. The second examines a different level of the system: the output an application receives. Given a banking support message and a fixed set of intents, can the system return the correct label in a form that software can use?\n\nI compared JEV’s choice API(typesafe/jev-1.13) with prompt-based label generation by Claude Haiku 4.5 (claude-haiku-4-5-20251001) and GPT-4.1 nano (gpt-4.1-nano-2025-04-14). The comparison evaluates these systems under the recorded configurations. Because their underlying models differ, it cannot isolate the causal effect of the output interface.\n\nThe main evaluation used selected intents from the BANKING77 test split. The four-category task contained 160 messages, with 40 examples each for refund requests, refunds not received, duplicate charges, and unrecognized card payments.\n\nThe eight-category task retained all 160 original messages and their labels, then added 40 examples each for pending, declined, and reverted card payments, and fees charged on card payments. This produced a balanced set of 320 messages. The saved preparation manifest records a random seed of 42. Because the original messages appear in both tasks, the four- and eight-category results are overlapping evaluations.\n\nA separate diagnostic set contained 24 constructed messages organized into 12 contrasting pairs. These tested distinctions such as requesting a refund versus waiting for an issued refund, or recognizing a payment versus reporting it as unauthorized. This set was evaluated with the four-category label vocabulary and reported separately from the BANKING77 results.\n\nA shared configuration defined the allowed labels and their descriptions. Claude and GPT received identical system prompts for each task, including the instruction to return exactly one allowed label, spelled as provided, without an explanation. Both were configured with temperature zero and an output limit of 64 tokens. JEV returned its prediction through a structured choice field.\n\nThe label vocabulary retained the original spelling and punctuation, including Refund_not_showing_up and reverted_card_payment?. These details matter when evaluating whether an output satisfies an exact label contract.\n\nI evaluated correctness and output validity separately. A prediction was correct when the extracted label exactly matched the reference label. An output was valid when that label belonged to the allowed set. A valid but incorrect label therefore counted as a classification error, while an output outside the allowed vocabulary failed both checks.\n\nThe primary results use strict label matching. A separate post hoc analysis examines specific capitalization and punctuation variants observed in Claude’s eight-category outputs; it does not replace the strict scores.\n\nAlongside these measures, I retained recorded client latency and available API usage information. Cost comparisons require explicit pricing assumptions and complete usage records; missing values are treated as unavailable. Calibration was not evaluated because comparable class-probability estimates were not collected across all three systems.\n\nThese measurements address two practical requirements of a classification system: selecting the right intent and returning an output the application can reliably consume.\n\nThe accuracy ranking changed between the four- and eight-category tasks. Claude performed best on the four-category set, correctly classifying 159 of 160 messages. JEV classified 157 correctly, and GPT classified 142. All three returned valid labels for every case in this task.\n\n*Accuracy uses strict label matching. The eight-way set includes all 160 four-way messages.*\n\nAll three systems also classified every boundary case correctly. These constructed pairs provided a check on specific distinctions, such as requesting a refund versus waiting for an issued refund. They did not distinguish the systems from one another, and success on 24 designed examples should not be interpreted as comprehensive robustness.\n\nOn the eight-category task, JEV achieved the highest strict accuracy. Claude’s score, however, requires a closer look at the outputs themselves.\n\nClaude returned 22 outputs outside the eight-category vocabulary. Twenty omitted the question mark from reverted_card_payment?, one capitalized its first letter, and one gave an explanation instead of an allowed label. Its output-validity rate was therefore 298/320, or 93.13%. JEV and GPT returned valid labels for all 320 cases.\n\nTo examine the contribution of these label-format differences, I conducted a post hoc analysis mapping the two observed variants — reverted_card_payment and Reverted_card_payment?—to the allowed label. This corrected 20 predictions and raised Claude’s accuracy to 291/320, or 90.94%. One additional prediction became valid but remained incorrect: the reference label was request_refund. The explanatory response remained invalid.\n\nThis adjustment changes the interpretation of the strict ranking. Claude moves above GPT after the specified normalization, while remaining below JEV. Because the mapping was chosen after inspecting the outputs, the adjusted score is a sensitivity analysis rather than a replacement for the primary result.\n\nFor an application, these are distinct failure modes. A label variant may be recoverable with an explicit mapping. A correctly formatted prediction for the wrong intent still sends the request to the wrong category. Reporting both validity and correctness makes that distinction visible.\n\nThe eight-way results also show why accuracy alone is insufficient for choosing a system.\n\nJEV combined the highest strict accuracy with complete output validity, but had the longest mean client latency. Claude had the shortest mean latency, while GPT combined complete validity with a higher strict accuracy than Claude. Whether the normalization step is acceptable changes the practical comparison between the two generative systems.\n\nThese timings include the conditions experienced by the client. The systems were evaluated in separate runs, so network and service conditions may contribute to the differences. They should not be read as an isolated measurement of model computation or as a guaranteed production response time.\n\nCost requires similar care. The archive contains token-usage records for Claude and GPT and credit usage for JEV’s eight-way run, but these are different billing units. Some earlier JEV records lack usage information. A monetary comparison therefore requires documented conversion rates and explicit treatment of missing records.\n\nThese trade-offs matter in workflows that repeatedly classify requests, check retrieved evidence, or select the next action. A decision-oriented component could handle some of these bounded judgments, with more demanding or uncertain cases passed to another model or a human reviewer. Whether this division of work is worthwhile depends on measured decision quality, latency, and total workflow cost.\n\nThe experiment demonstrates differences between these particular systems and configurations. It does not establish that a choice-oriented interface causes better decisions, or that generated labels must be unreliable. The application-level question is more concrete: how often does the system choose the right label, how often can the output be consumed as expected, and what time and cost are required to obtain it?\n\nThe two experiments examine different parts of an AI system: the representations formed inside a Transformer and the outputs delivered to an application. Their connection is an engineering question: which part of a computation is useful for the task, and how should that usefulness be evaluated?\n\nWhen extracting a representation, choosing a model is only the beginning. The layer and token position determine which state we observe. For retrieval or RAG, pooling and the similarity metric introduce further choices about how information is combined and compared. A representation that separates a few contrasting contexts is not automatically a good retrieval embedding; its value must be tested on the documents, queries, and relevance criteria of the intended application.\n\nThe Pythia experiment makes these choices visible. Cosine similarity, Euclidean distance, and vector norms describe different aspects of the same hidden states, and their patterns need not agree. For interpretability and debugging, examining them together helps locate where contextual differences emerge and how their geometry changes across depth. Establishing whether those differences cause a particular output would require interventions beyond the measurements used here.\n\nAt the application boundary, the choice concerns what the system must return. An explanation requires language; a routing step may require only a valid category. In the classification experiments, accuracy, output validity, and client latency exposed different trade-offs. Claude’s label variants also showed that parsing and normalization belong to the system being evaluated: accepting an alternative spelling can recover a usable answer, but it cannot repair an incorrect decision.\n\nA workflow could therefore assign bounded checks to one component and more demanding reasoning or explanation to another. Whether that arrangement helps must be measured across the full workflow, including errors, escalation, latency, and cost. The present experiments inform that design question without establishing a universally preferable architecture.\n\nBefore choosing a representation or an endpoint, practitioners can ask five questions:\n\nA hidden state is not “the embedding” of a token. It is one representation at one point in a computation, and whether it is useful depends on what you want to do with it.\n\nThe same is true of an output: generation is one endpoint of computation, not computation itself.\n\nThe [companion Colab notebook](https://colab.research.google.com/drive/1SyMyk5NwqYaprK2GujIkwQcvClOY4rNI?usp=sharing) includes code for the six-context Pythia experiment and recorded predictions from the banking-intent comparison. Readers can regenerate the figures and evaluation tables without paid API calls or API keys.\n\nEthayarajh, K. (2019). [How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings](https://aclanthology.org/D19-1006/). *Proceedings of EMNLP-IJCNLP*, 55–65.\n\nRogers, A., Kovaleva, O., & Rumshisky, A. (2020). [A Primer in BERTology: What We Know About How BERT Works](https://aclanthology.org/2020.tacl-1.54/). *Transactions of the Association for Computational Linguistics, 8*, 842–866.\n\nBiderman, S., et al. (2023). [Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling](https://proceedings.mlr.press/v202/biderman23a.html). *Proceedings of the 40th International Conference on Machine Learning, 202*, 2397–2430.\n\nCasanueva, I., Temčinas, T., Gerz, D., Henderson, M., & Vulić, I. (2020). [Efficient Intent Detection with Dual Sentence Encoders](https://aclanthology.org/2020.nlp4convai-1.5/). *Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI*, 38–45.\n\n[From Token to Trajectory](https://pub.towardsai.net/from-token-to-trajectory-8300f311f6a2) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/from-token-to-trajectory", "canonical_source": "https://pub.towardsai.net/from-token-to-trajectory-8300f311f6a2?source=rss----98111c9905da---4", "published_at": "2026-10-01 17:01:03+00:00", "updated_at": "2026-10-01 17:18:01.851441+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "ai-tools"], "entities": ["Pythia-160M", "JEV", "Claude", "GPT"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-token-to-trajectory", "markdown": "https://wpnews.pro/news/from-token-to-trajectory.md", "text": "https://wpnews.pro/news/from-token-to-trajectory.txt", "jsonld": "https://wpnews.pro/news/from-token-to-trajectory.jsonld"}}