{"slug": "from-joint-state-transition-prediction-to-language-a-minimal-predictive-of", "title": "From Joint State-Transition Prediction to Language: A Minimal Predictive Hypothesis of Intelligence", "summary": "A new paper proposes a minimal hypothesis of intelligence, suggesting that it may consist solely of predicting transitions between high-dimensional joint states, without requiring causality, logic, or symbolic reasoning as primitive mechanisms. The authors argue that language emerges from this predictive process, and they suggest a unified multimodal latent state space for AI, where perception, language, memory, action, and world dynamics are learned through state-transition prediction.", "body_md": "This paper proposes a minimal hypothesis connecting physical structure, biological intelligence, language, and artificial intelligence.\n\nThe central claim is that intelligence may not require causality, logic, symbolic reasoning, planning, or explicit object relations as primitive cognitive mechanisms. At its lowest level, intelligence may consist only of predicting transitions between high-dimensional joint states.\n\nReality is minimally assumed to admit local states that can participate in larger joint states and undergo state transitions. A nervous system, itself composed of many simultaneously active units, naturally supports distributed high-dimensional states and can learn to predict how such states change.\n\nLanguage is proposed to emerge from this predictive process rather than from a predesigned symbolic system. During practical interaction with the world, sounds, gestures, perceptions, actions, and bodily states occur together. When sounds become reliably predictive of other states, they acquire symbolic function. Once symbols begin predicting other symbols, prediction can operate in a compressed, recursively composable symbolic state space. On this view, explicit causality, logic, mathematics, planning, and science emerge from increasingly complex language-state prediction rather than from separate underlying cognitive mechanisms.\n\nThis hypothesis suggests a corresponding direction for artificial intelligence: a unified multimodal latent state space in which perception, language, memory, action, and world dynamics are learned through state-transition prediction.\n\nWe begin with a deliberately weak assumption about reality.\n\nReality can be represented, for an observer, as states that change. States can also contain distinguishable local structure and participate in larger joint states.\n\nLet `X_t` denote the state accessible to an intelligent system at time `t`.\n\nA state transition can be written simply as:\n\n`X_t → X_(t+1)`\n\nIf the transition is uncertain, prediction can be expressed as:\n\n`P(X_(t+1) | X_≤t)`\n\nNothing in this formulation requires an explicit causal relation.\n\nSuppose two local states, `A` and `B`, participate in the same joint state. Instead of assuming that `A` causes a change in `B`, we can minimally describe the observed transition as:\n\n`(A_t, B_t) → (A_(t+1), B_(t+1))`\n\nThe same idea extends to arbitrarily high-dimensional states.\n\nA **local world** can be understood as a relatively stable and partially autonomous region of state-transition structure. Different local worlds may have different internal regularities, may combine into larger joint states, and may interact without requiring a single predefined global decomposition of reality.\n\nThe theory therefore starts with three minimal assumptions:\n\nCausality, logic, objects, forces, and other higher-level relations are not required as primitive assumptions.\n\nA biological nervous system is naturally suited to representing joint states.\n\nA brain state is not a single indivisible symbol. At any moment, many neurons and neural populations are active simultaneously. A simplified neural state may therefore be represented as:\n\n`Z_t = (z_1, z_2, ..., z_n)`\n\nDifferent aspects of perception, memory, bodily condition, ongoing action, and previous experience can coexist within this distributed state.\n\nThe minimal cognitive operation proposed here is:\n\n`Z_≤t → predicted Z_(t+1)`\n\nor probabilistically:\n\n`P(Z_(t+1) | Z_≤t)`\n\nThe system predicts its next joint state.\n\nUnder the strong form of this hypothesis, intelligence does not require separate primitive mechanisms for:\n\nThese may appear as useful high-level descriptions, implementation details, or emergent regularities, while the underlying operation remains joint state-transition prediction.\n\nAction does not need to be treated as an external module attached to a world model.\n\nThe current joint state may include:\n\nWe may write:\n\n`Z_t = (world, body, memory, language, action, goal, ...)`\n\nThe next predicted state naturally contains both a future world state and the organism's own future action:\n\n`Z_t → Z_(t+1)`\n\nFrom an external engineering perspective, this may look like:\n\n`observation → policy → action`\n\nbut at a more unified level it can be interpreted as:\n\n`current joint state → next joint state`\n\nThe organism's next movement is simply one component of the future state being predicted.\n\nPlanning can then be interpreted as longer-horizon prediction:\n\n`Z_t → Z_(t+1) → Z_(t+2) → ... → Z_(t+k)`\n\nA sequence that appears externally as a plan may arise from prediction over future joint states containing both self and environment.\n\nThe value of diverse experience is not merely that it exposes a system to more situations.\n\nIf a system only memorizes:\n\n`X_1 → Y_1`\n\n`X_2 → Y_2`\n\n`X_3 → Y_3`\n\nit has learned particular transitions but little reusable structure.\n\nWhen one finite system must accurately predict a wide variety of transitions, it is pressured to compress repeated regularities into reusable latent structure.\n\nA useful summary is:\n\n`diverse experience → predictive compression → reusable latent transition structure`\n\nThis latent structure does not need to explicitly contain concepts such as:\n\nIt only needs to improve prediction across many situations.\n\nUnder this view, what humans later call \"knowledge\" may correspond to stable, generalizable predictive structure.\n\nLanguage need not begin as an intentionally designed representation system.\n\nConsider an organism interacting socially with its environment. Its experienced joint state contains:\n\nInitially, a sound is simply another event in the world.\n\nFor example, a recurring joint state may contain:\n\n`(water-like visual state, sound \"water\", bodily state, social context, ...)`\n\nThe predictive system gradually learns regularities such as:\n\n`water-like visual state → expected sound`\n\nand also:\n\n`sound → expected visual / bodily / situational state`\n\nNo symbolic designer is required.\n\nThe sound gains symbolic function because it becomes reliably connected, through prediction, with other parts of the joint state.\n\nSymbolic function is therefore proposed to emerge from stable predictive coupling.\n\nThe crucial transition occurs when a sound or other signal can partially substitute for the corresponding real-world state in prediction.\n\nAt first:\n\n`world state A → world state B`\n\nThen a sound becomes associated with world state `A`:\n\n`world state A ↔ sound A`\n\nEventually, hearing `sound A` can activate predictions that would normally follow from directly experiencing `world state A`.\n\nIn functional terms:\n\n`sound A ≈ predictive proxy for world state A`\n\nThe sound is not identical to the real state. It becomes useful because it can stand in for that state within the predictive process.\n\nThis creates a new possibility:\n\n`symbol → prediction about world`\n\nand eventually:\n\n`symbol → symbol`\n\nAt that point, prediction no longer requires the corresponding physical event to be present at every step.\n\nOnce symbolic states become sufficiently stable, they can begin predicting other symbolic states.\n\nThe system moves through three broad stages:\n\n`world → world prediction`\n\nthen:\n\n`world ↔ symbol`\n\nand finally:\n\nThe third stage is especially important.\n\nLanguage becomes a predictive environment of its own.\n\nA linguistic state can produce another linguistic state:\n\n`L_t → L_(t+1)`\n\nwhich becomes the input for another transition:\n\n`L_(t+1) → L_(t+2)`\n\nproducing a potentially long sequence:\n\n`L_t → L_(t+1) → L_(t+2) → ...`\n\nThis symbolic environment has unusual properties. Compared with raw sensory experience, symbols can be:\n\nPrediction therefore gains a low-cost symbolic substrate in which increasingly long and abstract state transitions can occur.\n\nThis hypothesis makes a stronger claim than saying that language merely expresses causal and logical structures already present inside cognition.\n\nIt proposes that explicit causality and logic may arise within the dynamics of language itself.\n\nConsider repeated linguistic transitions such as:\n\n`because A ... therefore B`\n\nor:\n\n`if A, then B; A; therefore B`\n\nThrough repeated prediction and generalization, such symbolic transition patterns become stable and composable.\n\nHumans later describe the resulting behavior using concepts such as:\n\nBut the latent predictive system does not necessarily contain separate hidden entities corresponding to \"causality\" or \"logic.\"\n\nThe stronger claim is:\n\n`logical behavior does not imply an explicit internal logic engine`\n\nand:\n\n`causal behavior does not imply an explicit internal causal relation`\n\nThe latent level may still consist only of high-dimensional predictive transition structure.\n\nUnder this interpretation, explicit logic and causality are stabilized forms of symbolic state transition.\n\nOnce symbols can predict symbols, language can become increasingly complex through its own continued operation.\n\nA simple symbolic state can generate another symbolic state, which itself becomes a new state available for prediction.\n\nFor example:\n\n`one → two → addition → multiplication → algebra → functions → calculus`\n\nThe later structures do not need to appear directly in raw sensory experience as ready-made objects.\n\nThey can emerge through repeated symbolic composition and prediction.\n\nLanguage therefore becomes a kind of **secondary predictive world**.\n\nIt remains grounded, at least historically and functionally, in interaction with reality, yet it can develop structures far beyond immediate perception.\n\nThis allows prediction to operate over:\n\nHigher intelligence may therefore depend less on the appearance of a new reasoning mechanism and more on the emergence of a new predictive state space.\n\nOnce linguistic and symbolic prediction becomes recursively composable, increasingly formal systems can emerge.\n\nCounting stabilizes numerical symbols.\n\nNumerical symbols participate in arithmetic.\n\nArithmetic supports algebra.\n\nFurther symbolic development produces geometry, calculus, formal logic, programming languages, and scientific theories.\n\nScience then creates a feedback loop between symbolic prediction and reality:\n\n`reality → prediction → language → symbolic prediction → prediction about reality → experiment`\n\nThe result of experiment changes later prediction and language.\n\nScientific concepts therefore do not need to be literal copies of reality's hidden ontology.\n\nThey can instead be treated as symbolic predictive structures constrained by reality.\n\nA scientific theory is valuable when it successfully:\n\nIn this sense, science can be understood as language becoming increasingly formal while remaining constrained by the world.\n\nThe local-world view adds an ontological layer to the hypothesis.\n\nReality does not need to be assumed to consist of one globally uniform set of primitive relations.\n\nInstead, different regions or systems may possess different local structures and transition regularities.\n\nA local world may be represented abstractly as:\n\n`W_i = (S_i, T_i)`\n\nwhere:\n\n`S_i` is a local state space,`T_i` describes local state transitions.\nWhen local worlds interact, a larger joint state can form:\n\n`(W_i, W_j) → joint state`\n\nThe resulting transition does not require us to assign a fundamental causal arrow between the two local worlds.\n\nWe may simply observe:\n\n`joint state_t → joint state_(t+1)`\n\nFrom the perspective of an observer, what later becomes described as an object, relation, force, cause, or rule may be a stable way of compressing repeated local-world interactions.\n\nThe local-world framework therefore provides a possible minimal ontology for the predictive theory:\n\n`local worlds → joint states → joint state transitions`\n\nCurrent AI systems already display separate pieces of this picture.\n\nLarge language models learn:\n\n`P(token_(t+1) | token_≤t)`\n\nand nevertheless exhibit behaviors associated with:\n\nWorld models learn regularities such as:\n\n`world state_t → world state_(t+1)`\n\nMultimodal systems connect text, image, audio, video, and other modalities.\n\nEmbodied agents additionally predict or generate actions.\n\nA natural long-term architecture suggested by this hypothesis is a unified latent state:\n\n`Z_t = (vision, audio, 3D, language, memory, body, action, goal, ...)`\n\nwith a shared predictive objective:\n\n`P(Z_(t+Δ) | Z_≤t)`\n\nThe same latent state would connect:\n\n`world ↔ language`\n\n`world ↔ action`\n\n`language ↔ action`\n\n`memory ↔ world`\n\nand potentially every other modality relevant to the agent.\n\nIn such a system, language prediction and world prediction would not be fundamentally separate kinds of intelligence.\n\nThey would be different regions or projections of one evolving predictive state.\n\nAction would also be part of the predicted future rather than a completely separate computational category.\n\nThe strongest scientific form of this proposal is:\n\nA sufficiently capable system trained to predict diverse multimodal joint state transitions can develop perception, action, linguistic abstraction, logical reasoning, causal reasoning, and planning without requiring these abilities to be implemented as separate primitive cognitive mechanisms.\n\nThis hypothesis is stronger than the general statement that prediction is important.\n\nIt predicts that increasingly capable unified predictive systems should be able to acquire behaviors traditionally attributed to separate cognitive modules without requiring those modules to be fundamental.\n\nThe hypothesis would be weakened if some important class of cognition consistently failed to emerge from sufficiently rich predictive learning and instead required a qualitatively distinct computational mechanism.\n\nPossible experimental comparisons could include:\n\nThe goal would not be to prove that explicit modules are never useful.\n\nThe question is whether they are **fundamental requirements for intelligence**.\n\nThe complete proposal can be compressed into the following chain:\n\n`local worlds`\n\n`↓`\n\n`joint states`\n\n`joint state transitions`\n\n`prediction of joint state transitions`\n\n`practical interaction with reality`\n\n`sounds and signals become predictive proxies`\n\n`symbols predict symbols`\n\n`language becomes a self-extending predictive state space`\n\n`logic, causality, mathematics, planning, and science emerge`\n\nThe entire process is hypothesized to remain grounded in one minimal mechanism:\n\n**joint state-transition prediction**\n\nThis paper proposes a minimal predictive hypothesis of intelligence.\n\nReality, for an observer, presents changing local and joint states. Biological intelligence learns to predict how those joint states change. During practical interaction, sounds and other signals become reliably coupled with parts of reality and gradually acquire symbolic function. Once symbols can substitute for real states in prediction, they begin predicting one another. Language then develops into a recursively self-extending predictive environment.\n\nWithin this symbolic environment, increasingly stable and complex transition structures appear. Humans describe some of them as causality, logic, mathematics, planning, and science.\n\nThe central hypothesis is therefore:\n\n**The latent foundation of intelligence may contain only joint state-transition prediction, while explicit causality and logic emerge later through the autonomous development of language prediction.**\n\nThis perspective suggests a unified direction for artificial intelligence as well:\n\n**A future general intelligence may be best modeled as one multimodal predictive system operating over a shared latent state in which world, language, memory, body, and action continuously predict one another.**\n\nThis hypothesis overlaps with several existing research traditions while making a stronger reductionist claim.\n\nThe distinctive claim developed here is that the latent cognitive level need not contain hidden causal or logical structures at all. It may consist only of predictive joint-state dynamics, while explicit causality and logic arise through the emergence and autonomous development of linguistic prediction.", "url": "https://wpnews.pro/news/from-joint-state-transition-prediction-to-language-a-minimal-predictive-of", "canonical_source": "https://dev.to/ylemkairos/from-joint-state-transition-prediction-to-language-a-minimal-predictive-hypothesis-of-intelligence-54a2", "published_at": "2026-09-08 09:49:53+00:00", "updated_at": "2026-09-08 10:03:06.703428+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/from-joint-state-transition-prediction-to-language-a-minimal-predictive-of", "markdown": "https://wpnews.pro/news/from-joint-state-transition-prediction-to-language-a-minimal-predictive-of.md", "text": "https://wpnews.pro/news/from-joint-state-transition-prediction-to-language-a-minimal-predictive-of.txt", "jsonld": "https://wpnews.pro/news/from-joint-state-transition-prediction-to-language-a-minimal-predictive-of.jsonld"}}