# From Joint State-Transition Prediction to Language: A Minimal Predictive Hypothesis of Intelligence

> Source: <https://dev.to/ylemkairos/from-joint-state-transition-prediction-to-language-a-minimal-predictive-hypothesis-of-intelligence-54a2>
> Published: 2026-09-08 09:49:53+00:00

This paper proposes a minimal hypothesis connecting physical structure, biological intelligence, language, and artificial intelligence.

The central claim is that intelligence may not require causality, logic, symbolic reasoning, planning, or explicit object relations as primitive cognitive mechanisms. At its lowest level, intelligence may consist only of predicting transitions between high-dimensional joint states.

Reality is minimally assumed to admit local states that can participate in larger joint states and undergo state transitions. A nervous system, itself composed of many simultaneously active units, naturally supports distributed high-dimensional states and can learn to predict how such states change.

Language is proposed to emerge from this predictive process rather than from a predesigned symbolic system. During practical interaction with the world, sounds, gestures, perceptions, actions, and bodily states occur together. When sounds become reliably predictive of other states, they acquire symbolic function. Once symbols begin predicting other symbols, prediction can operate in a compressed, recursively composable symbolic state space. On this view, explicit causality, logic, mathematics, planning, and science emerge from increasingly complex language-state prediction rather than from separate underlying cognitive mechanisms.

This hypothesis suggests a corresponding direction for artificial intelligence: a unified multimodal latent state space in which perception, language, memory, action, and world dynamics are learned through state-transition prediction.

We begin with a deliberately weak assumption about reality.

Reality can be represented, for an observer, as states that change. States can also contain distinguishable local structure and participate in larger joint states.

Let `X_t` denote the state accessible to an intelligent system at time `t`.

A state transition can be written simply as:

`X_t → X_(t+1)`

If the transition is uncertain, prediction can be expressed as:

`P(X_(t+1) | X_≤t)`

Nothing in this formulation requires an explicit causal relation.

Suppose two local states, `A` and `B`, participate in the same joint state. Instead of assuming that `A` causes a change in `B`, we can minimally describe the observed transition as:

`(A_t, B_t) → (A_(t+1), B_(t+1))`

The same idea extends to arbitrarily high-dimensional states.

A **local world** can be understood as a relatively stable and partially autonomous region of state-transition structure. Different local worlds may have different internal regularities, may combine into larger joint states, and may interact without requiring a single predefined global decomposition of reality.

The theory therefore starts with three minimal assumptions:

Causality, logic, objects, forces, and other higher-level relations are not required as primitive assumptions.

A biological nervous system is naturally suited to representing joint states.

A brain state is not a single indivisible symbol. At any moment, many neurons and neural populations are active simultaneously. A simplified neural state may therefore be represented as:

`Z_t = (z_1, z_2, ..., z_n)`

Different aspects of perception, memory, bodily condition, ongoing action, and previous experience can coexist within this distributed state.

The minimal cognitive operation proposed here is:

`Z_≤t → predicted Z_(t+1)`

or probabilistically:

`P(Z_(t+1) | Z_≤t)`

The system predicts its next joint state.

Under the strong form of this hypothesis, intelligence does not require separate primitive mechanisms for:

These may appear as useful high-level descriptions, implementation details, or emergent regularities, while the underlying operation remains joint state-transition prediction.

Action does not need to be treated as an external module attached to a world model.

The current joint state may include:

We may write:

`Z_t = (world, body, memory, language, action, goal, ...)`

The next predicted state naturally contains both a future world state and the organism's own future action:

`Z_t → Z_(t+1)`

From an external engineering perspective, this may look like:

`observation → policy → action`

but at a more unified level it can be interpreted as:

`current joint state → next joint state`

The organism's next movement is simply one component of the future state being predicted.

Planning can then be interpreted as longer-horizon prediction:

`Z_t → Z_(t+1) → Z_(t+2) → ... → Z_(t+k)`

A sequence that appears externally as a plan may arise from prediction over future joint states containing both self and environment.

The value of diverse experience is not merely that it exposes a system to more situations.

If a system only memorizes:

`X_1 → Y_1`

`X_2 → Y_2`

`X_3 → Y_3`

it has learned particular transitions but little reusable structure.

When one finite system must accurately predict a wide variety of transitions, it is pressured to compress repeated regularities into reusable latent structure.

A useful summary is:

`diverse experience → predictive compression → reusable latent transition structure`

This latent structure does not need to explicitly contain concepts such as:

It only needs to improve prediction across many situations.

Under this view, what humans later call "knowledge" may correspond to stable, generalizable predictive structure.

Language need not begin as an intentionally designed representation system.

Consider an organism interacting socially with its environment. Its experienced joint state contains:

Initially, a sound is simply another event in the world.

For example, a recurring joint state may contain:

`(water-like visual state, sound "water", bodily state, social context, ...)`

The predictive system gradually learns regularities such as:

`water-like visual state → expected sound`

and also:

`sound → expected visual / bodily / situational state`

No symbolic designer is required.

The sound gains symbolic function because it becomes reliably connected, through prediction, with other parts of the joint state.

Symbolic function is therefore proposed to emerge from stable predictive coupling.

The crucial transition occurs when a sound or other signal can partially substitute for the corresponding real-world state in prediction.

At first:

`world state A → world state B`

Then a sound becomes associated with world state `A`:

`world state A ↔ sound A`

Eventually, hearing `sound A` can activate predictions that would normally follow from directly experiencing `world state A`.

In functional terms:

`sound A ≈ predictive proxy for world state A`

The sound is not identical to the real state. It becomes useful because it can stand in for that state within the predictive process.

This creates a new possibility:

`symbol → prediction about world`

and eventually:

`symbol → symbol`

At that point, prediction no longer requires the corresponding physical event to be present at every step.

Once symbolic states become sufficiently stable, they can begin predicting other symbolic states.

The system moves through three broad stages:

`world → world prediction`

then:

`world ↔ symbol`

and finally:

The third stage is especially important.

Language becomes a predictive environment of its own.

A linguistic state can produce another linguistic state:

`L_t → L_(t+1)`

which becomes the input for another transition:

`L_(t+1) → L_(t+2)`

producing a potentially long sequence:

`L_t → L_(t+1) → L_(t+2) → ...`

This symbolic environment has unusual properties. Compared with raw sensory experience, symbols can be:

Prediction therefore gains a low-cost symbolic substrate in which increasingly long and abstract state transitions can occur.

This hypothesis makes a stronger claim than saying that language merely expresses causal and logical structures already present inside cognition.

It proposes that explicit causality and logic may arise within the dynamics of language itself.

Consider repeated linguistic transitions such as:

`because A ... therefore B`

or:

`if A, then B; A; therefore B`

Through repeated prediction and generalization, such symbolic transition patterns become stable and composable.

Humans later describe the resulting behavior using concepts such as:

But the latent predictive system does not necessarily contain separate hidden entities corresponding to "causality" or "logic."

The stronger claim is:

`logical behavior does not imply an explicit internal logic engine`

and:

`causal behavior does not imply an explicit internal causal relation`

The latent level may still consist only of high-dimensional predictive transition structure.

Under this interpretation, explicit logic and causality are stabilized forms of symbolic state transition.

Once symbols can predict symbols, language can become increasingly complex through its own continued operation.

A simple symbolic state can generate another symbolic state, which itself becomes a new state available for prediction.

For example:

`one → two → addition → multiplication → algebra → functions → calculus`

The later structures do not need to appear directly in raw sensory experience as ready-made objects.

They can emerge through repeated symbolic composition and prediction.

Language therefore becomes a kind of **secondary predictive world**.

It remains grounded, at least historically and functionally, in interaction with reality, yet it can develop structures far beyond immediate perception.

This allows prediction to operate over:

Higher intelligence may therefore depend less on the appearance of a new reasoning mechanism and more on the emergence of a new predictive state space.

Once linguistic and symbolic prediction becomes recursively composable, increasingly formal systems can emerge.

Counting stabilizes numerical symbols.

Numerical symbols participate in arithmetic.

Arithmetic supports algebra.

Further symbolic development produces geometry, calculus, formal logic, programming languages, and scientific theories.

Science then creates a feedback loop between symbolic prediction and reality:

`reality → prediction → language → symbolic prediction → prediction about reality → experiment`

The result of experiment changes later prediction and language.

Scientific concepts therefore do not need to be literal copies of reality's hidden ontology.

They can instead be treated as symbolic predictive structures constrained by reality.

A scientific theory is valuable when it successfully:

In this sense, science can be understood as language becoming increasingly formal while remaining constrained by the world.

The local-world view adds an ontological layer to the hypothesis.

Reality does not need to be assumed to consist of one globally uniform set of primitive relations.

Instead, different regions or systems may possess different local structures and transition regularities.

A local world may be represented abstractly as:

`W_i = (S_i, T_i)`

where:

`S_i` is a local state space,`T_i` describes local state transitions.
When local worlds interact, a larger joint state can form:

`(W_i, W_j) → joint state`

The resulting transition does not require us to assign a fundamental causal arrow between the two local worlds.

We may simply observe:

`joint state_t → joint state_(t+1)`

From the perspective of an observer, what later becomes described as an object, relation, force, cause, or rule may be a stable way of compressing repeated local-world interactions.

The local-world framework therefore provides a possible minimal ontology for the predictive theory:

`local worlds → joint states → joint state transitions`

Current AI systems already display separate pieces of this picture.

Large language models learn:

`P(token_(t+1) | token_≤t)`

and nevertheless exhibit behaviors associated with:

World models learn regularities such as:

`world state_t → world state_(t+1)`

Multimodal systems connect text, image, audio, video, and other modalities.

Embodied agents additionally predict or generate actions.

A natural long-term architecture suggested by this hypothesis is a unified latent state:

`Z_t = (vision, audio, 3D, language, memory, body, action, goal, ...)`

with a shared predictive objective:

`P(Z_(t+Δ) | Z_≤t)`

The same latent state would connect:

`world ↔ language`

`world ↔ action`

`language ↔ action`

`memory ↔ world`

and potentially every other modality relevant to the agent.

In such a system, language prediction and world prediction would not be fundamentally separate kinds of intelligence.

They would be different regions or projections of one evolving predictive state.

Action would also be part of the predicted future rather than a completely separate computational category.

The strongest scientific form of this proposal is:

A sufficiently capable system trained to predict diverse multimodal joint state transitions can develop perception, action, linguistic abstraction, logical reasoning, causal reasoning, and planning without requiring these abilities to be implemented as separate primitive cognitive mechanisms.

This hypothesis is stronger than the general statement that prediction is important.

It predicts that increasingly capable unified predictive systems should be able to acquire behaviors traditionally attributed to separate cognitive modules without requiring those modules to be fundamental.

The hypothesis would be weakened if some important class of cognition consistently failed to emerge from sufficiently rich predictive learning and instead required a qualitatively distinct computational mechanism.

Possible experimental comparisons could include:

The goal would not be to prove that explicit modules are never useful.

The question is whether they are **fundamental requirements for intelligence**.

The complete proposal can be compressed into the following chain:

`local worlds`

`↓`

`joint states`

`joint state transitions`

`prediction of joint state transitions`

`practical interaction with reality`

`sounds and signals become predictive proxies`

`symbols predict symbols`

`language becomes a self-extending predictive state space`

`logic, causality, mathematics, planning, and science emerge`

The entire process is hypothesized to remain grounded in one minimal mechanism:

**joint state-transition prediction**

This paper proposes a minimal predictive hypothesis of intelligence.

Reality, for an observer, presents changing local and joint states. Biological intelligence learns to predict how those joint states change. During practical interaction, sounds and other signals become reliably coupled with parts of reality and gradually acquire symbolic function. Once symbols can substitute for real states in prediction, they begin predicting one another. Language then develops into a recursively self-extending predictive environment.

Within this symbolic environment, increasingly stable and complex transition structures appear. Humans describe some of them as causality, logic, mathematics, planning, and science.

The central hypothesis is therefore:

**The latent foundation of intelligence may contain only joint state-transition prediction, while explicit causality and logic emerge later through the autonomous development of language prediction.**

This perspective suggests a unified direction for artificial intelligence as well:

**A future general intelligence may be best modeled as one multimodal predictive system operating over a shared latent state in which world, language, memory, body, and action continuously predict one another.**

This hypothesis overlaps with several existing research traditions while making a stronger reductionist claim.

The distinctive claim developed here is that the latent cognitive level need not contain hidden causal or logical structures at all. It may consist only of predictive joint-state dynamics, while explicit causality and logic arise through the emergence and autonomous development of linguistic prediction.
