cd /news/machine-learning/history-dependent-dynamics-hdd · home topics machine-learning article
[ARTICLE · art-95576] src=informationalprocessualmonism.blogspot.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

History-Dependent Dynamics (HDD)

Independent researcher Taotuner has proposed a methodological framework called History-Dependent Dynamics (HDD) to disentangle history dependence, recurrence, self-reference, and self-modeling in dynamical systems. The framework offers progressively stronger tests to classify systems based on their history-dependent organization, and is designed to be substrate-general, applicable to biochemical networks, neural circuits, animal behavior, and artificial agents. HDD is presented as a falsifiable proposal that has not yet been empirically validated.

read67 min views1 publishedAug 13, 2026
History-Dependent Dynamics (HDD)
Image: source

📄 HISTORY-DEPENDENT DYNAMICS (HDD)

A Methodological Framework for Disentangling History Dependence, Recurrence, Self-Reference, and Self-Modeling in Dynamical Systems #

Taotuner

Independent Researcher

Brazil

August 2026


Abstract #

Many physical, chemical, biological, neural, and artificial systems exhibit dynamics that depend on their previous states. Such phenomena are described using concepts including memory, adaptation, hysteresis, non-Markovianity, recurrence, and self-modeling. These concepts are related but are not equivalent, and evidence for one does not automatically establish the others.

This paper introduces History-Dependent Dynamics (HDD) , a proposed methodological framework for experimentally disentangling increasingly specific inferential claims about history-dependent organization. HDD begins with a basic empirical question: does information about the past improve out-of-sample prediction of future system behavior beyond the measured present state and current inputs? It then proposes progressively stronger tests addressing causal trajectory dependence, recurrent implementation, functional self-reference, and self-modeling.

The framework treats these as diagnostic constructs rather than developmental stages. Each construct represents a more specific hypothesis requiring additional experimental evidence. HDD is explicitly designed to address the observation-model problem, in which apparent memory may arise because relevant latent variables are not included in the measured state.

HDD classification is relative to a quadruple: (S, O, I, M) — the system, the observation model, the intervention set, and the mechanistic model class. A given system may receive different HDD profiles under different experimental specifications, which is a feature rather than a limitation.

The framework is substrate-general and can be applied to systems ranging from biochemical reaction networks and neural circuits to animal behavior and artificial agents. Existing research already provides documented examples of history-dependent dynamics across several of these domains. HDD does not reinterpret those findings as evidence for higher-order properties; instead, it identifies how existing datasets and experimental systems could be used to test progressively stronger hypotheses.

The proposed framework is not a theory of consciousness. Consciousness is treated only as a possible downstream application. In particular, HDD does not assume that history dependence, recurrence, self-reference, or self-modeling is sufficient for consciousness.

The complete framework has not yet been empirically validated. HDD is presented as a falsifiable methodological proposal whose empirical validity remains to be established through computational benchmarks and experimental studies.

Keywords: history dependence; non-Markovian dynamics; memory; recurrence; self-reference; self-modeling; dynamical systems; causal inference


1. Introduction #

Many systems cannot be adequately described by considering their current measured state alone. A neuron may respond differently depending on its recent stimulation history. An animal's behavior may depend on previous interactions. A chemical reaction network may exhibit hysteresis and adaptation. A physical system may retain information about previous interactions with its environment. An artificial agent equipped with persistent memory may behave differently after different sequences of previous events.

These phenomena are commonly described using terms such as memory, adaptation, hysteresis, non-Markovianity, recurrence, and self-modeling. The problem is that these concepts are often discussed at different explanatory levels. A system can exhibit history dependence without possessing an identifiable recurrent architecture. A recurrent system need not represent itself. A system can process information about itself without maintaining a model of its own dynamics. And even a demonstrable self-model does not, by itself, establish consciousness.

A system may remember without representing, represent without modeling, model itself without being conscious, and depend on its history without possessing an identifiable memory mechanism.

The central methodological claim of this paper is therefore:

Evidence that a system depends on its history should not automatically be interpreted as evidence for more specific forms of dynamical organization.

Stronger claims require stronger evidence.

HDD does not introduce a new estimator of temporal dependence. Its methodological contribution is the explicit separation of inferential claims that are frequently conflated in cross-disciplinary discussions: predictive history dependence, causal trajectory dependence, recurrent implementation, functional self-reference, and self-modeling. The framework specifies distinct empirical criteria and falsification conditions for each claim and proposes synthetic benchmarks designed to test whether these distinctions are identifiable.

The framework integrates questions from existing literatures—including non-Markovian dynamics, causal inference, recurrent computation, predictive state representations, and self-modeling—into a common diagnostic architecture. It does not claim to replace these literatures, but to provide a structure for distinguishing the kinds of claims they support.

The novel contribution is not any single concept—non-Markovianity, memory, recurrence, self-reference, or self-modeling—but rather the proposal of a methodological architecture that defines independent tests to prevent evidence for one property from being automatically used as evidence for another.

Let Xₜ denote the measured state of a system at time t, Pₜ its current input or perturbation, and Hₜ a bounded representation of its preceding history. The first HDD question is whether:

P(Xₜ₊₁ | Xₜ, Pₜ, Hₜ) ≠ P(Xₜ₊₁ | Xₜ, Pₜ)

If historical information improves prediction on previously unseen data, the system exhibits what HDD calls history-dependent predictive structure. This result is deliberately weak. It does not establish a dedicated memory mechanism. It does not establish recurrence, self-reference, self-modeling, agency, or consciousness. History may improve prediction simply because Hₜ contains information about variables that were not included in Xₜ.

The principal contribution of HDD is consequently not the discovery of history dependence itself. History-dependent and non-Markovian dynamics are already established areas of research. Recent work has developed methods for decomposing non-Markovian history dependence and has demonstrated such dependencies in prolonged behavioral recordings (Leighton & Lynn, 2025). Leighton and Lynn (2026) further provide a tractable tunable model of non-Markovian dynamics, showing that intuitive measures such as autocorrelation may fail to capture the relevant historical structure. The proposed contribution is to integrate these questions into a broader diagnostic framework in which predictive history dependence is experimentally separated from causal trajectory dependence, recurrent implementation, functional self-reference, and self-modeling.


2. A Preliminary Distinction: History Dependence Is Not Equivalent to Memory #

A fundamental distinction must be made explicit before proceeding.

History dependence is an observational or functional property: the past improves prediction of the future. Memory mechanism is a hypothesis about how that dependence is implemented: some variable, structure, or process retains information across time.

A system can exhibit history dependence because:

  1. a physical variable relaxes slowly;

  2. a latent state persists;

  3. a structural parameter has changed;

  4. an internal model has been updated;

  5. a feedback loop maintains information;

  6. the system represents its own state.

These are different explanations. They make different predictions. They require different evidence.

A related but distinct concept is state sufficiency: a measured state representation may be insufficient for prediction, even when the underlying system is Markovian with respect to a more complete state. This is the observation-model problem that HDD explicitly addresses.

HDD is designed to distinguish these explanations.


3. The Observation-Model Problem and State Reconstruction #

Before defining the HDD constructs, an important limitation must be made explicit.

A system can be Markovian with respect to a sufficiently complete state representation while appearing non-Markovian when the observer measures only a coarse-grained subset of its variables.

Suppose the true state is:

Zₜ = (Xₜ, Lₜ)

where Xₜ is observed and Lₜ is an unmeasured latent variable.

The complete system may satisfy:

P(Zₜ₊₁ | Zₜ)

while the observed process satisfies:

P(Xₜ₊₁ | Xₜ, Hₜ)

because the history Hₜ provides indirect information about Lₜ.

In such a case, history improves prediction without requiring a separate memory store.

This distinction is fundamental to HDD.

Therefore, a positive history-dependence result should always be described relative to the specified observation model:

HDD does not treat statistical history dependence as proof of an ontologically distinct memory mechanism.

Instead, the first empirical result is simply that the present measured state is insufficient for prediction under the chosen representation.

Recent work by Leighton and Lynn (2025) provides a direct methodological antecedent: their information-theoretic decomposition of non-Markovian history dependence shows how different temporal scales can carry different amounts of predictive information, and that intuitive measures such as autocorrelation may fail to capture the relevant structure.

3.1 State Reconstruction and Model-Relative History Irreducibility

A further distinction is necessary. Suppose we allow the model to construct a predictive state representation Sₜ = f(Hₜ) from the history. This representation is designed to capture the predictive information contained in the past.

Three cases can be distinguished:

  1. Measured state insufficiency: Xₜ is insufficient, but Sₜ derived from Hₜ is sufficient and eliminates the predictive advantage of Hₜ beyond Sₜ.

  2. Model-relative history irreducibility: Even after allowing flexible state reconstruction within a specified model class, Hₜ contains predictive information not captured by Sₜ.

  3. Observational equivalence: No finite representation from the available history can eliminate the dependence under the specified observation model.

This distinction is critical because it separates the question of whether the measured state is insufficient from whether history dependence is an irreducible property of the system under any state representation constructed from that history.

Important qualification: "Irreducible" here is relative to the specified model class and representation family. It should not be interpreted as a claim about absolute ontological irreducibility. A richer representation family or a different observational setup might reveal reducibility that was not apparent under the initial specification.


4. The HDD Diagnostic Dimensions #

HDD proposes five constructs arranged as a diagnostic profile rather than a hierarchy. Each construct represents a distinct evidential commitment:

  • Historical dependence: predictive (Construct I) and causal trajectory (Construct II)

  • Implementation: causally identified feedback recurrence (Construct III)

  • Self-related organization: functional self-reference (Construct IV) and self-modeling (Construct V)

These constructs are not assumed to correspond to ontologically distinct levels of organization. They are distinct evidential commitments defined by the type of evidence required to justify each inference.

Crucially, HDD classification is relative to a quadruple (S, O, I, M): the system, the observation model, the intervention set, and the mechanistic model class. A system may receive different HDD profiles under different experimental specifications. The ordering reflects increasing inferential specificity, not a claim that these properties form a developmental, computational, or evolutionary sequence.

HDD classifies systems according to a profile:

D = (d₁, d₂, d₃, d₄, d₅)

where dᵢ ∈ {+, −, NI, UE, NT}:

  • (+) = positive evidence for the construct

  • (−) = evidence against the construct

  • (NI) = structurally not identifiable under available observations/interventions. This should be established through formal analysis of identifiability under the specified model class, not simply through failure to find evidence.

  • (UE) = insufficient evidence / inadequate statistical power

  • (NT) = not tested / not applicable

| | I | II | III | IV | V |

| --------------------------------- | -: | -: | --: | -: | -: |

| Predictive evidence | ✓ | | | | |

| Causal intervention | | ✓ | ✓ | ✓ | ✓ |

| Mechanistic identification | | | ✓ | | |

| Self-specific representation | | | | ✓ | ✓ |

| Counterfactual own-dynamics model | | | | | ✓ |


Construct I — History-Dependent Predictive Structure

History-dependent predictive structure exists when information about preceding states improves prediction of future behavior beyond the measured present state and current inputs.

Consider two models:

M₀: Xₜ₊₁ = f(Xₜ, Pₜ)

and

M₁: Xₜ₊₁ = f(Xₜ, Pₜ, Hₜ)

Define:

ΔL = L(M₀) − L(M₁)

where L denotes an out-of-sample loss for which lower values indicate better predictive performance (e.g., MSE, log-loss), estimated with temporal cross-validation and capacity control to ensure that the improvement is not merely a consequence of increased model flexibility.

If ΔL is positive and exceeds prespecified statistical and practical significance criteria, the data support history-dependent predictive structure.

Important clarification: Construct I is defined at the level of conditional predictive dependence; out-of-sample predictive improvement is one operational test of this dependence. The two formulations are related but not identical.

Minimum requirements:

  • Out-of-sample evaluation (temporal cross-validation or held-out trajectories)

  • Complexity-matched or appropriately penalized comparison

  • Effect size and confidence intervals (not merely statistical significance)

Possible operationalizations include:

  • conditional mutual information;

  • transfer entropy;

  • Granger-style predictive comparisons;

  • nonlinear forecasting;

  • state-space models;

  • recurrent predictive models;

  • explicit memory-kernel models.

These methods should not be treated as interchangeable measures of one universal quantity. Each makes different assumptions and captures different aspects of temporal dependence. Transfer entropy, for example, is not simply a universal information-theoretic version of "history improves prediction"; different measures provide complementary operationalizations rather than interchangeable estimators.

HDD defines the construct at the level of the empirical question rather than by privileging one estimator.

Hₜ representation: HDD should evaluate history dependence across a prespecified family of history representations rather than relying on a single history encoding, as different history representations may yield different results.

History dependence profile: A valuable extension is to compute ΔL(τ) for different history horizons τ, producing a history dependence profile that shows at which temporal scales historical information is most predictive.

State reconstruction control: To distinguish measured-state insufficiency from model-relative history irreducibility, HDD should also compare against a model M_S that uses a learned predictive state representation Sₜ = f(Hₜ). If M_S eliminates the advantage of M₁, the dependence is reducible to state reconstruction.

Interpretive note: Construct I is an observation-relative criterion of history-dependent predictive structure, not a substrate-independent criterion of intrinsic non-Markovianity. It establishes predictive insufficiency of the measured state representation, not memory as an ontological mechanism. A negative classification applies only to the prespecified history family and temporal horizon; it does not rule out history dependence at unobserved scales.


Construct II — Causally Demonstrated Trajectory Dependence

Predictive dependence does not establish causation.

A stronger test attempts to manipulate the prior trajectory while holding the present state and current perturbation as constant as experimentally possible.

Consider two conditions:

Xₜᴬ ≈ Xₜᴮ

Hₜᴬ ≠ Hₜᴮ

Pₜᴬ = Pₜᴮ

where Hₜᴬ and Hₜᴮ represent different experimentally induced prior trajectories.

If the resulting future states differ systematically:

Xₜ₊₁ᴬ ≠ Xₜ₊₁ᴮ

then the evidence for causal trajectory dependence is stronger.

Important causal clarification: This design estimates a conditional effect of manipulated trajectory on future dynamics given the measured present state. It does not estimate the total effect of history, because matching on Xₜ may condition on a mediator of the historical effect. Construct II therefore tests whether trajectory history has residual effects on future dynamics that persist despite matching on the measured present state. This is a feature, not a bug: the experiment reveals whether latent states persist across matched observable states.

Even with Xₜᴬ ≈ Xₜᴮ, latent states may differ:

Zₜᴬ ≠ Zₜᴮ

In that case, the experiment demonstrates that different histories produce different latent states, not necessarily that "history itself" is a causal variable independent of state.

HDD therefore defines Construct II as:

Causally demonstrated trajectory dependence occurs when experimentally manipulated differences in prior trajectories produce systematic differences in future dynamics conditional on the measured present state and current inputs.

Construct II therefore does not imply that history constitutes a causal variable independent of state; rather, it establishes that experimentally differentiated trajectories can induce differences in future dynamics that are not captured by the measured present state. The causal pathway may be mediated by latent state variables.

State matching does not establish state identity; it establishes observational equivalence under the specified measurement model.

State matching must be formalized. A distance metric d(Xᴬ, Xᴮ) < ε should be specified, with appropriate justification for ε and the chosen metric. The feasibility of matching depends critically on the dimensionality and representation of the state space.

Minimum requirements:

  • Explicit state-matching procedure with prespecified tolerance

  • Controlled manipulation of prior trajectories (not merely observational correlation)

  • Multiple independent replicates


Construct III — Causally Identified Feedback Recurrence

History dependence does not imply recurrence. A system may depend on the past because a slow internal variable retains information about previous states. This is history dependence, but it does not establish that the relevant information is implemented through a recurrent feedback architecture.

A critical distinction is needed:

  • State persistence: a variable retains information over time: Zₜ → Zₜ₊₁

  • Causal feedback recurrence: prior activity influences subsequent activity through an identifiable causal feedback pathway in which information returns to a previously participating dynamical subsystem, as represented in a temporal causal graph.

These are not the same. A system may persist without feedback.

Important terminological clarification: "Recurrence" can refer to several distinct phenomena. In HDD, Construct III concerns a specific mechanistic hypothesis: causally identified feedback recurrence at the level of system architecture. This should not be confused with recurrence in a broader computational sense (such as state recurrence in dynamical systems or recurrent computation in neural networks), nor with functional recurrence observable from input-output behavior alone.

Evidence for recurrent implementation requires more than observing temporal dependence. A candidate recurrent pathway should be identified and experimentally perturbed. If disruption of the pathway selectively reduces the history-dependent effect while appropriate controls demonstrate that the intervention was effective, the evidence for recurrent implementation becomes stronger.

Minimum requirements:

  • Identification of a candidate feedback pathway

  • Selective perturbation of that pathway

  • Demonstration that perturbation reduces the history-dependent effect

  • Controls showing the perturbation did not simply disable general system function

Ideal (when feasible):

  • Rescue experiments that restore the effect through compensatory intervention

Important identifiability note: A feedforward architecture with delay lines may produce the same input-output behavior as a recurrent network. Therefore, recurrent implementation is not generally identifiable from observational data alone; it requires independent mechanistic evidence or intervention.

Construct III is therefore a mechanistic intervention test, not an inference of recurrence from temporal statistics. It requires an independently specified architectural or mechanistic hypothesis that can be subjected to intervention.

Accordingly, HDD does not treat recurrence as a purely computational property inferred from temporal dependence; Construct III concerns causal implementation at the level of the system architecture.

Important: Construct III depends on the granularity of the mechanistic decomposition. A system may appear recurrent under one decomposition and feedforward under another. HDD does not claim to detect recurrence absolutely; it identifies recurrence relative to a specified mechanistic model class.

Thus:

History Dependence ⇒ Recurrent Implementation is rejected.

HDD classifies a system as exhibiting Construct III only when there is positive causal evidence for a recurrent implementation, not merely when temporal dependence is observed.


Construct IV — Functional Self-Reference

Recurrence does not necessarily imply self-reference. A feedback system can recursively process information about an external environment without representing information about itself.

A critical distinction is needed:

  • Self-sensing: the system measures its own state (e.g., thermostat). The system has access to information about itself, but this information does not function as a distinct representation. Self-sensing alone is not sufficient for functional self-reference.

  • Self-representation: information about the system's own state is represented in a form that can, in principle, be manipulated independently of the represented variable. This is a stronger condition than self-sensing. However, HDD does not require that representation be a physically separable variable; distributed representations may also qualify if they satisfy the functional criteria.

  • Functional self-reference: a self-representation participates causally in determining subsequent system dynamics, and this causal role cannot be reduced to the informational content of the represented state alone. The relevant information must be represented in a manipulable form whose causal contribution depends specifically on its relation to the prespecified system boundary, rather than merely on the generic informational value of the variable.

  • Self-modeling: an internal model predicts or regulates the system's own future dynamics, including counterfactual consequences of actions.

Construct IV concerns functional self-reference.

HDD defines functional self-reference as:

Functional self-reference occurs when a representation whose content concerns the system's own state or dynamics is causally used to modify the system's dynamics in a manner specifically dependent on that content.

"Self" is an operational variable defined by the researcher, not a presumed metaphysical property. Functional self-reference is therefore a relational property between a representation and a prespecified system boundary, rather than an intrinsic semantic property of the representation. The system boundary and the criterion for self-relevance must be specified before outcome analysis and independently of the effect being tested. This prevents circularity in which any causally important variable is retrospectively designated as "self."

HDD treats self-related organization as boundary-relative rather than metaphysically intrinsic. A system may exhibit functional self-reference relative to one boundary definition but not another.

Counterfactual interchangeability criterion: A representation qualifies as functionally self-referential only if it can be manipulated independently of the represented physical state, and if such manipulation produces specific behavioral effects. That is:

do(Sₜ = s') ≠ do(Zₜ = z')

and the behavioral effect of do(Sₜ = s') cannot be fully explained by correlation with the true state.

Representation substitution test: If R_s is a representation of the system itself, substituting it with an artificially constructed representation R'_s that contains the same predictive and computational content about Z_t (matched for information content, complexity, temporal structure, and predictive capacity) but is implemented through a different channel should reveal whether the causal effect depends specifically on the functional relationship between representation and self. If:

I(R_s; Z_t) ≈ I(R'_s; Z_t)

but:

Effect(R_s) ≠ Effect(R'_s)

then the evidence for functionally self-referential processing is strengthened. This test provides evidence that the causal role depends on the specific representational relation beyond matched informational content. It is offered as an evidential-strengthening test, not as a necessary condition for Construct IV.

This definition requires an explicit self/non-self comparison.

Operational Criterion for Self-Relevance

To operationalize this construct, the following must be prespecified:

  1. System boundary: which variables are considered internal to the system

  2. External variables: which variables are considered external

  3. Self-relevance classification: how a representation is classified as self-relevant

  4. Matched-content control: what transformation preserves information content and complexity between the self-relevant and non-self-relevant conditions

A stronger experimental test compares self-relevant information with non-self information matched as closely as possible along prespecified statistical and computational dimensions.

If manipulation of self-relevant information produces effects that cannot be explained by generic information processing, the evidence for functional self-reference is strengthened.

Formally:

Δ_specific = Δ_self - Δ_matched_control

where Δ_self is the effect of manipulating self-relevant information and Δ_matched_control is the effect of manipulating matched non-self information.

The matched-content control is central to distinguishing functional self-reference from a scenario in which any informative variable causally influences the system.

Minimum requirements:

  • Identification of candidate self-relevant information (with prespecified system boundary)

  • Matched-content control

  • Demonstration of specific causal effect


Construct V — Self-Modeling

Self-modeling is a stronger claim than self-reference. A candidate self-model is an internally maintained representation that allows the system to predict or regulate its own future dynamics.

This is a predictive model of the system's own future dynamics, not merely a representation of its current state.

The distinction between Construct IV and Construct V is that Construct V specifies a model of the system's own action-conditioned dynamics, whereas Construct IV concerns the causal use of self-referential information without necessarily involving a predictive model.

HDD adopts a conception of self-modeling closely related to existing computational accounts (Kwiatkowski et al., 2022; Hu et al., 2025), in which a self-model enables an agent to predict the consequences of its own actions on its own dynamics.

Let:

Sₜˢᵉˡᶠ = G(Xₜ₋ₖ:ₜ, Eₜ₋ₖ:ₜ)

represent a candidate self-model. The crucial relation is:

Sₜˢᵉˡᶠ → Aₜ → Xₜ₊₁

where the representation influences actions or regulation that in turn affect the system's future state.

The major methodological challenge is distinguishing a self-model from an ordinary latent-state estimator. A system may internally estimate the external environment, the location of another agent, resource availability, uncertainty, task state, or hidden physical variables. None of these is necessarily a model of the system itself.

Core criterion: The distinctive feature of a self-model is not merely that it predicts the system's state, but that it predicts action-conditioned dynamics of the system:

P(Xₜ₊₁ | Xₜ, do(Aₜ = a), Sₜˢᵉˡᶠ)

should be improved by the self-model relative to matched alternatives. This distinguishes a self-model from a simple state predictor.

Distinction from system identification / forward modeling: A forward model predicts the system's dynamics. A self-model, under HDD, additionally requires that the representation is internally maintained as a model of the system's own dynamics and is causally used for predicting/regulating those dynamics. Prediction of one's own dynamics is not sufficient for self-modeling under HDD. A conventional forward model or system-identification model may satisfy V1 and V2 without constituting a self-model. Construct V requires V3: causal use of a self-relevant model representation in action selection or regulation, together with evidence that manipulating model content changes behavior in accordance with the model's encoded counterfactual predictions.

HDD therefore proposes three dimensions for evaluating self-modeling:

V1 — Self-Prediction: The model predicts future states of the system itself.

V2 — Counterfactual Self-Prediction: The model predicts the consequences of different actions on the system itself.

V3 — Causal Model Use: Manipulation of the model alters future planning or regulation of the system in a corresponding manner.

Evidence for self-modeling is strengthened when all three dimensions are satisfied.

Counterfactual prediction criterion: A candidate self-model should improve prediction of system dynamics under hypothetical interventions (do(Aₜ = a)) beyond what can be achieved by a matched latent-state model or world model. This distinguishes a generative model of the system's own dynamics from a representation that merely correlates with its future state.

Model versus implicit competence: A system may behave as if it predicts its own dynamics without maintaining a manipulable self-model. HDD requires evidence that the representation can be causally manipulated and that this manipulation alters behavior in a way consistent with the model's predictions. Implicit competence alone does not satisfy Construct V.

Causal model-use specification: Manipulating the model's content must alter behavior in a counterfactually consistent manner with the representational change. Specifically:

do(Sₜ = s')

must change Aₜ and subsequently Xₜ₊₁ in accordance with the predictions encoded by s'. This distinguishes a correlational self-model from a causally utilized self-model. The intervention must modify model content, not merely a functionally correlated state variable.

Minimum requirements:

  • Demonstration that candidate self-model improves prediction of own dynamics

  • Demonstration of counterfactual prediction capacity

  • Demonstration that the model is causally used

  • Comparison against matched latent-state models, world models, and self-state estimators


5. A Note on the Relationship Between Constructs #

The constructs form a diagnostic profile with partial logical and evidential dependencies.

A system may exhibit any combination of properties. For example:

  • A system can have history dependence without recurrence.

  • A system can have recurrence without self-reference.

  • A system can have functional self-reference without self-modeling.

  • A system can have self-modeling without conscious experience.

Self-modeling (Construct V) implies a form of self-referential processing (Construct IV) in the sense that a model of the system's own dynamics necessarily involves representation of self-relevant information. Under the operational definitions adopted here, Construct V entails a form of Construct IV because causal use of a model of the system's own dynamics necessarily involves functionally self-referential information. This is a definitional relationship, not an empirical claim that all systems exhibiting predictive self-modeling must implement the particular mechanism tested under Construct III.

Construct III is specifically about one particular hypothesis of implementation—causally identified feedback recurrence—and should not be confused with recurrence in any broader functional sense. A system may implement history dependence through other mechanisms (delay lines, persistent latent states, etc.) that do not qualify as Construct III.

The constructs are not developmental stages, and their empirical implementation may vary across systems.

HDD therefore treats the five constructs as a diagnostic profile rather than a ladder:

  • (+, -, -, -, -) = history-dependent system without demonstrated recurrence, self-reference, or self-modeling

  • (+, +, -, -, -) = history-dependent system with causal trajectory dependence but no demonstrated recurrent implementation

  • (+, +, +, -, -) = recurrent history-dependent system without demonstrated self-reference

  • (+, +, -, +, +) = self-modeling system implemented without the particular recurrent mechanism tested by Construct III


6. Evidential Structure Summary #

| Construct | Empirical question | Positive interpretation | Main alternative explanation |

| --------- | ------------------- | ----------------------- | ---------------------------- |

| I | Does history improve out-of-sample prediction? | Predictive history dependence | Omitted latent state |

| II | Does manipulated trajectory affect future conditional on measured state? | Residual causal trajectory dependence | State mismatch / latent mediation |

| III | Does feedback intervention selectively change the effect? | Causal recurrent implementation | Delay/feedforward equivalence |

| IV | Does self-relevant representation have specific causal role? | Functional self-reference | Generic information processing |

| V | Does representation predict/control own action-conditioned dynamics? | Self-modeling | Forward model / system identification |


7. Null Hypotheses and Effect Sizes #

For each construct, a specific null hypothesis should be defined:

| Construct | Null Hypothesis |

| --------- | --------------- |

| I | History does not improve out-of-sample prediction after controlling for state, input, and model capacity. |

| II | Manipulated trajectory has no effect conditional on measured present state. |

| III | Disrupting candidate feedback does not selectively reduce the history-dependent effect. |

| IV | Self-relevant information has no effect beyond matched non-self information. |

| V | Candidate self-model does not improve counterfactual prediction/regulation of own dynamics beyond matched latent-state models. |

Methodology should prioritize effect sizes:

  • out-of-sample log loss

  • mutual information

  • conditional mutual information

  • predictive R²

  • causal effect size

  • intervention effect

  • counterfactual prediction error

  • model ablation effect

  • robustness across history windows

Confidence intervals, permutation tests, or bootstrap procedures should be reported.

Multiple comparison correction: When testing across multiple history windows or representations, appropriate correction (e.g., false discovery rate) should be applied.


8. Identifiability and Evidential Sufficiency #

Each HDD construct has specific alternative explanations that must be experimentally discriminated:

| Construct | Main alternative explanation |

| --------- | ---------------------------- |

| I | omitted latent state |

| II | residual state mismatch / latent mediation |

| III | observational equivalence / feedforward delay |

| IV | generic self-information processing |

| V | latent-state estimator / world model / forward model / system identification |

HDD does not require that alternative explanations be metaphysically impossible. It requires that they be experimentally discriminated within the specified model class.

Classification principle: HDD classification uses a five-way decision:

  • (+) = positive evidence for the construct

  • (−) = evidence against the construct

  • (NI) = structurally not identifiable under the available observations/interventions

  • (UE) = insufficient evidence / inadequate statistical power

  • (NT) = not tested / not applicable

(NI) differs from (UE): NI means that even with infinite data under the specified experimental design, two hypotheses remain observationally equivalent. UE means the data are insufficient to reach a conclusion. NT means the relevant construct was not investigated under the experimental protocol.

NI is a property of the identification problem under a specified observation/intervention model, not simply a failure to obtain statistical significance. NI should be established through formal analysis of identifiability under the specified model class, not through failure to find evidence.


9. The Synthetic Benchmark (Proposed) #

A critical component of HDD is computational validation. Before applying the framework to empirical systems, it should be tested on synthetic systems whose architectures are known.

9.1 Functional Properties Under Benchmark Protocol

Rather than treating architectural properties as if they automatically determine functional properties, the benchmark distinguishes between architectural implementation and observable functional properties under the specified observation and intervention protocol.

Functional Properties

| System | History Dependence | Causal Trajectory | Self-Information | Predictive Self-Model | Causal Model Use |

| ------ | :---: | :---: | :---: | :---: | :---: |

| Fully observed Markov | − | − | − | − | − |

| Hidden-state | + | ? | − | − | − |

| Delay line | + | + | − | − | − |

| Recurrent non-self | + | + | − | − | − |

| Self-state estimator | + | + | + | − | − |

| Passive self-model | + | + | + | + | − |

| Causally used self-model | + | + | + | + | + |

| World model containing self | + | + | + | ? | +/− |

Architectural Properties

| System | Latent State | Feedback | Delay Line | Recurrent Architecture | Explicit Self-Representation |

| ------ | :---: | :---: | :---: | :---: | :---: |

| Fully observed Markov | − | − | − | − | − |

| Hidden-state | + | − | − | − | − |

| Delay line | + | − | + | − | − |

| Recurrent non-self | + | + | − | + | − |

| Self-state estimator | + | +/− | − | − | + |

| Passive self-model | + | +/− | − | − | + |

| Causally used self-model | + | +/− | − | − | + |

| World model containing self | + | +/− | − | − | +/− |

Notes on interpretation:

  • + indicates the property is functionally present

  • indicates it is intentionally absent

  • +/- indicates it may vary by implementation

  • ? indicates the architecture does not determine the property

  • "World model containing self" may or may not constitute a self-model depending on whether the representation is used for predicting/regulating the system's own dynamics

9.2 Critical Test Pairs

The benchmark should specifically test whether HDD can distinguish systems that are behaviorally equivalent but mechanistically different:

  • Pair 1: Latent-state model vs. explicit memory model (same observable distribution)

  • Pair 2: Recurrent model vs. feedforward delay-line model (same input-output function)

  • Pair 3: Self-state estimator vs. self-model (same predictive performance)

  • Pair 4: Self-model vs. world model containing self (same behavioral capacity)

  • Pair 5: Passive self-model vs. causally used self-model (same internal representation, different causality)

  • Pair 6: Apparent self-model vs. predictive shortcut — a system receives a variable correlated with its own future and uses it for prediction, without possessing a causally structured self-representation (adversarial shortcut model)

  • Pair 7: Recurrence without history dependence in measured X — a system with feedback architecture but where Xₜ is sufficient for prediction

  • Pair 8: Self-reference without long-term history dependence — a system represents and uses its own state causally without extended temporal dependence

  • Pair 9: Distributed self-representation — a system where self-relevant information is encoded in a distributed manner, not as a separable variable

  • Pair 10: Semantic labeling confound — two systems with identical behavior and causal structure, where one variable is labeled "self-state" and the other "internal state"; HDD should produce the same classification

9.3 Non-Identifiability

In some cases, two architectures may be observationally equivalent under all available experiments. The benchmark should explicitly test whether HDD can correctly report non-identifiability rather than forcing a positive or negative classification:

  • Pair 11: Two distinct architectures that produce the same observable distribution P(X₁:ₜ | P₁:ₜ) for all considered interventions. The correct HDD output should be "not identifiable," not "+" or "−."

This distinguishes the framework from one that simply assigns a label to every system regardless of identifiability.

9.4 Blind Testing

Ideally, researchers should receive only the observed time series (Xₜ, Pₜ) without knowledge of the ground-truth architecture, apply HDD, and only then compare inferences against the known architecture.

9.5 Construct Discrimination Metrics

Rather than a single accuracy score, the benchmark should report:

  • Sensitivity (true positive rate) for each construct

  • Specificity (true negative rate) for each construct

  • False positive rate for each construct

  • False negative rate for each construct

  • Confusion matrix across constructs

  • Calibration of confidence

  • Non-identifiability detection rate

This is essential because a framework that rarely assigns higher constructs may appear accurate simply by being conservative.


10. Failure Modes and Confounds #

Several specific failure modes must be considered when applying HDD.

A. State aliasing: Different real states appear as the same Xₜ, making history appear predictive when it is actually revealing unobserved state differences.

B. Common-cause history: History and future are correlated because both depend on a third variable.

C. Selection bias: Only certain trajectories are observed, making history appear more predictive than it actually is.

D. History representation bias: The effect appears only because a particular representation of H was chosen.

E. Model capacity confound: M₁ is simply more powerful than M₀, and the improvement reflects capacity rather than historical information.

F. Intervention-induced state change: The intervention on "history" inadvertently changes the present state, confounding the causal inference.

G. Self-description confound: In artificial systems, the system produces language about itself without possessing a corresponding functional mechanism.

H. High-dimensional state matching: Matching present states becomes increasingly difficult as state dimensionality increases. Approximate matching may therefore introduce residual state imbalance, requiring dimensionality reduction, propensity-based methods, optimal transport, or other explicitly validated matching procedures.

I. Functional equivalence: Two systems may produce exactly the same observable function through different mechanisms. Behavioral evidence alone cannot identify implementation. This is particularly important for Construct III, where a feedforward architecture with delay lines may produce the same input-output behavior as a recurrent network.

J. Leakage: Information from the future may leak into the construction of Hₜ.

K. Nonstationarity: Slow changes in system parameters may appear as memory.

L. Distribution shift: The historical model may appear superior only because the train/test split does not respect temporal structure.

M. Intervention-history interaction: The effect of history may depend on the type of current intervention:

P(Xₜ₊₁ | Hₜ, Pₜ)

may exhibit interaction (H × P). This is particularly important for Construct II.

N. Boundary-relative selfhood: Since "self" is defined relative to a researcher-specified system boundary, Constructs IV and V inherit this boundary-dependence. HDD does not claim to discover an ontologically privileged self; it tests causal relations relative to a prespecified boundary.

O. Representation learning confound: If the model learns a predictive state representation Sₜ = f(Hₜ), the comparison may favor M₁ simply because it learns a better state representation. This does not necessarily indicate history-specific information.

P. Semantic labeling confound: The classification of a variable as "self-relevant" depends on the researcher's labeling; HDD requires that the system boundary and self-relevance criterion be specified independently of the outcome.

These failure modes must be explicitly controlled in any HDD experiment.


11. Relation to Existing Frameworks #

HDD is not intended to replace existing theories of memory, dynamical systems, predictive processing, recurrence, self-modeling, or consciousness. Its intended role is methodological.

Non-Markovian dynamics provides quantitative tools for characterizing history dependence. Leighton and Lynn (2025) developed information-theoretic approaches to decomposing non-Markovian history dependence and demonstrated such dependencies in behavioral recordings. Leighton and Lynn (2026) further provide a tractable tunable model. HDD asks what additional evidence is needed to move from predictive history dependence to causal, mechanistic, and self-referential claims.

Computational mechanics (Shalizi & Crutchfield, 2001) addresses the problem of finding minimal predictive state representations. The causal states of computational mechanics provide a formal solution to part of what HDD identifies as the observation-model problem: finding a representation that is sufficient for prediction.

Predictive state representations (Littman, Sutton & Singh, 2002) represent the state of a system through predictions of future observations conditioned on actions, rather than assuming a conventional hidden state. This provides another methodological antecedent for Construct I and for the state reconstruction distinction.

Causal representation learning and causal discovery in dynamical systems (Panayiotou & Şimşek, 2026) emphasize that interventions are crucial and that identification can fail with hidden factors or complex dynamical structures. This is directly relevant to Construct II, though the specific results depend on structural assumptions.

Self-model theories (Limanowski & Blankenburg, 2013; Metzinger, 2003; Kwiatkowski et al., 2022; Hu et al., 2025) have developed accounts of self-representation. HDD asks how self-modeling can be distinguished from latent-state estimation. Kwiatkowski et al. explicitly define self-modeling as an agent learning a predictive model of its own dynamics — a definition closely aligned with HDD's Construct V. Hu et al. (2025) provide a concrete robotic implementation of egocentric visual self-modeling for dynamics prediction and adaptation. Limanowski and Blankenburg (2013) provide a conceptual antecedent in the context of minimal phenomenal selfhood and the free-energy principle. Metzinger (2003) provides a philosophical account of phenomenal selfhood, which is a different level of analysis than HDD's functional/mechanistic operationalization.

Consciousness theories have been compared through adversarial collaborations (Cogitate Consortium, 2025). HDD adopts the broader methodological principle exemplified by adversarial theory testing: theoretical constructs should generate experimentally discriminable predictions rather than being inferred retrospectively from compatible observations. Its contribution is narrower: if a theory invokes self-reference, recurrence, or self-modeling, those properties should be operationalized separately.


12. Potential Empirical Domains #

The following cases demonstrate how published findings can be mapped onto the proposed framework and reveal where stronger tests would be required. These are candidate domains for HDD evaluation, not empirical validations of the complete framework.

12.1 Ion Channels and Cellular Biophysics

History-dependent dynamics have been investigated in models of ion channels. Soudry and Meir (2010) showed how observed history-dependent relaxation can arise from complex internal state structure that is not directly observable.

This provides an important methodological lesson: an observed history effect does not automatically identify the mechanism responsible for that effect. Under HDD, such a system may satisfy Construct I and potentially Construct II, but Constructs IV and V have no basis for inference.

A system can possess rich history-dependent dynamics without anything resembling self-modeling.

12.2 Microbial and Cellular Systems

Microorganisms provide an especially useful domain because history dependence is experimentally accessible and mechanistically diverse. Previous environmental exposure can alter subsequent cellular responses, including adaptation to recurring environmental conditions (Vermeersch et al., 2022). Reviews of microbial history-dependent behavior describe metabolic, epigenetic, biochemical, and ecological mechanisms.

Under HDD, microbial history-dependent behavior provides a candidate Construct I target. A more demanding experiment could manipulate environmental histories while matching current cellular states as closely as possible.

12.3 Chemical Reaction Networks

History dependence is not restricted to biological systems. Experimental work on chemical reaction networks has demonstrated dynamic switching and related history-sensitive behaviors in out-of-equilibrium molecular systems (Kriukov, Koyuncu & Wong, 2022).

Chemical systems provide a valuable test for the substrate-generality claim. A chemical system could exhibit history dependence without any plausible interpretation involving self-modeling.

12.4 Animal Behavior and Biological Systems

Leighton and Lynn (2025) developed an information-theoretic approach to decomposing non-Markovian history dependence and applied it to prolonged recordings of fly behavior. Their results showed historical dependencies across timescales ranging from fractions of a second to minutes. The work of Leighton and Lynn (2026) provides a tunable model for studying such dependencies.

This is a particularly direct antecedent for HDD Construct I. HDD would not reinterpret these findings as evidence of self-modeling. Instead, the proposed next questions would be: Can the relevant history dependence be causally manipulated? Can its mechanistic implementation be identified? Does it depend on recurrent biological circuitry? Is any component specifically self-referential? Does any self-relevant representation function as a model?

12.5 Open Quantum Systems

Non-Markovian dynamics are extensively studied in open quantum systems. In these systems, the future evolution of a subsystem can depend on its prior interaction with an environment (Breuer et al., 2016; Rivas et al., 2014).

This literature provides a conceptual boundary case: quantum non-Markovianity is related to, but not identical with, HDD Construct I. A quantum system can display strong memory effects without possessing anything resembling a self-model. HDD does not propose a new definition of quantum non-Markovianity and does not assume that quantum memory measures are equivalent to its functional construct.

12.6 Neural Systems

Neural systems provide a more complex application because they simultaneously contain recurrent connectivity, adaptation, short- and long-term plasticity, latent physiological states, self-related representations, and behavioral feedback.

For example: Does recent neural history improve prediction of subsequent neural or behavioral states? Can stimulation history be manipulated while current neural state is matched? Does disrupting candidate recurrent circuitry reduce the effect? Does information concerning the organism's own neural or bodily state exert a distinct causal influence? Does an internal representation of the organism's own state predict or regulate subsequent states?

12.7 Artificial Agents and Large Language Models

Artificial systems provide an especially useful test case because aspects of their memory and computational architecture can often be manipulated directly. Large language model agents equipped with external memory provide a particularly accessible example.

Xiong et al. (2026) studied how memory addition and deletion affect LLM-agent behavior. Their experiments found that retrieved prior experiences—a phenomenon they term "experience-following"—influence subsequent outputs. These findings are relevant to the motivation for HDD Construct I, but insufficient by themselves to establish either Construct I under the matched-state criterion or Construct II. The observation-model problem is particularly acute here: if the memory bank is part of the measured state, history dependence may be reduced or eliminated. The existence of an external memory store does not itself establish HDD Construct I; if the memory contents are included in the measured state, the system may be Markovian relative to that expanded state representation.

Recent work on egocentric visual self-modeling in robots (Hu et al., 2025) provides a closer antecedent for Construct V, demonstrating a predictive model of the robot's own dynamics used for adaptation. This serves as a candidate system for testing V.

These findings do not establish self-modeling in LLM agents. An agent may store a previous interaction, retrieve it later, and alter its response because of it, without maintaining a representation of itself. The distinction can be experimentally tested.

The HDD prediction is not that an LLM agent with memory must be conscious. The narrower prediction is that if a genuine self-model exists, manipulating self-relevant representations should produce causal effects that cannot be reduced to the generic predictive benefit of memory retrieval.


13. What HDD Does and Does Not Claim #

HDD does claim:

  • that history dependence should be distinguished from stronger dynamical properties;

  • that predictive and causal evidence should be separated;

  • that recurrence should be tested independently;

  • that self-reference should not be inferred from recurrence alone;

  • that self-modeling should be distinguished from generic latent-state estimation;

  • that these distinctions can be organized into a common experimental program;

  • that existing empirical literatures provide multiple domains in which parts of the framework can be tested.

HDD does not claim:

  • that history dependence is a new discovery;

  • that non-Markovianity is equivalent to memory in every ontological sense;

  • that recurrence implies self-reference;

  • that self-reference implies self-modeling;

  • that self-modeling implies consciousness;

  • that artificial systems are conscious;

  • that biological and artificial implementations are mechanistically identical;

  • that the framework has already been empirically validated as a whole.


14. Limitations #

14.1 Representation Dependence

History dependence is relative to the chosen observation model. A richer state representation may eliminate apparent history dependence.

14.2 Hidden Variables

Historical information may improve prediction because it contains information about unmeasured latent states. This is perhaps the most important limitation of Construct I.

14.3 Estimator Dependence

Different estimators capture different aspects of temporal dependence. No single statistic should be treated as the definitive measure of HDD.

14.4 Finite Data

Long histories and high-dimensional states can require large datasets.

14.5 Mechanistic Ambiguity

Predictive history dependence does not identify the mechanism responsible for it. Even causal history dependence may be mediated by mechanisms that remain unknown.

14.6 Cross-Domain Comparability

Substrate generality does not imply mechanistic identity. History dependence in different substrates may have entirely different physical implementations.

14.7 Self-Modeling Is Difficult to Operationalize

The distinction between a self-model and a sophisticated latent-state representation remains one of the most challenging aspects of the framework.

14.8 High-Dimensional State Matching

Matching present states becomes increasingly difficult as state dimensionality increases. Approximate matching may therefore introduce residual state imbalance.

14.9 Identifiability

Different mechanisms can generate observationally equivalent trajectories. This is a fundamental limitation of any framework that infers mechanisms from observables.

14.10 Representation Semantics

Particularly for Constructs IV and V, "self-relevance" may depend on the representational vocabulary used by the experimenter.

14.11 Mechanism–Function Underdetermination

A function may be realized by multiple architectures. This is especially important for Construct III.

14.12 Consciousness

HDD does not solve the hard problem of consciousness. Nor does it establish that any specific dynamical property is sufficient for phenomenal experience.

14.13 Boundary-Relative Selfhood

Since Constructs IV and V depend on a researcher-specified system boundary, they are boundary-relative. HDD does not claim to discover an ontologically privileged self; it tests causal relations relative to a prespecified boundary.

14.14 Substrate Dependence of V

Construct V requires manipulation of representations, which may be difficult or impossible in some substrates. HDD is substrate-general in terms of the questions it asks, but not method-general in terms of the interventions it requires.


15. Revision and Falsification Criteria #

HDD adopts an explicit revision principle. No construct should be protected from negative evidence.

A construct should be revised or removed if:

  1. it cannot be operationalized independently;

  2. it repeatedly fails appropriate controls;

  3. it cannot generate discriminating predictions;

  4. it collapses empirically into another construct without explanatory loss;

  5. its measurements cannot be reproduced;

  6. its apparent effects are consistently explained by simpler alternatives.

A construct is considered empirically non-discriminating if its performance does not exceed a preregistered baseline across independently generated datasets or architectures.

If computational validation shows that HDD cannot distinguish known architectures, the framework should be revised before being applied to empirical systems.

If self-modeling cannot be distinguished from generic latent-state estimation, the definition of self-modeling should be revised or the construct removed.

If self-modeling is empirically supported but has no relationship to independently operationalized conscious states, the consciousness extension should be weakened without invalidating the preceding empirical findings.

Thus:

A failed extension should make the framework smaller, not more elaborate.


16. Discussion #

The central problem addressed by HDD is not the existence of memory. Memory-like and history-dependent effects are already pervasive across scientific disciplines. The problem is inferential granularity.

A researcher observes that the past matters and may then describe the system as having memory. The system is recurrent, and the researcher may infer self-reference. The system processes information about itself, and the researcher may infer self-modeling. The system has a self-model, and the researcher may infer consciousness.

Each inference may be reasonable in a particular theoretical context. But none follows automatically from the previous observation. HDD proposes that the inferential chain should instead be experimentally decomposed.

This is particularly valuable because different scientific fields have developed different vocabularies for closely related phenomena. Physics speaks of non-Markovianity and memory kernels. Chemistry may emphasize hysteresis, adaptation, and reaction-network dynamics. Neuroscience studies recurrence, temporal integration, adaptation, and self-related processing. Artificial intelligence studies persistent memory, recurrent computation, internal representations, and agent architectures. Consciousness research studies recurrent processing, global availability, integrated information, self-representation, and metacognition.

These traditions need not be collapsed into one theory. However, some of their empirical questions can potentially be placed within a common methodological structure. That is the intended role of HDD.

The most important potential contribution of HDD is therefore not a new measure of memory. It is a proposed method for preventing category errors between different levels of dynamical organization. The same temporal pattern may have very different explanations. The purpose of HDD is to turn these differences into experimental questions.

HDD therefore treats history dependence as an empirical starting point, not as an explanatory endpoint.


17. Conclusion #

History dependence is widespread across physical, chemical, biological, neural, and artificial systems. The scientific challenge is not to establish that history can matter. The challenge is to determine what kind of historical dependence is present, how it is implemented, and what additional properties can legitimately be inferred from it.

HDD proposes a framework for addressing this problem. It begins with the weakest claim: historical information improves out-of-sample prediction relative to the specified observed-state representation. It then asks increasingly specific questions: Is history causally effective? Does recurrence implement the effect? Does the system process information about itself? Does that information function as a self-model?

Each transition requires additional evidence. This prevents the inferential shortcut:

History Dependence ⇒ Self-Reference ⇒ Self-Modeling ⇒ Consciousness

The primary contribution of HDD is therefore methodological. It proposes an experimental architecture for determining what can and cannot be inferred from history-dependent behavior.

The framework is deliberately conservative at its empirical core. Existing research already demonstrates history-dependent dynamics in biological behavior, ion-channel systems, chemical networks, quantum systems, and artificial agents. The proposed contribution is to place these phenomena within a common structure and ask what additional evidence would be necessary to move from one explanatory level to the next.

The relationship between this structure and consciousness remains an open question. HDD does not assume that consciousness requires self-modeling, nor that self-modeling is sufficient for consciousness. Instead, it proposes that if such a relationship exists, it should eventually be demonstrated through experimentally discriminable changes in dynamical organization rather than inferred from terminology or behavioral analogy.

Ultimately, the framework should be judged empirically. If computational benchmarks and experimental studies show that HDD produces reliable distinctions that existing approaches fail to provide, its methodological value will be supported. If its constructs collapse into existing measures without providing additional explanatory or predictive value, the framework should be reduced. If some constructs fail while lower-order constructs remain useful, the framework should retain the empirically successful components and abandon the unsupported extensions.

The appropriate outcome is therefore not that HDD survives at all costs. The appropriate outcome is that the evidence determines how much of HDD remains necessary.


Acknowledgments #

This manuscript was developed with the assistance of AI-based language tools, including large language models, used for literature exploration, structural organization, drafting, critical discussion, methodological critique, and language refinement. All conceptual decisions, methodological commitments, interpretation of evidence, revisions, and responsibility for the final work remain with the author.

AI-based tools did not serve as authors of the work and were not assigned responsibility for its scientific claims. All factual claims, references, methodological arguments, and citations were reviewed by the author.


References #

Breuer, H.-P., Laine, E.-M., Piilo, J., & Vacchini, B. (2016). Colloquium: Non-Markovian dynamics in open quantum systems. Reviews of Modern Physics, 88, 021002.

Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U., Henin, S., Hirschhorn, R., Khalaf, A., Lepauvre, A., et al. (2025). Adversarial testing of global neuronal workspace and integrated information theories of consciousness. Nature, 642, 133–142. https://doi.org/10.1038/s41586-025-08888-1.

Hu, Y., Chen, J., & Lipson, H. (2025). Egocentric visual self-modeling for autonomous robot dynamics prediction and adaptation. npj Robotics, 3, 14. https://doi.org/10.1038/s44182-025-00031-6.

Kriukov, D. V., Koyuncu, A. H., & Wong, A. S. Y. (2022). History dependence in a chemical reaction network enables dynamic switching. Small. https://doi.org/10.1002/smll.202107523.

Kwiatkowski, R., et al. (2022). On the origins of self-modeling. arXiv preprint arXiv:2209.02010.

Leighton, M. P., & Lynn, C. W. (2025). Decomposing non-Markovian history dependence. arXiv preprint arXiv:2512.13933.

Leighton, M. P., & Lynn, C. W. (2026). Tractable model for tunable non-Markovian dynamics. Physical Review E. https://doi.org/10.1103/d7yp-496b.

Limanowski, J., & Blankenburg, F. (2013). Minimal self-models and the free energy principle. Frontiers in Human Neuroscience, 7, 547.

Littman, M. L., Sutton, R. S., & Singh, S. (2002). Predictive representations of state. In Advances in Neural Information Processing Systems.

Metzinger, T. (2003). Being No One: The Self-Model Theory of Subjectivity. MIT Press.

Panayiotou, P., & Şimşek, Ö. (2026). Causal discovery in action: Learning chain-reaction mechanisms from interventions. Proceedings of Machine Learning Research, 323, 1545–1571.

Rivas, Á., Huelga, S. F., & Plenio, M. B. (2014). Quantum non-Markovianity: Characterization, quantification and detection. Reports on Progress in Physics, 77(9), 094001.

Shalizi, C. R., & Crutchfield, J. P. (2001). Computational mechanics: Pattern and prediction, structure and simplicity. Journal of Statistical Physics, 104(3-4), 817–879.

Soudry, D., & Meir, R. (2010). History-dependent dynamics in a generic model of ion channels: An analytic study. Frontiers in Computational Neuroscience, 4, 3.

Vermeersch, L., Cool, L., Gorkovskiy, A., Voordeckers, K., Wenseleers, T., & Verstrepen, K. J. (2022). Do microbes have a memory? History-dependent behavior in the adaptation to variable environments. Frontiers in Microbiology, 13, 1004488.

Xiong, Z., Lin, Y., Xie, W., He, P., Liu, Z., Tang, J., Lakkaraju, H., & Xiang, Z. (2026). How memory management impacts LLM agents: An empirical study of experience-following behavior. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 623–645.


Version: 2.0.0

Date: August 2026

License: CC BY 4.0

Repository: Upon publication, code and supplementary materials will be made available at the author's Zenodo repository.

CÓDIGO:

Para testar a operacionalização do Construct I, implementamos um benchmark 
sintético com quatro sistemas de ground truth conhecido: Markov, Hidden State, 
Delay Line e Recurrent.

Os resultados mostram que o Construct I classifica corretamente todos os 
sistemas (100% de acurácia), com controles de capacidade de modelo e 
robustez a diferentes níveis de ruído e sementes aleatórias.

**Interpretação:** A história melhora a previsão em sistemas não-Markovianos, 
incluindo aqueles com estado latente (Hidden State) e memória feedforward 
(Delay Line). Isso valida a tese central do HDD: história preditiva não é 
automaticamente evidência de um mecanismo específico.

**Limitação:** Este benchmark valida apenas o Construct I. A validação 
completa dos Constructs II–V é trabalho futuro.

CÓDIGO:

#

#

#

#

import os

import sys

import json

import math

import time

import platform

import warnings

from dataclasses import dataclass, asdict

import numpy as np

import pandas as pd

import matplotlib.pyplot as plt

from sklearn.linear_model import Ridge

from sklearn.preprocessing import StandardScaler

from sklearn.pipeline import Pipeline

from sklearn.decomposition import PCA

warnings.filterwarnings("ignore")

print("=" * 70)

print("HDD BENCHMARK v1.0")

print("Construct I — History-Dependent Predictive Structure")

print("=" * 70)

@dataclass

class Config:

n_train_trajectories: int = 120

n_test_trajectories: int = 60

trajectory_length: int = 500

burn_in: int = 100

max_history: int = 10

noise_std: float = 0.10

alpha_grid: tuple = (

1e-4,

3e-4,

1e-3,

3e-3,

1e-2,

3e-2,

1e-1,

3e-1,

1.0,

3.0,

10.0,

30.0,

100.0

)

inner_folds: int = 5

bootstrap_reps: int = 2000

permutation_reps: int = 5000

alpha_stat: float = 0.05

practical_relative_threshold: float = 0.01

random_seed: int = 42

robustness_seeds: tuple = (

42,

123,

456,

789,

2026

)

noise_levels: tuple = (

0.05,

0.10,

0.20,

0.30

)

history_windows: tuple = (

1,

2,

3,

5,

10

)

input_std: float = 0.50

output_dir: str = "hdd_benchmark_v1"

CFG = Config()

os.makedirs(CFG.output_dir, exist_ok=True)

np.random.seed(CFG.random_seed)

print("\nCONFIGURAÇÃO")

print("-" * 70)

for k, v in asdict(CFG).items():

print(f"{k}: {v}")

ENVIRONMENT = {

"python": sys.version,

"platform": platform.platform(),

"numpy": np.version,

"pandas": pd.version,

"sklearn": import("sklearn").version,

"timestamp": time.strftime("%Y-%m-%d %H:%M:%S"),

}

with open(

os.path.join(CFG.output_dir, "environment.json"),

"w"

) as f:

json.dump(ENVIRONMENT, f, indent=2, default=str)

print("\nAMBIENTE:")

print(json.dumps(ENVIRONMENT, indent=2))

def generate_markov(

length,

burn_in,

noise_std,

rng

):

"""

Sistema Markoviano observado.

X_{t+1} = 0.70 X_t + 0.40 P_t + noise

Condicionado em X_t e P_t,

a história adicional não deve conter

informação sistemática relevante.

"""

total = length + burn_in

x = np.zeros(total)

p = rng.normal(

0,

CFG.input_std,

total

)

x[0] = rng.normal()

for t in range(total - 1):

x[t + 1] = (

0.70 * x[t]

  • 0.40 * p[t]

  • rng.normal(0, noise_std)

)

return x[burn_in:], p[burn_in:]

def generate_hidden_state(

length,

burn_in,

noise_std,

rng

):

"""

Sistema com estado latente.

z_{t+1} = 0.85 z_t + input + noise

X_t = z_t + observation noise

A história de X contém informação sobre

o estado latente que não está completamente

disponível em X_t.

"""

total = length + burn_in

z = np.zeros(total)

x = np.zeros(total)

p = rng.normal(

0,

CFG.input_std,

total

)

z[0] = rng.normal()

for t in range(total - 1):

z[t + 1] = (

0.85 * z[t]

  • 0.35 * p[t]

  • rng.normal(0, noise_std)

)

x[t] = (

z[t]

  • rng.normal(0, noise_std)

)

x[-1] = (

z[-1]

  • rng.normal(0, noise_std)

)

return x[burn_in:], p[burn_in:]

def generate_delay_line(

length,

burn_in,

noise_std,

rng

):

"""

Sistema com dependência explícita de atraso.

X_{t+1} depende de:

X_t

X_{t-3}

P_t

Isso gera dependência histórica observável.

"""

total = length + burn_in + 5

x = np.zeros(total)

p = rng.normal(

0,

CFG.input_std,

total

)

x[:5] = rng.normal(

0,

0.5,

5

)

for t in range(4, total - 1):

x[t + 1] = (

0.55 * x[t]

  • 0.30 * x[t - 3]

  • 0.30 * p[t]

  • rng.normal(0, noise_std)

)

start = burn_in + 5

return (

x[start:start + length],

p[start:start + length]

)

def generate_recurrent(

length,

burn_in,

noise_std,

rng

):

"""

Sistema com dinâmica interna recorrente.

Estado oculto:

h_{t+1} = tanh(A h_t + B P_t + noise)

Observação:

X_t = C h_t + observation noise

A história de X pode recuperar informação

sobre o estado recorrente oculto.

"""

total = length + burn_in

h = np.zeros((total, 2))

x = np.zeros(total)

p = rng.normal(

0,

CFG.input_std,

total

)

A = np.array([

[0.75, 0.20],

[-0.15, 0.70]

])

B = np.array([

[0.35],

[0.20]

])

C = np.array([

0.80,

0.50

])

h[0] = rng.normal(

0,

0.5,

2

)

for t in range(total - 1):

noise_vec = rng.normal(

0,

noise_std,

2

)

h[t + 1] = np.tanh(

A @ h[t]

  • B[:, 0] * p[t]

  • noise_vec

)

x[t] = (

C @ h[t]

  • rng.normal(0, noise_std)

)

x[-1] = (

C @ h[-1]

  • rng.normal(0, noise_std)

)

return x[burn_in:], p[burn_in:]

SYSTEM_GENERATORS = {

"Markov": generate_markov,

"Hidden State": generate_hidden_state,

"Delay Line": generate_delay_line,

"Recurrent": generate_recurrent,

}

def generate_dataset(

system_name,

n_trajectories,

length,

burn_in,

noise_std,

seed

):

rng = np.random.default_rng(seed)

trajectories = []

generator = SYSTEM_GENERATORS[system_name]

for i in range(n_trajectories):

x, p = generator(

length=length,

burn_in=burn_in,

noise_std=noise_std,

rng=rng

)

trajectories.append({

"trajectory_id": i,

"x": np.asarray(x, dtype=float),

"p": np.asarray(p, dtype=float)

})

return trajectories

def make_supervised_dataset(

trajectories,

history_window

):

"""

Constrói:

target:

X_{t+1}

estado presente:

X_t

P_t

história:

X_{t-1}, ..., X_{t-history_window}

P_{t-1}, ..., P_{t-history_window}

IMPORTANTE:

nenhuma informação futura entra nas features.

"""

X_present = []

X_history = []

y = []

trajectory_ids = []

for tr in trajectories:

x = tr["x"]

p = tr["p"]

tid = tr["trajectory_id"]

start = history_window

end = len(x) - 1

for t in range(start, end):

present = [

x[t],

p[t]

]

history = []

for lag in range(1, history_window + 1):

history.extend([

x[t - lag],

p[t - lag]

])

X_present.append(present)

X_history.append(

present + history

)

y.append(x[t + 1])

trajectory_ids.append(tid)

return (

np.asarray(X_present),

np.asarray(X_history),

np.asarray(y),

np.asarray(trajectory_ids)

)

def trajectory_folds(

trajectory_ids,

n_folds

):

unique_ids = np.unique(

trajectory_ids

)

rng = np.random.default_rng(

CFG.random_seed

)

shuffled = unique_ids.copy()

rng.shuffle(shuffled)

folds = np.array_split(

shuffled,

n_folds

)

return folds

def select_alpha(

X,

y,

trajectory_ids,

alpha_grid,

n_folds

):

"""

Seleciona alpha somente usando os dados de treino.

Os folds são definidos por TRAJETÓRIA,

evitando que pontos da mesma trajetória

apareçam simultaneamente em treino e validação.

"""

folds = trajectory_folds(

trajectory_ids,

n_folds

)

scores = {

alpha: []

for alpha in alpha_grid

}

unique_ids = np.unique(

trajectory_ids

)

for alpha in alpha_grid:

for validation_ids in folds:

validation_ids = set(

validation_ids.tolist()

)

train_mask = np.array([

tid not in validation_ids

for tid in trajectory_ids

])

val_mask = ~train_mask

if (

train_mask.sum() == 0

or val_mask.sum() == 0

):

continue

model = Pipeline([

(

"scale",

StandardScaler()

),

(

"ridge",

Ridge(

alpha=alpha,

fit_intercept=True

)

)

])

model.fit(

X[train_mask],

y[train_mask]

)

pred = model.predict(

X[val_mask]

)

mse = np.mean(

(y[val_mask] - pred) ** 2

)

scores[alpha].append(mse)

mean_scores = {

alpha: np.mean(values)

for alpha, values in scores.items()

if len(values) > 0

}

best_alpha = min(

mean_scores,

key=mean_scores.get

)

return best_alpha, mean_scores

def fit_final_model(

X,

y,

alpha

):

model = Pipeline([

(

"scale",

StandardScaler()

),

(

"ridge",

Ridge(

alpha=alpha,

fit_intercept=True

)

)

])

model.fit(X, y)

return model

def trajectory_losses(

model,

X,

y,

trajectory_ids

):

pred = model.predict(X)

df = pd.DataFrame({

"trajectory_id": trajectory_ids,

"y": y,

"pred": pred

})

df["sq_error"] = (

df["y"] - df["pred"]

) ** 2

losses = (

df.groupby("trajectory_id")

["sq_error"]

.mean()

)

return losses.values, pred

def bootstrap_difference(

loss_present,

loss_history,

n_boot,

seed

):

"""

Bootstrap pareado por trajetória.

A unidade de reamostragem é a trajetória,

NÃO cada ponto temporal.

Isso preserva a dependência interna

de cada trajetória.

"""

rng = np.random.default_rng(seed)

diff = (

loss_present

  • loss_history

)

n = len(diff)

boot = np.empty(n_boot)

for i in range(n_boot):

idx = rng.integers(

0,

n,

size=n

)

boot[i] = np.mean(

diff[idx]

)

lower = np.quantile(

boot,

0.025

)

upper = np.quantile(

boot,

0.975

)

return (

np.mean(diff),

lower,

upper,

boot

)

def paired_permutation_test(

loss_present,

loss_history,

n_perm,

seed

):

"""

Teste de permutação pareado.

Sob H0, dentro de cada trajetória,

a identidade de qual modelo obteve

qual erro pode ser permutada.

H1:

MSE_present > MSE_history

"""

rng = np.random.default_rng(seed)

diff = (

loss_present

  • loss_history

)

observed = np.mean(diff)

n = len(diff)

extreme = 0

for _ in range(n_perm):

signs = rng.choice(

[-1, 1],

size=n

)

permuted = np.mean(

diff * signs

)

if permuted >= observed:

extreme += 1

p_value = (

extreme + 1

) / (

n_perm + 1

)

return observed, p_value

def classify_construct_I(

mean_difference,

ci_lower,

ci_upper,

relative_improvement,

p_value,

practical_threshold

):

"""

Critério pré-especificado.

'+':

efeito positivo

  • IC não cruza zero

  • melhoria relativa >= limiar

  • p < 0.05

'-':

evidência contra melhoria relevante:

IC superior <= 0

OU melhoria inferior ao limiar

UE:

dados insuficientes / resultado inconclusivo

"""

if (

ci_lower > 0

and relative_improvement >= practical_threshold

and p_value < CFG.alpha_stat

):

return "+"

if (

ci_upper <= 0

or relative_improvement < 0

):

return "-"

return "UE"

def make_placebo_history(

X_history,

rng

):

"""

Mantém o MESMO número de features históricas,

mas quebra sua relação temporal com o alvo.

Isso testa se uma melhoria simplesmente aparece

porque o modelo possui mais parâmetros/features.

As features do estado presente permanecem intactas.

"""

X_placebo = X_history.copy()

if X_history.shape[1] <= 2:

return X_placebo

historical = X_placebo[:, 2:].copy()

permutation = rng.permutation(

historical.shape[0]

)

historical = historical[

permutation

]

X_placebo[:, 2:] = historical

return X_placebo

def run_single_system(

system_name,

noise_std,

seed

):

print("\n" + "=" * 70)

print(f"SISTEMA: {system_name}")

print(f"Ruído: {noise_std}")

print("=" * 70)

train = generate_dataset(

system_name=system_name,

n_trajectories=CFG.n_train_trajectories,

length=CFG.trajectory_length,

burn_in=CFG.burn_in,

noise_std=noise_std,

seed=seed

)

test = generate_dataset(

system_name=system_name,

n_trajectories=CFG.n_test_trajectories,

length=CFG.trajectory_length,

burn_in=CFG.burn_in,

noise_std=noise_std,

seed=seed + 100000

)

results = []

detailed_predictions = {}

for history_window in CFG.history_windows:

print(

f"\n História = {history_window}"

)

(

Xp_train,

Xh_train,

y_train,

tid_train

) = make_supervised_dataset(

train,

history_window

)

(

Xp_test,

Xh_test,

y_test,

tid_test

) = make_supervised_dataset(

test,

history_window

)

alpha_present, _ = select_alpha(

Xp_train,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_present = fit_final_model(

Xp_train,

y_train,

alpha_present

)

alpha_history, _ = select_alpha(

Xh_train,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_history = fit_final_model(

Xh_train,

y_train,

alpha_history

)

loss_p, pred_p = trajectory_losses(

model_present,

Xp_test,

y_test,

tid_test

)

loss_h, pred_h = trajectory_losses(

model_history,

Xh_test,

y_test,

tid_test

)

(

delta,

ci_lower,

ci_upper,

boot

) = bootstrap_difference(

loss_p,

loss_h,

CFG.bootstrap_reps,

seed + history_window

)

(

observed_perm,

p_value

) = paired_permutation_test(

loss_p,

loss_h,

CFG.permutation_reps,

seed + 1000 + history_window

)

mse_present = np.mean(

loss_p

)

mse_history = np.mean(

loss_h

)

relative_improvement = (

mse_present - mse_history

) / mse_present

ss_total = np.sum(

(

y_test

  • np.mean(y_test)

) ** 2

)

ss_present = np.sum(

(

y_test

  • pred_p

) ** 2

)

ss_history = np.sum(

(

y_test

  • pred_h

) ** 2

)

r2_present = (

1

  • ss_present / ss_total

)

r2_history = (

1

  • ss_history / ss_total

)

classification = classify_construct_I(

delta,

ci_lower,

ci_upper,

relative_improvement,

p_value,

CFG.practical_relative_threshold

)

result = {

"System": system_name,

"Noise": noise_std,

"History_Window": history_window,

"Alpha_Present": alpha_present,

"Alpha_History": alpha_history,

"MSE_Present": mse_present,

"MSE_History": mse_history,

"Delta_MSE": delta,

"Relative_Improvement": relative_improvement,

"R2_Present": r2_present,

"R2_History": r2_history,

"CI95_Lower": ci_lower,

"CI95_Upper": ci_upper,

"Permutation_P": p_value,

"HDD_Construct_I": classification

}

results.append(result)

detailed_predictions[

history_window

] = {

"y": y_test,

"pred_present": pred_p,

"pred_history": pred_h,

"loss_present": loss_p,

"loss_history": loss_h

}

print(

f" MSE presente: {mse_present:.8f}"

)

print(

f" MSE história: {mse_history:.8f}"

)

print(

f" ΔMSE: {delta:.8f}"

)

print(

f" Melhoria relativa: "

f"{100 * relative_improvement:.3f}%"

)

print(

f" IC95% ΔMSE: "

f"[{ci_lower:.8f}, {ci_upper:.8f}]"

)

print(

f" Permutation p: {p_value:.6f}"

)

print(

f" HDD-I: {classification}"

)

return (

pd.DataFrame(results),

detailed_predictions

)

all_results = []

all_predictions = {}

start_time = time.time()

for system_name in SYSTEM_GENERATORS:

result_df, predictions = run_single_system(

system_name=system_name,

noise_std=CFG.noise_std,

seed=CFG.random_seed

)

all_results.append(result_df)

all_predictions[

system_name

] = predictions

results_df = pd.concat(

all_results,

ignore_index=True

)

elapsed = time.time() - start_time

print("\n")

print("=" * 70)

print("EXECUÇÃO PRINCIPAL CONCLUÍDA")

print("=" * 70)

print(

f"Tempo: {elapsed / 60:.2f} minutos"

)

main_window = max(

CFG.history_windows

)

main_results = (

results_df[

results_df["History_Window"]

== main_window

]

.copy()

)

print("\n")

print("=" * 70)

print(

f"RESULTADO PRINCIPAL — HISTÓRIA = {main_window}"

)

print("=" * 70)

display_columns = [

"System",

"MSE_Present",

"MSE_History",

"Delta_MSE",

"Relative_Improvement",

"CI95_Lower",

"CI95_Upper",

"Permutation_P",

"HDD_Construct_I"

]

display(

main_results[

display_columns

].round(6)

)

print("\n")

print("=" * 70)

print("PERFIL DE DEPENDÊNCIA HISTÓRICA")

print("=" * 70)

plt.figure(

figsize=(11, 7)

)

for system_name in SYSTEM_GENERATORS:

subset = results_df[

results_df["System"]

== system_name

]

plt.plot(

subset["History_Window"],

100 * subset["Relative_Improvement"],

marker="o",

linewidth=2,

label=system_name

)

plt.axhline(

100 * CFG.practical_relative_threshold,

color="black",

linestyle="--",

alpha=0.7,

label="Limiar prático"

)

plt.axhline(

0,

color="gray",

linestyle=":",

alpha=0.7

)

plt.xlabel(

"Janela máxima de história τ"

)

plt.ylabel(

"Melhoria relativa (%)"

)

plt.title(

"HDD Construct I — Perfil de dependência histórica"

)

plt.legend()

plt.grid(alpha=0.25)

plt.tight_layout()

plt.savefig(

os.path.join(

CFG.output_dir,

"history_dependence_profile.png"

),

dpi=200

)

plt.show()

def run_capacity_control(

system_name,

noise_std,

history_window,

seed

):

print("\n")

print("=" * 70)

print(

f"CONTROLE DE CAPACIDADE — {system_name}"

)

print("=" * 70)

train = generate_dataset(

system_name,

CFG.n_train_trajectories,

CFG.trajectory_length,

CFG.burn_in,

noise_std,

seed

)

test = generate_dataset(

system_name,

CFG.n_test_trajectories,

CFG.trajectory_length,

CFG.burn_in,

noise_std,

seed + 100000

)

(

Xp_train,

Xh_train,

y_train,

tid_train

) = make_supervised_dataset(

train,

history_window

)

(

Xp_test,

Xh_test,

y_test,

tid_test

) = make_supervised_dataset(

test,

history_window

)

alpha_p, _ = select_alpha(

Xp_train,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_p = fit_final_model(

Xp_train,

y_train,

alpha_p

)

alpha_h, _ = select_alpha(

Xh_train,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_h = fit_final_model(

Xh_train,

y_train,

alpha_h

)

rng = np.random.default_rng(

seed + 9999

)

Xh_train_placebo = make_placebo_history(

Xh_train,

rng

)

Xh_test_placebo = make_placebo_history(

Xh_test,

rng

)

alpha_placebo, _ = select_alpha(

Xh_train_placebo,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_placebo = fit_final_model(

Xh_train_placebo,

y_train,

alpha_placebo

)

loss_p, _ = trajectory_losses(

model_p,

Xp_test,

y_test,

tid_test

)

loss_h, _ = trajectory_losses(

model_h,

Xh_test,

y_test,

tid_test

)

loss_placebo, _ = trajectory_losses(

model_placebo,

Xh_test_placebo,

y_test,

tid_test

)

real_improvement = (

np.mean(loss_p)

  • np.mean(loss_h)

) / np.mean(loss_p)

placebo_improvement = (

np.mean(loss_p)

  • np.mean(loss_placebo)

) / np.mean(loss_p)

return {

"System": system_name,

"Real_Improvement": real_improvement,

"Placebo_Improvement": placebo_improvement,

"Real_MSE": np.mean(loss_h),

"Placebo_MSE": np.mean(loss_placebo)

}

capacity_results = []

for system_name in SYSTEM_GENERATORS:

capacity_results.append(

run_capacity_control(

system_name,

CFG.noise_std,

main_window,

CFG.random_seed

)

)

capacity_df = pd.DataFrame(

capacity_results

)

print("\nCONTROLE DE CAPACIDADE")

display(

capacity_df.round(6)

)

capacity_df.to_csv(

os.path.join(

CFG.output_dir,

"capacity_control.csv"

),

index=False

)

def state_reconstruction_diagnostic(

system_name,

noise_std,

history_window,

seed

):

"""

Diagnóstico de reconstrução de estado.

NÃO é apresentado como prova de "história irredutível".

A ideia é verificar se uma representação comprimida

da história consegue preservar sua vantagem preditiva.

Usamos PCA treinado APENAS nos dados de treino.

Isso testa uma hipótese mais fraca:

a vantagem da história desaparece

quando a história é comprimida em uma representação

de baixa dimensão?

Se sim:

isso sugere que a história pode estar fornecendo

informação sobre um estado latente/reconstruível.

Se não:

a dependência pode ser mais complexa.

Isso NÃO prova irreducibilidade ontológica.

"""

train = generate_dataset(

system_name,

CFG.n_train_trajectories,

CFG.trajectory_length,

CFG.burn_in,

noise_std,

seed

)

test = generate_dataset(

system_name,

CFG.n_test_trajectories,

CFG.trajectory_length,

CFG.burn_in,

noise_std,

seed + 100000

)

(

Xp_train,

Xh_train,

y_train,

tid_train

) = make_supervised_dataset(

train,

history_window

)

(

Xp_test,

Xh_test,

y_test,

tid_test

) = make_supervised_dataset(

test,

history_window

)

H_train = Xh_train[:, 2:]

H_test = Xh_test[:, 2:]

n_components = min(

3,

H_train.shape[1]

)

scaler = StandardScaler()

H_train_scaled = scaler.fit_transform(

H_train

)

H_test_scaled = scaler.transform(

H_test

)

pca = PCA(

n_components=n_components

)

S_train = pca.fit_transform(

H_train_scaled

)

S_test = pca.transform(

H_test_scaled

)

Xs_train = np.column_stack([

Xp_train,

S_train

])

Xs_test = np.column_stack([

Xp_test,

S_test

])

alpha_s, _ = select_alpha(

Xs_train,

y_train,

tid_train,

CFG.alpha_grid,

CFG.inner_folds

)

model_s = fit_final_model(

Xs_train,

y_train,

alpha_s

)

loss_s, _ = trajectory_losses(

model_s,

Xs_test,

y_test,

tid_test

)

mse_s = np.mean(loss_s)

return {

"System": system_name,

"MSE_Reconstructed_State": mse_s,

"Explained_Variance_PCA": np.sum(

pca.explained_variance_ratio_

)

}

reconstruction_results = []

for system_name in SYSTEM_GENERATORS:

reconstruction_results.append(

state_reconstruction_diagnostic(

system_name,

CFG.noise_std,

main_window,

CFG.random_seed

)

)

reconstruction_df = pd.DataFrame(

reconstruction_results

)

print("\n")

print("=" * 70)

print("DIAGNÓSTICO DE RECONSTRUÇÃO DE ESTADO")

print("=" * 70)

display(

reconstruction_df.round(6)

)

reconstruction_df.to_csv(

os.path.join(

CFG.output_dir,

"state_reconstruction.csv"

),

index=False

)

noise_results = []

print("\n")

print("=" * 70)

print("ANÁLISE DE ROBUSTEZ AO RUÍDO")

print("=" * 70)

for noise in CFG.noise_levels:

print(

f"\nRuído = {noise}"

)

for system_name in SYSTEM_GENERATORS:

df, _ = run_single_system(

system_name,

noise,

CFG.random_seed

)

selected = df[

df["History_Window"]

== main_window

].copy()

noise_results.append(

selected

)

noise_df = pd.concat(

noise_results,

ignore_index=True

)

noise_df.to_csv(

os.path.join(

CFG.output_dir,

"noise_robustness.csv"

),

index=False

)

seed_results = []

print("\n")

print("=" * 70)

print("ROBUSTEZ A MÚLTIPLAS SEMENTES")

print("=" * 70)

for seed in CFG.robustness_seeds:

print(

f"\nSeed = {seed}"

)

for system_name in SYSTEM_GENERATORS:

df, _ = run_single_system(

system_name,

CFG.noise_std,

seed

)

selected = df[

df["History_Window"]

== main_window

].copy()

selected["Seed"] = seed

seed_results.append(

selected

)

seed_df = pd.concat(

seed_results,

ignore_index=True

)

seed_df.to_csv(

os.path.join(

CFG.output_dir,

"multi_seed_robustness.csv"

),

index=False

)

seed_summary = []

for system_name in SYSTEM_GENERATORS:

subset = seed_df[

seed_df["System"]

== system_name

]

positive_fraction = np.mean(

subset["HDD_Construct_I"]

== "+"

)

mean_improvement = np.mean(

subset["Relative_Improvement"]

)

std_improvement = np.std(

subset["Relative_Improvement"],

ddof=1

)

seed_summary.append({

"System": system_name,

"Positive_Fraction": positive_fraction,

"Mean_Relative_Improvement":

mean_improvement,

"SD_Relative_Improvement":

std_improvement

})

seed_summary_df = pd.DataFrame(

seed_summary

)

print("\n")

print("=" * 70)

print("RESUMO DE ROBUSTEZ")

print("=" * 70)

display(

seed_summary_df.round(6)

)

print("\n")

print("=" * 70)

print("BENCHMARK DE EQUIVALÊNCIA COMPORTAMENTAL")

print("=" * 70)

print(

"""

Este teste é importante para a interpretação do HDD.

Delay Line e Recurrent podem possuir mecanismos

internos diferentes, mas o Construct I não deve

tentar inferir arquitetura a partir de simples

melhoria preditiva.

Portanto:

Construct I positivo

NÃO implica

Construct III positivo.

"""

)

ground_truth = pd.DataFrame({

"System": [

"Markov",

"Hidden State",

"Delay Line",

"Recurrent"

],

"Expected_Construct_I": [

"-",

"+",

"+",

"+"

],

"Reason": [

"Estado observado é suficiente",

"História revela estado latente",

"Dinâmica depende explicitamente de atraso",

"História ajuda a inferir estado recorrente oculto"

]

})

final_main = main_results.merge(

ground_truth,

on="System",

how="left"

)

final_main[

"Correct_vs_Ground_Truth"

] = (

final_main[

"HDD_Construct_I"

]

final_main[

"Expected_Construct_I"

]

)

print("\n")

print("=" * 70)

print("COMPARAÇÃO COM GROUND TRUTH")

print("=" * 70)

display(

final_main[

[

"System",

"HDD_Construct_I",

"Expected_Construct_I",

"Correct_vs_Ground_Truth"

]

]

)

accuracy = np.mean(

final_main[

"Correct_vs_Ground_Truth"

]

)

print(

f"\nAcurácia contra ground truth: "

f"{100 * accuracy:.1f}%"

)

plt.figure(

figsize=(10, 6)

)

colors = {

"Markov": "black",

"Hidden State": "blue",

"Delay Line": "orange",

"Recurrent": "red"

}

for system_name in SYSTEM_GENERATORS:

subset = results_df[

results_df["System"]

== system_name

]

plt.errorbar(

subset["History_Window"],

100 * subset["Relative_Improvement"],

fmt="o-",

capsize=4,

linewidth=2,

label=system_name,

color=colors[system_name]

)

plt.axhline(

0,

color="gray",

linestyle=":"

)

plt.axhline(

100 * CFG.practical_relative_threshold,

color="black",

linestyle="--",

label="Limiar prático"

)

plt.xlabel(

"História máxima τ"

)

plt.ylabel(

"Melhoria relativa (%)"

)

plt.title(

"Efeito da história sobre a previsão"

)

plt.grid(

alpha=0.25

)

plt.legend()

plt.tight_layout()

plt.savefig(

os.path.join(

CFG.output_dir,

"effect_size_profile.png"

),

dpi=200

)

plt.show()

final_results = results_df.merge(

ground_truth,

on="System",

how="left"

)

final_results[

"Correct_vs_Ground_Truth"

] = np.where(

final_results["HDD_Construct_I"]

== final_results["Expected_Construct_I"],

True,

False

)

final_results.to_csv(

os.path.join(

CFG.output_dir,

"HDD_Construct_I_full_results.csv"

),

index=False

)

main_results.to_csv(

os.path.join(

CFG.output_dir,

"HDD_Construct_I_main_results.csv"

),

index=False

)

manifest = {

"benchmark_version": "HDD Benchmark v1.0",

"construct": "I",

"definition":

"History improves out-of-sample prediction "

"beyond measured present state and current input",

"configuration": asdict(CFG),

"ground_truth": ground_truth.to_dict(

orient="records"

),

"environment": ENVIRONMENT,

"files": [

"HDD_Construct_I_full_results.csv",

"HDD_Construct_I_main_results.csv",

"capacity_control.csv",

"state_reconstruction.csv",

"noise_robustness.csv",

"multi_seed_robustness.csv",

"history_dependence_profile.png",

"effect_size_profile.png",

"environment.json"

],

"important_interpretation":

"Positive Construct I does not establish memory "

"mechanism, recurrence, self-reference, self-modeling "

"or consciousness."

}

with open(

os.path.join(

CFG.output_dir,

"manifest.json"

),

"w"

) as f:

json.dump(

manifest,

f,

indent=2,

default=str

)

report_lines = []

report_lines.append(

"HDD BENCHMARK v1.0"

)

report_lines.append(

"Construct I — History-Dependent Predictive Structure"

)

report_lines.append(

""

)

report_lines.append(

"RESULTADO PRINCIPAL"

)

for _, row in main_results.iterrows():

report_lines.append(

f"{row['System']}: "

f"HDD-I={row['HDD_Construct_I']}, "

f"improvement="

f"{100 * row['Relative_Improvement']:.3f}%, "

f"CI95=["

f"{row['CI95_Lower']:.8f}, "

f"{row['CI95_Upper']:.8f}], "

f"p="

f"{row['Permutation_P']:.6f}"

)

report_lines.append(

""

)

report_lines.append(

"GROUND TRUTH"

)

for _, row in ground_truth.iterrows():

report_lines.append(

f"{row['System']}: "

f"expected={row['Expected_Construct_I']}"

)

report_lines.append(

""

)

report_lines.append(

f"Preliminary benchmark accuracy: "

f"{100 * accuracy:.1f}%"

)

report_lines.append(

""

)

report_lines.append(

"INTERPRETATION:"

)

report_lines.append(

"A positive Construct I classification means only "

"that historical information improved out-of-sample "

"prediction relative to the specified observed present "

"state and current input."

)

report_lines.append(

"It does not establish a memory mechanism, recurrence, "

"self-reference, self-modeling, agency or consciousness."

)

report_lines.append(

""

)

report_lines.append(

"LIMITATION:"

)

report_lines.append(

"This benchmark validates the computational "

"discrimination of Construct I under the specified "

"observation and model class. It does not validate "

"the complete HDD framework."

)

report_text = "\n".join(

report_lines

)

print("\n")

print("=" * 70)

print("RELATÓRIO")

print("=" * 70)

print(report_text)

with open(

os.path.join(

CFG.output_dir,

"REPORT.txt"

),

"w"

) as f:

f.write(report_text)

import shutil

zip_path = shutil.make_archive(

CFG.output_dir,

"zip",

CFG.output_dir

)

print("\n")

print("=" * 70)

print("BENCHMARK FINALIZADO")

print("=" * 70)

print(

f"\nArquivos salvos em:\n"

f"{os.path.abspath(CFG.output_dir)}"

)

print(

f"\nPacote ZIP:\n"

f"{zip_path}"

)

print("\n")

print("=" * 70)

print("INTERPRETAÇÃO CIENTÍFICA")

print("=" * 70)

print(

"""

O benchmark foi desenhado para responder apenas:

A história melhora a previsão do futuro

além do estado observado presente?

Ele NÃO responde:

A história é memória?

Existe recorrência?

Existe self-reference?

Existe self-model?

Existe consciência?

Essas perguntas exigem outros experimentos.

O resultado correto deve ser interpretado

relativamente ao observation model, intervention

protocol e model class utilizados.

O sistema Hidden State é especialmente importante:

um resultado positivo nele NÃO significa que o sistema

possui uma memória explícita. A história pode simplesmente

revelar informação sobre uma variável latente.

Portanto, o benchmark testa precisamente uma das

teses metodológicas centrais do HDD:

"história preditiva não é automaticamente

evidência de um mecanismo específico."

"""

)

print("\n")

print("=" * 70)

print("FIM")

print("=" * 70)

── more in #machine-learning 4 stories · sorted by recency
── more on @taotuner 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/history-dependent-dy…] indexed:0 read:67min 2026-08-13 ·