REI
Unlocking Plasticity Rules
4:54 PM | Aug 20, 2026

// Unlocking Plasticity Rules

Temporal Context Projection and Counterfactual Utility Plasticity are live in the Adapt-1 Preview API today. What each rule changes, and what the current evidence establishes.

Disclaimer: CUP and TCP are learning rules inside Adapt-1, with their own source code and mathematical formulation to be published. If you are unfamiliar with the architecture, Adapt-1 Preview Part 1, Adapt-1 Preview Part 2, and the Behavioral Report are important background reads.

Useful parts can fail together

A planner can exploit a small error in learned physics, search toward a trajectory that looks optimal only inside the model, and hand the resulting impossible state to a value estimator that confidently rewards it.

Under different conditions, a short model rollout combined with a value estimate can outperform either mechanism alone. Separate component tests leave their conditional joint utility unmeasured.

The same robot mechanisms can be complementary or interfere.
Figure 1. Composition outcome. The world model, planner, and value estimator each produce useful predictions. A short rollout and value estimate produce positive joint utility in one regime. Planner search amplifies a model error in the second regime, and the estimator rewards an unreachable state.

Each component can remain useful while the composition fails. Most systems fix the relationships among components before outcomes can evaluate them. Those relationships must instead change in response to measured outcomes.

Structural choices ordinary updates leave fixed

An online learner can update parameters while its computational structure remains fixed. It may revise a policy, fit a model, update stored state, or change predictor confidence. The architecture generally predefines the information admitted into current state and the permitted interactions among internal predictors.

Optimization begins after those declarations. A model can optimize over its supplied state vector, and an ensemble can learn weights over supplied predictions. These operations leave upstream state aliasing and interaction-specific predictor utility outside their update targets.

Under partial observability, these declarations determine the evidence represented at decision time and the units that can receive credit from a later outcome. Discarded information remains unavailable during the decision. Per-source posterior scores cannot encode an error that occurs only when sources are combined.

An online structure learner operating under delayed outcomes must construct a causally valid state from bounded experience and estimate interaction-specific predictive utility. Persistent state must remain bounded while both structures and source models continue to change.

An Adapt-1 Domain can enable contextual policy, episodic and associative memory, learned models, transition state, relational state, sequential credit, and an adaptive posterior. Several persistent mechanisms already update from eligible events and attributed outcomes. The two learning rules extend plasticity to the representation used for a decision and the topology among registered predictors.

What ordinary online learning often leaves fixed.
Figure 2. Declared structure. Online updates change memory statistics, model parameters, policy state, and source reliability. The declared state representation and composition rule stay fixed.

Observation-equivalent process states

Under partial observability, the current observation may fail to identify the situation that produced it. Two moments can share the same observation and require different actions because their transition sequences produced different process states. State aliasing has merged situations the task requires the learner to distinguish.

Additional updates preserve that contradiction when the representation stays unchanged. The learner needs a bounded causal procedure that admits selected committed history before prediction. Future observations remain unavailable, and the decision representation stays linked to the later outcome that evaluates it. Outcomes can reveal which eligible temporal dependency is useful.

The same failed test can follow either a dependency update or a local interface edit in a software-agent trace. The required intervention depends on the transition sequence even when the terminal frame is identical. Optimization over the final error message leaves the missing process state unavailable.

Observation-equivalent states in a software-agent trace.
Figure 3. Process-state ambiguity. A software agent receives the same terminal frame after two causal trajectories. The trajectory determines the process state associated with that observation.

Temporal Context Projection

Temporal Context Projection (TCP) addresses this representation boundary. It acts before a decision by allowing bounded, causally ordered observations from the same episode to participate in the current representation.

TCP admits eligible history before the decision while leaving the rewarding lag undeclared. Adapt-1's existing substrate discovers which projected context predicts the measured outcome through ordinary online updates.

A declared temporal horizon and selected fields bound the eligible projection. The decision representation preserves causal order and stays linked to the feedback that evaluates it.

Temporal Context Projection.
Figure 4. State construction. The maximum lag bounds causal history available before a decision. Ordinary learning discovers which eligible prior field predicts later feedback.

TCP on RepeatPrevious

POPGym RepeatPrevious turns state aliasing into a controlled task. On RepeatPreviousMedium, the correct action at time t is the observation seen at t−32. Present-only input is therefore insufficient by construction. Identical current inputs can have incompatible targets because their histories differ.

The Domain exposed a generic maximum temporal lag of 64 while leaving the rewarding lag undeclared. TCP placed bounded prior observations in the current representation, where ordinary online learning discovered the useful dependency from outcomes.

Five TCP-only runs each received ten episodes, totaling 1,030 interactions per run. The final-episode mean return was 0.9444, and the last-five-episode mean was 0.7578. Seed 0 then entered frozen evaluation with feedback, writes, and adaptation disabled. It returned 1.0000 across sixteen unseen episodes, and its state fingerprint stayed unchanged.

RepeatPreviousEasy provides the causal ablation. The present-only condition ended at −0.5109 mean return and 0.2445 action accuracy. TCP reached 1.0000 return and accuracy in every one of five runs, then remained perfect across frozen evaluation. TCP plus Counterfactual Utility Plasticity (CUP) reached the same ceiling. Because TCP-only already reached ceiling and produced lower probability loss, the return gain is fully attributable to TCP.

POPGym RepeatPreviousEasy frozen results.
Figure 5. Causal state result. TCP changed the decision representation and produced the return gain. The TCP-only condition already reached ceiling, so CUP receives no causal attribution for the 32-step solution.

Published results provide difficulty and sample-budget context. In the original POPGym study, twelve of thirteen reported models had negative MMER on RepeatPreviousMedium after fifteen million timesteps, with LMU as the positive exception. A later TMLR table reports vTransformer and GRU at 1.000 on RepeatPreviousEasy after one million environment steps. Adapt-1 TCP reached 1.000 after 510 online interactions per run. The protocols differ in optimization, source lineage, metrics, and uncertainty reporting. Direct rank claims would require normalized evaluation.

SystemRepeatPreviousEasyTraining and evaluation protocol
Adapt-1 TCP1.000 ± 0.000 SD510 online interactions per run; 16 frozen episodes × 5 runs
vTransformer1.000 ± 0.000 SE1,000,000 environment steps; 16 test episodes × 5 seeds
GRU1.000 ± 0.000 SE1,000,000 environment steps; 16 test episodes × 5 seeds
Mamba0.993 ± 0.001 SE1,000,000 environment steps; 16 test episodes × 5 seeds
Published memoryless−0.434 ± 0.013 SE1,000,000 environment steps; 16 test episodes × 5 seeds

Within the matched experiment, changing the representation boundary made the aliased mapping learnable. TCP supplied temporal identity, and Adapt-1's existing substrate learned the 32-step dependency.

Interaction-specific predictive utility

The robot failure illustrates a broader composition problem. A relationship among predictors can add value beyond its parts. It can also duplicate evidence or introduce systematic conflict, and its effect can change with the environment.

Accurate predictors can still repeat the same upstream error when combined. Breiman's analysis of random forests formalized this strength–correlation tradeoff.

Per-predictor reliability scores cannot represent those conditional effects; the missing utility belongs to the relationship.

Counterfactual Utility Plasticity

Counterfactual Utility Plasticity (CUP) makes the relationships among an agent's registered prediction sources learnable. Ordinary posterior learning estimates how much to trust each source on its own. CUP goes further: after the sources commit their predictions and decision-linked outcome feedback arrives, it asks whether a combination predicted the outcome better or worse than its members and every smaller combination can already explain. It accumulates that relationship-specific evidence over time.

When the evidence is strong enough, a helpful relationship is consolidated into bounded persistent state and can alter later posterior fusion. A harmful relationship is marked as inhibited, while older evidence decays so relationships can be revised when conditions change. CUP therefore learns not only which sources are reliable, but which sources are complementary, redundant, or harmful together.

RepeatPreviousMedium exposes CUP state among multiple registered predictors. Wisconsin isolates predictor composition through static prequential classification against fixed fusion and simple selection.

Persistent marginal and joint utility

RepeatPreviousMedium records CUP's structural state while TCP and the adaptive substrate remain causally responsible for the 32-step solution. The trace isolates the persistent predictor relationships that emerged during the run.

RepeatPreviousMedium exposed two relevant sources: contextual_memory and learned_model. The exported CUP state retained marginal and joint relations for those sources.

In seed 0, learned_model ended consolidated with positive utility, while contextual_memory × learned_model ended inhibited. In seed 3, contextual_memory was inhibited alone, while the pair was consolidated. A source's marginal utility and conditional utility can therefore have opposite signs.

Two learned divisions of predictive labor on RepeatPreviousMedium.
Figure 6. Inspectable topology. CUP learned separate persistent utility for marginal and joint relations. One run inhibited the pair, and another consolidated it.

Each exported relation records its participants, interaction order, support, uncertainty, status, and persistent weight. Its weight can modify later composition. The exported state shows which internal predictors and predictor relationships earned supportive or inhibitory influence.

Across five paired ten-episode Medium runs, TCP plus CUP raised mean learning-curve return from 0.2839 to 0.3167, an 11.5% relative increase, and won four of five seeds. The last-five mean rose from 0.7578 to 0.8067. The paired 95% interval crossed zero, and both variants reached 1.0000 in the frozen seed-0 test.

CUP's additional acquisition effect on RepeatPreviousMedium.
Figure 7. Acquisition interval. Across five paired runs, CUP raised the average acquisition curve. The paired 95% interval crossed zero, leaving the acquisition effect unresolved.

Five runs leave the acquisition effect unresolved. They do establish the structural result: CUP learned persistent marginal and joint influence among separately attributable predictors while the primary learner continued to operate.

Adaptive topology against simpler controls

Wisconsin Diagnostic Breast Cancer isolates online predictor composition without the long-memory requirement. The first twelve declared features generated singleton and pair coalition predictors. Across thirty randomized prequential orders, every probability was issued before its label arrived. The first 100 observations were excluded as warm-up, and feature selection remained independent of labels.

Across the matched ladder, only the composition rule changed: the exact disabled fallback, uniform fusion of the same coalition bank, selection of the single coalition with the lowest past prequential loss, or CUP's adaptive topology.

CUP improved probability quality over the disabled path and fixed uniform fusion. Against uniform fusion, it reduced mean log loss by 0.0201 in all thirty orders. Fixed uniform retained a 0.0131 balanced-accuracy advantage.

Wisconsin matched CUP control ladder.
Figure 8. Matched composition controls. CUP improved probability quality over uniform fusion in every randomized order. Past-loss selection of one coalition achieved lower log loss, and fixed uniform produced the highest balanced accuracy.

Best-coalition selection produced 0.2493 mean log loss, compared with CUP's 0.3087, and had lower loss in all thirty orders. Its balanced-accuracy advantage over CUP was small and statistically unresolved. Selection fit this stable, highly redundant regime better on log loss.

Selection does not erase CUP's structural result. CUP learned source relationships online and improved log loss over fixed fusion in every order. The control shows that adaptive topology must be evaluated against both fusion and selection.

Controlled shifts provide a stronger test of CUP's distinctive behavior. A useful relation can become redundant or harmful after a regime change, and a weak source can become conditionally necessary. Such shifts exercise consolidation, inhibition, decay, and revision through behavior.

Agent problems addressed by TCP and CUP

TCP addresses process-state ambiguity in agents. Long-running agents encounter identical interface observations produced by different process states. A failed test after a dependency update can require dependency rollback, while the same test after a local schema edit can require interface repair. A latency alert following a deployment and one following a traffic surge also require different interventions.

A transcript preserves history as stored data. TCP makes bounded committed history eligible before the decision, allowing the existing substrate to learn which temporal distinctions identify the next action. A coding agent can learn which earlier edit defines the current failure, and an embodied agent can use recent contact to disambiguate the current sensor frame. Trajectory information erased from the present observation thus enters the decision state.

TCP turns trajectory into actionable process state.
Figure 9. Learned process state. TCP makes bounded committed history available to downstream learning before a decision. The learned representation can distinguish terminal observations produced by different trajectories.

CUP operates on interaction-specific predictive utility among registered sources. Those sources commit predictions before the outcome. Measured feedback evaluates their combinations and writes the resulting relations into persistent topology.

Shared upstream signals can cause two predictors to repeat the same error. A tool-derived source may repair a learned model's blind spot in a specific regime. Marginal source utility can differ from its conditional utility.

CUP gives each relationship persistent state. Agents can use this state to discount correlated agreement, consolidate combinations that add predictive value, inhibit damaging combinations, and revise relations when the environment changes.

Each relation retains its participants, order, status, support, and uncertainty. Agent debugging can inspect the learned interaction, the direction of its utility, the evidence supporting its status, and its response to a regime change.

CUP writes interaction-specific utility into persistent topology.
Figure 10. Persistent interaction utility. CUP writes interaction-specific utility among registered sources into persistent topology. Outcome updates assign supportive or inhibitory status and revise stale relations after utility changes.

Software development and incident response contain hidden process-state ambiguity and interaction-specific source utility. The same problems appear in scientific control, long-horizon browsing, and robotics. Outcome feedback can update both structures while the base learners continue to operate.

During continued deployment, outcome updates can admit a previously missing temporal dependency into current state or change the persistent influence of a predictor relationship. The Domain bounds both updates; exported state preserves the evidence needed for inspection.

Ablation replication: https://github.com/0xReisearch/adapt1-behavioral-study

References