REI
The Dynamics of, Sequential, Transitional and Contextual Online Learning
6:30 PM | Aug 12, 2026

// The Dynamics of, Sequential, Transitional and Contextual Online Learning

Adapt-1 is available in open beta through the Adapt-1 app. Setup instructions and Domain guidance are in the documentation.

This article shows five learning procedures. Each section states the input, the persistent change, the later output, and the limit of the evidence. Retained run files supply the traces.

Earlier articles describe the architecture in more detail: Introducing Adapt-1 Preview and Adaptive State from Partially Observed Streams.

Adapt-1 was built around a specific objective: useful behavior as quickly as possible from the smallest amount of experience available, update from live outcomes as they arrive, and become queryable throughout that process. The emphasis is on sample-efficient adaptation, low mistake rates during learning, and retaining only the information that remains useful for the task.

Evaluations covered: Symbolic Alchemy, ALFWorld, CausaLab, BOP-Ask and Transition mini sim.

Five stages of behavioral evidence: prior state, eligible evidence, persistent revision, later output, and external consequence.
Eligible evidence updates retained state before a later output receives an external consequence.

Case 01 / Rapid online adaptation

Rapid online adaptation in Symbolic Alchemy

For this release, we use Adapt-1 as a fast online learner. It must use limited interaction to improve later queries and reduce avoidable mistakes.

25official chemistry episodes
5,000live environment decisions
251.6mean return
2,143delayed-credit updates

Symbolic Alchemy tests this role with a new hidden chemistry in every episode. Each avoidable action uses part of a fixed decision budget.

Later sections separate these behaviors. We start with Alchemy because one run brings several of them together.

Hidden chemistry defines potion effects and transition constraints. It also defines the stone value function. Each episode contains ten trials under one chemistry.

A trial provides new stones and potions and allows 20 decisions. One chemistry therefore remains active across 200 decisions.

Adapt-1 started with empty learned state and no Alchemy task pretraining. Reported results cover episodes 0 through 24 from the retained archive.

Reported window Value
Episodes 0 through 24
Live environment decisions 5,000
Total return 6,290
Mean return 251.6
Initial learned state Empty
Alchemy task pretraining None

A twenty-sixth retained episode remains outside the reported mean, trace aggregates, and figures.

Adapt-1 kept learning throughout each chemistry. New transitions and rewards updated learned state as the run continued.

Stable structure and hidden chemistry

We supply stable task grammar and ontology through the Domain. It declares 39 public observation fields and 40 native slot actions.

Structural declarations include six signed-axis effects, a bijective assignment, reversible effects, connected topology constraints, and eight objective hypotheses.

Current chemistry stays hidden. Adapt-1 must infer the active potion mapping, transition graph, objective, and useful action sequences from live transitions.

For each live query, Core uses the current public observation and retained evidence to select a native action. The next state and source reward then update the learner.

Each episode begins a fresh session for local chemistry evidence. Persistent learner state continues through the selected 25-episode sequence.

Symbolic Alchemy task structure and online learning target.
Stable structure reduces task ambiguity. Live transitions still determine the active chemistry.

Adaptation within a chemistry

In episode 0, we observe a rapid move from unresolved effects to supported predictions.

Alchemy selected fieldsEpisode 0 / trial 1 / early chemistry learning
Symbolic Alchemy / trial 1steps 0 through 6
EPISODE 0 | TRIAL 1

STEP 0   potion slot 11 → stone slot 2
         explore | support 0 | state changed

STEP 1   potion slot 7 → stone slot 0
         explore | support 0 | state changed

STEP 2   potion slot 3 → stone slot 1
         explore | support 0 | state changed

STEP 3   potion slot 0 → stone slot 0
         explore | support 0 | state changed

STEP 4   potion slot 6 → stone slot 0
         explore | support 1 | state changed

STEP 5   potion slot 1 → stone slot 0
         predicted | exact_state | support 1

STEP 6   potion slot 8 → stone slot 0
         predicted | transferred_transform | support 1
  1. 1
    Unresolved effects. Early selections record exploration and state changes.
  2. 2
    Exact-state support. Step 5 uses evidence from the current state.
  3. 3
    Transferred prediction. Step 6 applies an observed transform to another state.
Selected retained fields. Action, support, prediction basis, and state-change labels come from the retained Alchemy excerpt.

We see five unresolved potion choices at the start. Step 5 then uses exact-state evidence.

At step 6, Core transfers an observed transform to another applicable state. Trial 2 then presents new stones and potions under the same chemistry.

Alchemy selected fieldsEpisode 0 / trial 2 / inference and delayed credit
Symbolic Alchemy / trial 2steps 20 through 26
EPISODE 0 | TRIAL 2

STEP 20  potion slot 7 → stone slot 1
         predicted | transferred_transform | support 1

STEP 21  potion slot 3 → stone slot 1
         predicted | constraint_elimination | support 3

STEP 22  potion slot 5 → stone slot 1
         predicted | transferred_transform | support 1

STEP 23  stone slot 1 → cauldron
         reward 1 | delayed-credit updates 23

STEP 24  potion slot 1 → stone slot 2
         predicted | transferred_transform | support 1

STEP 25  potion slot 10 → stone slot 2
         predicted | transferred_transform | support 2

STEP 26  stone slot 2 → cauldron
         reward 15 | delayed-credit updates 2
  1. 1
    Constraint elimination. Three supporting observations narrow the potion effect.
  2. 2
    First delayed outcome. Reward 1 updates 23 earlier decisions.
  3. 3
    Later delayed outcome. Reward 15 produces two additional credit updates.
Selected retained fields. The excerpt keeps the action order, prediction basis, rewards, and delayed-credit counts.

At step 21, Core uses constraint elimination after three supporting observations. Later cauldron rewards update earlier choices through sequential credit.

Reward 1 at step 23 produces 23 delayed-credit updates. Reward 15 at step 26 produces two more.

These updates let later outcomes revise decisions that created them.

Across 25 episodes, we see the same pattern. Trial 1 contains 148 explore rows among 500 decisions, or 29.6 percent.

Trial 5 contains five such rows, or 1.0 percent. Trials 6 through 10 contain six rows in 2,500 decisions.

Symbolic Alchemy unresolved exploration by trial.
Unresolved selections fall early within each chemistry. Delayed-credit updates remain active throughout the episode.

Fast reduction matters because each trial allows only 20 decisions. Later queries can use supported predictions after uncertainty falls.

Across the selected window, 2,143 delayed-credit updates connect later outcomes to earlier choices.

Adapt-1 uses its native substrate and learner stack. Candidate learners can be neural or non-neural, and each starts this task with empty learned state. As evidence accumulates, Core trains the candidates, gates their contribution, keeps supported learners, and suppresses those that lack support. The same live substrate continues to serve queries at unchanged latency during these updates. Introducing Adapt-1 Preview describes the full architecture.

Early queries resolve potion effects. Later queries reuse learned transitions and inferred constraints while delayed outcomes continue to revise policy evidence.

Pinon et al.

Pinon et al. provide the strongest published comparison for task-data and inference efficiency.

Their method learns a Transformer dynamics model from one million Symbolic Alchemy episodes. Each training episode contains 200 environment steps.

Task-specific training therefore covers about 200 million environment steps before evaluation. Learned dynamics then support tree-search planning.

Their strongest configuration uses 10,000 tree expansions for every real environment action. It reaches 251.5 ± 4.5.

Measure Adapt-1 Pinon et al.
Mean return 251.6 251.5 ± 4.5
Alchemy steps before evaluation 0 About 200,000,000
Reported live steps 5,000 Reported separately from training
Training episodes None 1,000,000
Learned state at start Empty Trained Transformer dynamics model
Planner Bounded Domain transition planner 10,000 tree expansions per action

Adapt-1 reaches the same raw score region after 5,000 live decisions from empty learned state. Alchemy interaction counts differ by about 40,000 times.

This ratio covers task-specific environment interaction. System development and Domain engineering remain outside it.

Pinon reports tree expansions. Adapt records a bounded planner configuration, so inference budgets stay in their published units.

We can compare task data directly and report inference budgets separately. Pinon shows how much Alchemy training previously supported a score near 252.

AlKhamissi et al.

AlKhamissi et al. provide the strongest comparison for representation and behavioral scaffolding.

Their best system reaches 272.85 ± 1.78. Its modified observation counts remaining potions by hue.

A modified action selects a potion hue and a stone. A wrapper then chooses an available potion instance with that hue.

Structured episodic memory stores the prior stone state, potion color, reward, and next stone state. Trials use 15 decisions.

Additional penalties cover null transitions, invalid choices, used stones, and repeated colors on one stone. Their ablations show large changes.

AlKhamissi et al. configuration Mean return
Modified representation and penalties 272.85 ± 1.78
Modified representation, no penalties 236.06 ± 2.18
Canonical input, output, and memory 158.91 ± 1.60

Adapt-1 keeps the public 39-field slot observation, 40 native slot actions, official 20-decision trials, and source reward.

Core selects a specific potion slot. Alchemy learning begins at episode 0.

Domain design gives Adapt-1 more explicit declarative structure. Its action interface provides less behavioral abstraction than AlKhamissi's modified setup.

From their ablations, we observe that representation and reward design can move performance by more than 100 points.

This comparison clarifies how explicit ontology, native action selection, and online learning shape the Adapt result.

Result scope

One retained sequence over 25 official chemistry episodes supports this result. A repeated-seed interval is unavailable.

Published protocols differ in task interface, training, planner, trial length, and reward signal. A matched leaderboard requires a shared protocol.

This run starts from a supplied ontology. Adapt-1 discovers the active chemistry online and uses it during the episode.

Within that scope, Adapt-1 achieves a 251.6 mean after 5,000 live decisions from empty learned state.

We observe quick uncertainty reduction inside each chemistry and continued delayed-credit updates throughout use.

Repeated queries update the learner. Later choices use the resulting evidence before the interaction budget is exhausted.

Case 02 / Online extraction and frozen execution

Online extraction followed by frozen execution

ALFWorld gives an agent household goals, text observations, and legal actions. Some goals require two objects. Other goals require cleaning, heating, or cooling before placement.

3,520official training games
80,635recorded decisions
93.46%valid_seen
90.87%valid_unseen

Adapt-1 uses a clear teaching boundary. First, the harness supplies expert reference actions and observations. Adapt-1 extracts plans and policy parameters from this stream.

The acquisition stream contains 3,520 official training games and 80,635 decisions. Adapt-1 converts 2,797 successful episodes into reusable plans. Adapt-1 then trains the candidate policy for 45 epochs.

Evaluation starts after this work. Five seeds use the frozen learned state. Each seed runs 140 valid_seen games and 134 valid_unseen games.

During evaluation, the evaluator receives the current observation and the admissible commands. Expert actions and all learner updates stay disabled during evaluation.

Evaluation measures the quality of the earlier extraction and tests retained task structure with the current room state.

Three ALFWorld task views show two pillows and their sofa destination.
These frames show the Pick Two task family. They give task context. They are not the exact text evaluation episode.

One trace shows partial completion in a goal that requires two pillows in one sofa.

Retained evaluation excerptEpisode 2 / selected actions, observations, and terminal fields
ALFWorld / held-out Pick Two successseed-0 · valid_unseen
ALFWORLD / HELD-OUT PICK TWO SUCCESS
seed-0 | valid_unseen | episode 2
goal: put two pillow in sofa

STEP 0   go to sofa 1
STEP 1   go to armchair 1
STEP 2   take pillow 2 from armchair 1
STEP 3   go to sofa 1
STEP 4   move pillow 2 to sofa 1
STEP 5   go to sidetable 1
STEP 6   look
STEP 7   go to shelf 9
STEP 8   go to sofa 1
STEP 9   go to shelf 5
STEP 10  go to shelf 6
STEP 11  go to shelf 8
STEP 12  go to armchair 1
STEP 13  take pillow 1 from armchair 1
STEP 14  go to sofa 1
STEP 15  move pillow 1 to sofa 1

STEP 9 OBSERVATION
You arrive at sofa 1. On the sofa 1, you see a creditcard 1,
a pillow 2, and a remotecontrol 2.

STEP 13 OBSERVATION
You arrive at armchair 1. On the armchair 1, you see a pillow 1.

STEP 15 OUTCOME
reward 1.0 | terminal true | won true
  1. 1
    First placement remains visible. The later sofa observation records pillow 2 at the destination.
  2. 2
    The second object is selected. The current observation identifies pillow 1 at the armchair.
  3. 3
    The second placement completes the goal. The terminal fields record reward 1 and a win.
Exact selected actions and cited observations. Other observations, plan identifiers, and internal obligation fields are absent from the excerpt.

Placing pillow 2 changes the room state. Here, we can observe how the evaluator uses that change. A later observation confirms that pillow 2 is already on the sofa. Next, the evaluator selects pillow 1 and completes the task.

Evidence from this trace supports current-state reuse during frozen execution. No retrieved plan identifier or internal completed-object list appears in the trace.

A second trace shows a change in object identity during one task.

Retained evaluation excerptEpisode 86 / object identity changes during execution
ALFWorld / object switchseed-0 · valid_unseen
ALFWORLD / OBJECT SWITCH AND DESTINATION REUSE
seed-0 | valid_unseen | episode 86
goal: heat some apple and put it in fridge

STEP 4   take apple 3 from sinkbasin 1
STEP 5   clean apple 3 with sinkbasin 1
STEP 6   go to microwave 1
STEP 7   go to fridge 1
STEP 8   open fridge 1
STEP 9   move apple 3 to fridge 1
STEP 10  take apple 1 from fridge 1
STEP 11  go to microwave 1
STEP 12  heat apple 1 with microwave 1
STEP 13  go to fridge 1
STEP 14  move apple 1 to fridge 1

STEP 9 OBSERVATION
You open the fridge 1. The fridge 1 is open. In it, you see a apple 1,
[remaining listed objects omitted]

STEP 10 OBSERVATION
You move the apple 3 to the fridge 1.

STEP 14 OUTCOME
reward 1.0 | terminal true | won true
  1. 1
    Apple 3 reaches the fridge. The destination is used before the task finishes.
  2. 2
    Execution changes object identity. The evaluator selects apple 1 from the updated fridge state.
  3. 3
    The heated apple completes the task. The terminal fields record reward 1 and a win.
Exact selected fields. Steps 0 through 3 and the marked remainder of the fridge object list are omitted.

Apple 3 reaches the destination. Task execution continues. Next, the evaluator selects apple 1 from the same fridge. After heating apple 1, the evaluator returns it to the fridge.

All five retained seeds contain this exact selected-action sequence.

A failure trace shows a remaining limit in a task that requires a cooled potato in a microwave.

Retained failure excerptEpisode 15 / ranking fields and terminal failure
ALFWorld / cooling failureseed-0 · valid_unseen
ALFWORLD / COOLING FAILURE BOUNDARY
seed-0 | valid_unseen | episode 15
goal: cool some potato and put it in microwave

STEP 1   take potato 2 from sinkbasin 1
STEP 2   clean potato 2 with sinkbasin 1
STEP 3   go to fridge 1
STEP 4   go to microwave 1
         observation: The fridge 1 is closed.
STEP 5   open microwave 1
STEP 6   move potato 2 to microwave 1
STEP 7   go to garbagecan 1
         reward 0.0 | terminal false | won false

STEP 4 RANKING
1  go to microwave 1
   final 0.7172916666666667 | plan 0.8090087079216964 | milestone 0.95
2  open fridge 1
   final 0.565210515069136 | plan 0.5022426592193358 | milestone 0.7270876968049386
3  go to countertop 1
   final 0.5526616915422886 | plan 0.5718588831224252 | milestone 0.6686567164179105
4  go to countertop 2
   final 0.5143283582089553 | plan 0.5718588831224252 | milestone 0.6686567164179105
5  go to garbagecan 1
   final 0.5020398009950249 | plan 0.569521487673977 | milestone 0.6507462686567165

STEP 49  go to stoveburner 1
         reward 0.0 | terminal true | won false
  1. 1
    The cooling path remains available. The observation says the fridge is closed before the ranking is produced.
  2. 2
    The ranking favors the destination. Movement to the microwave ranks above opening the fridge.
  3. 3
    The milestone order fails. The episode ends after 50 actions with reward 0.
Exact selected actions, all five retained ranking rows, and the terminal row. Steps 8 through 48 are omitted.

Here, the evaluator ranks movement to the microwave above opening the fridge. Adapt-1 puts the potato in the destination without a cooling action. After 50 actions, the episode ends.

Frozen task knowledge can still produce the wrong milestone order.

Across five seeds, Adapt-1 reaches 93.46% on valid_seen and 90.87% on valid_unseen. Sample standard deviations are 1.13 and 0.97 percentage points.

ALFWorld frozen results by split and task family.
Ten retained evaluation files supply the chart data. Each split gives equal weight to the six task families.

MemHarness studies the same general problem. Prior experience must remain useful when the current state changes. MemHarness and the first Adapt article appeared on 31 July 2026.

MemHarness uses Qwen2.5-7B-Instruct as its policy backbone. MemHarness then uses two benchmark-specific training stages before evaluation.

Stage one uses a supervised cold start. MemHarness uses 400 examples for each benchmark. One subset contains 200 multi-turn interaction trajectories. Another subset contains 200 trajectory-to-memory examples.

AgentGym supplies the source trajectories. GPT-5.1 creates retrieval queries, memory guidance, and target summaries from those trajectories.

MemHarness trains the policy on both subsets for two epochs. GPT-5.1 is used offline for data construction and is absent from reinforcement learning and evaluation.

Stage two uses Group Relative Policy Optimization, or GRPO. A successful episode gives reward 10. A failed episode gives reward 0. A small format reward also applies.

During GRPO, the memory bank starts empty and grows from the policy's successful and failed trajectories. Each memory contains a natural-language principle and its source observation.

During training, MemHarness retains half of the generated trajectories for memory distillation. When possible, it balances successful and failed trajectories. MemHarness also tracks memory utility and removes low-utility entries.

BGE-M3 retrieves the top three memories. For each action, the policy compares each source state with the current history. Before action generation, the policy can retain, revise, or reject a memory.

ALFWorld ablations show the effect of each stage. A cold-start model scores 7.6%. GRPO without memory scores 76.4%. Raw memory replay scores 70.1%. Full MemHarness scores 85.2%.

Without test-time memory, the trained MemHarness policy scores 83.0%. With retrieval but no reconstruction, its score is 79.6%. A generic language-model reconstruction scores 77.7%.

The out-of-distribution evaluation uses the same trained policy. Room layouts and object placements are unseen during training. Full MemHarness scores 85.9%. Without memory, the score is 83.0%. Raw memory replay scores 76.3%.

MemHarness produces the 85.2% result after both teaching stages. Reinforcement learning supplies most of the task skill. Reconstruction training also improves the policy. Test-time memory adds a smaller gain in the reported ALFWorld ablation.

System or ablation Evaluation Success rate Evaluation path
Adapt-1 valid_seen, five frozen seeds 93.46% Frozen typed evaluator
Adapt-1 valid_unseen, five frozen seeds 90.87% Frozen typed evaluator
MemHarness Main evaluation 85.2% Retrieved memory and policy reconstruction
MemHarness OOD evaluation 85.9% Same trained policy on unseen layouts and placements
MemHarness without test-time memory Main ablation 83.0% Trained policy without retrieval
MemHarness without reconstruction Main ablation 79.6% Retrieved memory without learned reconstruction
MemHarness without reconstruction OOD ablation 82.4% Retrieved memory without learned reconstruction
Generic language-model reconstruction Main ablation 77.7% Generic reconstruction with the actor held fixed
GRPO without memory Main ablation 76.4% Reinforcement-learned policy without test-time memory
Raw memory replay Main ablation 70.1% Unreconstructed retrieved memory

An ablation also tests the source observation stored with each memory. Removing that source state lowers ALFWorld success from 85.2% to 80.0%. A random source state raises the rejection rate from 8.7% to 13.3%.

Adapt-1 also receives teaching before its frozen ALFWorld score. Its teaching uses expert actions, plan construction, and candidate-policy training. Its held-out evaluator makes no language-model call.

These scores provide context only. Training and evaluation paths differ. During evaluation, both methods apply earlier experience to the current state.

Case 03 / Structural revision

Learning causal structure from interventions

CausaLab places the learner in a synthetic laboratory. In each trial, the learner changes one variable and observes the returned transition. After the interventions, the learner predicts a held-out resonance frequency.

500/500pre-clamp predictions within 5 Hz
489/500environment completions
0.930macro all-edge F1
238/500exact complete graphs
1.354mean directed SHD

CausaLab also scores the recovered causal graph. Scoring both outputs separates endpoint prediction from mechanism recovery.

Each Adapt-1 episode starts with fresh task state and fresh Domain state. Adapt-1 follows a fixed intervention schedule and uses the full legal budget.

After each intervention, Adapt-1 receives the measured transition and updates an explicit graph and a numerical equation. A final query reads the completed episode state without another update.

One trace shows a false target relation and its later correction.

Retained-field comparison6nodes_32 / contiguous trials 4 through 7 and final result
CausaLab / false target parent correctedseed 1 · 6nodes_32
CAUSALAB / FALSE TARGET PARENT CORRECTED
seed 1 | 6nodes_32 | contiguous trials 4 through 7

TRIAL 4
request pressure=10
returned pressure 6→10, moisture 70→74, conductivity 56→68,
radiation 32→44, resonanceFreq 297→353

TRIAL 5
request radiation=90
returned radiation 44→90 and resonanceFreq 353→445

TRIAL 6
request conductivity=30
returned conductivity 68→30 and resonanceFreq 445→369

TRIAL 7
request moisture=10
returned moisture 74→10, conductivity 30→-34,
resonanceFreq 369→241

FINAL
visible conductivity 260 | moisture 119 | ph 57 | pressure 57 | radiation 189
predicted 1087.9999999999993 | truth 1088.0 | task_success true
all-edge F1 1.0 | SHD 0
  1. 1
    An early intervention changes several variables. The returned transition supplies evidence for the projected model.
  2. 2
    A later transition conflicts with the early relation. The graph revision described below follows this measurement.
  3. 3
    The final endpoint and graph are correct. The prediction matches 1088 Hz with F1 1 and SHD 0.
The dark surface contains requested interventions, returned transitions, visible values, and final outcomes. Graph add/remove language in the surrounding prose is a post-hoc comparison across retained projections.

After trial 4, the projected equation includes moisture as a direct parent of frequency. Later transitions conflict with that relation. An audit comparison of successive retained projections locates the revision after trial 6.

After trial 4, the comparison adds moisture with coefficient -3.999999999988816. After trial 5, it adds radiation → resonanceFreq and retains the moisture term at -1.0000000000002802.

Trial 6 removes the moisture edge and adds a pressure edge. The trial 7 comparison adds pressure → conductivity. At the end, all eight edges are correct for this episode.

The audit comparison supplies the graph-change terms. Candidate-model scores and intermediate readiness do not appear in the retained material.

A second trace separates numerical accuracy, graph accuracy, and environment completion.

Retained failure excerpt6nodes_2 / numerical prediction, structure, and public action result
CausaLab / reactor clampseed 1 · 6nodes_2
CAUSALAB / EXACT NUMBER, STRUCTURAL ERROR, REACTOR CLAMP
seed 1 | 6nodes_2

FINAL MODEL
resonanceFreq = base + c_density*density + c_temperatureC*temperatureC
base 16.000000001304898
c_density 0.999999999993735
c_temperatureC 3.0000000000002593
all-edge precision 0.6666666666666666 | recall 1.0 | F1 0.8 | SHD 4

HELD-OUT CASE
visible conductivity 1497 | density 247 | moisture 62 |
pressure 490 | temperatureC 3789
predicted_frequency 11630.000000000047
true_frequency 11630.0 | absolute_error 4.729372449219227E-11

ACTION STEP 47
You said: {'value': 11630.000000000047}
Allowable range: 0 to 10,000 Hz
Current resonance frequency: 10000 Hz | Frequency set to 10000 Hz
task_success false
  1. 1
    The equation retains graph errors. The final model has F1 0.8 and SHD 4.
  2. 2
    The number is precise. The held-out error is below floating-point display precision.
  3. 3
    The public action is clamped. The environment caps the submitted frequency at 10,000 Hz and records failure.
Exact final-model, held-out, submitted-action, and environment-outcome fields. Evidence identifiers are omitted.

The numerical prediction matches the target to floating-point precision. However, four directed errors remain in the recovered graph. CausaLab clamps the submitted value and records failure.

These results describe different parts of the episode. Here, the equation predicts the target. Four graph errors remain, and CausaLab rejects the submitted action.

A third retained sequence shows hypothesis growth on a seven-node task.

Retained hypothesis-growth excerpt7nodes_6 / trials 1 through 7 and final held-out result
CausaLab / 7nodes_6 hypothesis growthseed 1 · selected retained fields
CAUSALAB / 7NODES_6 HYPOTHESIS GROWTH
seed 1 | requested interventions and returned transitions

TRIAL 1
request conductivity=90
returned conductivity 28→90; quantumSize 41→103; ph 150→336;
radiation 42→104; resonanceFreq 326→698

TRIAL 2
request moisture=70
returned moisture 18→70; temperatureC 53→157; resonanceFreq 698→802

TRIAL 3
request ph=70
returned ph 336→70; resonanceFreq 802→536

TRIAL 4
request quantumSize=30
returned quantumSize 103→30; resonanceFreq 536→317

TRIALS 5–6
request radiation=90; then temperatureC=90
returned radiation 104→90; resonanceFreq stays 317;
then temperatureC 157→90; resonanceFreq stays 317

TRIAL 7
request conductivity=70
returned conductivity 90→70; quantumSize 30→10; ph 70→10;
radiation 90→70; resonanceFreq 317→197

FINAL
visible conductivity 27 | moisture 62 | ph 138 | quantumSize 110 |
radiation 45 | temperatureC 178
predicted 608.999999999998 | truth 609.0 | task_success true
all-edge F1 1.0 | SHD 0
  1. 1
    The first intervention changes several variables. The returned transition begins the episode evidence sequence.
  2. 2
    A later intervention exposes upstream effects. Conductivity changes quantum size, pH, radiation, and resonance frequency.
  3. 3
    The final endpoint and graph are correct. The prediction matches 609 Hz with F1 1 and SHD 0.
Requested values, returned transitions, and final fields are retained. Trials 8 through 24 are omitted; graph-growth wording remains in the paper-background audit text below.

Audit comparisons across its projected models show the target equation grow from a base term to moisture, pH, and quantum-size terms. After trial 7, the graph adds conductivity → ph, conductivity → quantumSize, and conductivity → radiation.

Trials 8 through 24 are omitted. Requested values, returned transitions, and final fields remain inside the trace. The graph-growth terms come from comparisons across successive retained projections.

Across 500 episodes, all 500 pre-clamp predictions are within 5 Hz. CausaLab accepts 489 submitted answers. Macro all-edge F1 is 0.930.

Adapt-1 recovers the exact full graph in 238 episodes. Mean directed structural Hamming distance, or SHD, is 1.354.

Graph size changes the structural result. Exact recovery is 100% at three nodes. It falls to 10% at seven nodes. Target prediction stays within tolerance across the retained set.

CausaLab endpoint and mechanism results by graph size.
Separate chart series show target prediction, environment completion, edge recovery, exact graph recovery, and SHD.

One external comparison is useful for the seven-node public tasks. With a fixed schedule, Adapt-1 reports 96% environment accuracy, 0.852 all-edge F1, and 3.620 SHD.

CausaLab reports GPT-5.2-high at 64% accuracy, 0.745 F1, and 4.761 SHD. By comparison, the GPT agent selects interventions and can stop early. Adapt-1 uses a fixed full-budget schedule.

These runs test online structural revision from measured interventions. Adapt-1 does not select its experiments.

Case 04 / Contextual reinforcement

Contextual reinforcement from measured outcomes

BOP-Ask tests object-interaction reasoning. Its tasks include paths, projected boxes, grasp points, poses, and spatial relations.

566paired cases
436policy changes
323policy changes with same reward
113policy changes with changed reward

In this observation sample, perception stays fixed. For each case, the harness supplies cached detections and the task input. Adapt-1 selects a declared policy and emits a benchmark response.

After each response, the measured reward updates support for the selected policy in its context. Each response receives an immediate outcome. The sample uses no delayed trajectory credit.

A representative BOP-Ask grasp task with five target points.
This image gives task context and does not show one of the paired cases below.

Two clean-state executions form the sample. Both use the same prompts, ground truth, and cached perception. Both runs share exactly 566 cases.

Policy selection changes in 436 pairs. Reward stays the same in 323 of those pairs. Reward changes in 113 pairs. No pair has the same policy label with a different reward.

Paired outcome Reward same Reward changed
Policy same 130 0
Policy changed 323 113

Among coordinate-bearing pairs, 21 policy changes keep the same geometry and reward. Another 128 policy changes alter the geometry and keep the same reward.

Case one shows a policy-label change with the same output.

Paired retained outputsOrdinal 17 / policy divergence with identical geometry
BOP / ordinal 17same output · rewards 1 / 1
ORDINAL 17

A policy distinct_direct_path | selected_by core_exploration_tie_break
B policy attribute_direct_path | selected_by core_unique_argmax

both prediction:
(1169,945) → (1070,858) → (971,770) →
(872,683) → (773,595) → (674,508)

both distance_error 47.80231056374646 | both reward 1.0
  1. 1
    Policy labels and selection bases differ. One run uses an exploration tie-break and the other uses a unique argmax.
  2. 2
    The emitted route matches. Both policies emit the same six-point path.
  3. 3
    Error and reward also match. Both retained outcomes receive reward 1.
Exact policy, selection, geometry, error, and reward fields from the paired retained rows.

Policy labels differ. Both runs emit the same six-point path. Both runs also receive the same benchmark reward.

Case two shows an output change with the same reward.

Paired retained outputsOrdinal 58 / route divergence with matching reward
BOP / ordinal 58different routes · rewards 1 / 1
ORDINAL 58

A policy  attribute_direct_path | selected_by core_unique_argmax
A:        (753,345) → (795,372) → (836,399) →
          (878,427) → (920,454) → (962,481)
A distance_error 57.22142351080793 | reward 1.0

B policy  attribute_clearance_path | selected_by core_unique_argmax
B:        (753,345) → (755,434) → (776,492) →
          (818,519) → (880,515) → (962,481)
B distance_error 85.89788172440417 | reward 1.0
  1. 1
    The selected policies differ. Both selections use a unique argmax.
  2. 2
    The emitted geometry changes. The paths share endpoints and follow different routes.
  3. 3
    Both routes pass the benchmark. The retained errors differ while both rewards remain 1.
Exact policy names, selection fields, coordinates, errors, and rewards from the paired retained rows.

Both paths have the same start and end points. One policy emits a direct route. Run B emits a curved route. Both outputs receive reward 1.

Case three shows two different failed outputs.

Paired retained outputsOrdinal 133 / different projected boxes with matching failure
BOP / ordinal 133different boxes · rewards 0 / 0
ORDINAL 133

A policy attribute_perspective_cuboid | action_transition predicted
A bbox [[714,990],[996,990],[996,840],[714,840],
       [680,963],[962,963],[962,813],[680,813]]

B policy confidence_box_cuboid | action_transition predicted
B bbox [[375,423],[900,423],[900,111],[375,111],
       [375,423],[900,423],[900,111],[375,111]]

A action_count 8 | reward 0.0
B action_count 8 | reward 0.0
  1. 1
    The selected cuboid policies differ. Both retained rows record a predicted action transition.
  2. 2
    The eight-point outputs differ. Each run emits a distinct projected box.
  3. 3
    Both outputs fail. Each response contains eight actions and receives reward 0.
Exact policy, transition, complete geometry, action-count, and reward fields from the paired retained rows.

Run A and run B project different boxes. Both receive reward 0. Different policies can still receive the same reward.

Two paired BOP outputs show route and projected-box divergence.
Two successful routes appear on the left. Two failed projected boxes appear on the right.

Here, we can compare policy, output, and reward as separate fields. A policy change does not always change the output or reward.

A read-only explanation shows the evidence state for an earlier decision. At the cold decision, all priors are equal.

Read-only decision explanationCold state / equal priors and explicit ignorance
BOP / cold trajectory decisionstate version 1 · sealed
COLD DECISION
state_version 1 | sealed true
has_learned_state false
feedback_policy samples 0 | structured_transition samples 0
retrieved_memories 0 | supporting_memories 0 | counterevidence_memories 0
novelty 1.000000 | out_of_distribution true

trajectory_direct    | policy attribute_direct_path    | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072
trajectory_distinct  | policy distinct_direct_path     | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072
trajectory_clearance | policy attribute_clearance_path | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072

mode ucb | tie_size 3
selected_policy attribute_direct_path | selected_hypothesis trajectory_direct
selected_by core_exploration_tie_break

posterior_alpha 1 | posterior_beta 1 | posterior_std 0.288675
calibration accumulating | calibration support 0
model_score null | model_weight 0

update_memory_state false
read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: no_matching_group
  1. 1
    No outcome evidence exists. The decision has no learned feedback or retrieved memory.
  2. 2
    The candidates are tied. Equal priors, intervals, and scores lead to the exploration tie-break.
  3. 3
    The explanation is read-only. The sealed decision is read without changing memory state.
Selected fields from the sealed cold explanation response. Unrelated hypotheses, schema fields, and empty arrays are omitted.

No measured policy outcome exists at this point. All three policies have the same expected reward and selection score. Adapt-1 uses the exploration tie-break.

A checkpoint from the same diagnostic stream shows weak evidence after 49 measured outcomes.

Read-only decision explanationCheckpoint 50 / weak evidence remains visible
BOP / checkpoint-50 grasp decisionstate version 97 · sealed
CHECKPOINT 50
checkpoint_target 50 | completed_rows_observed_before_request 100
decision_id_matches true | state_version 97 | sealed true

has_learned_state true | status ready
feedback_policy samples 49 | structured_transition samples 47
active_regime regime-7 | regime_probability 0.938525 | novelty 0.134538

grasp_pca_minor | policy attribute_pca_minor      | expected_reward 0.375705 | selection_score 0.519342 | observations 3 | successes 0
grasp_pca_major | policy attribute_pca_major      | expected_reward 0.342406 | selection_score 0.497427 | observations 2 | successes 0
grasp_learned   | policy learned_normalized_grasp | expected_reward 0.324961 | selection_score 0.475656 | observations 2 | successes 0
grasp_axis      | policy confidence_axis_grasp    | expected_reward 0.293606 | selection_score 0.438027 | observations 2 | successes 0

mode ucb | selected_policy attribute_pca_minor | selected_hypothesis grasp_pca_minor
selected_by core_exploration
warning all_candidates_below_neutral

credible_interval [0,0.777888] | posterior_std 0.205195
calibration accumulating | calibration support 3
supporting_memories 0 | counterevidence_memories 3
model_score null | model_weight 0

read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: insufficient_support
  1. 1
    The state contains measured feedback. The explanation reports 49 outcome samples and the active regime.
  2. 2
    Core selects under weak support. The selected policy preserves the below-neutral warning.
  3. 3
    Uncertainty and counterevidence remain attached. The response retains its interval, calibration state, and three counterevidence records.
Selected fields from the sealed checkpoint-50 response. It concerns a different input from the cold and checkpoint-100 records.

The selected grasp policy has three observations and no successes. The explanation preserves the below-neutral warning and three counterevidence records.

A later checkpoint shows retained outcome support.

Read-only decision explanationCheckpoint 100 / retained contextual support
BOP / checkpoint-100 spatial decisionstate version 200 · sealed
CHECKPOINT 100
checkpoint_target 100 | completed_rows_observed_before_request 100
decision_id_matches true | state_version 200 | sealed true

has_learned_state true | status ready
feedback_policy samples 99 | structured_transition samples 100
active_regime regime-6 | regime_probability 0.264833 | novelty 0.287247

spatial_attribute  | policy attribute_centers  | expected_reward 0.750000 | selection_score 0.795812 | observations 12 | successes 11
spatial_confidence | policy confidence_centers | expected_reward 0.606246 | selection_score 0.788030 | observations 1  | successes 1
spatial_distinct   | policy distinct_centers   | expected_reward 0.500000 | selection_score 0.702072 | observations 0  | successes 0

mode ucb | selected_policy attribute_centers | selected_hypothesis spatial_attribute
selected_by core_exploration

raw_expected_reward 0.652943 | calibrated_contextual_reward 0.750000
credible_interval [0.252910,1] | posterior_std 0.204099
calibration calibrated | calibration support 12
supporting evidence IDs 12 | counterevidence IDs 1
model_score null | model_weight 0

QUERY-MATCHED RETAINED FEEDBACK
pass01-handal-0002 | relation right | policy attribute_centers | outcome success
pass01-handal-0009 | relation right | policy attribute_centers | outcome success
pass01-handal-0019 | relation right | policy attribute_centers | outcome success
pass01-handal-0060 | relation right | policy attribute_centers | outcome success
pass01-handal-0081 | relation right | policy attribute_centers | outcome success
pass01-handal-0097 | relation right | policy attribute_centers | outcome success

ACTIVATED CONNECTION
source right | target attribute_centers | score 0.7179726538693342
structure_kind statistical_association | learned true | predictively_validated false

read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: insufficient_support
  1. 1
    The selected policy has outcome support. Twelve observations contain eleven successes.
  2. 2
    The query activates retained feedback. Six successful right-relation outcomes match the selected policy.
  3. 3
    The connection keeps its status. The response labels the connection as an unvalidated statistical association.
Selected fields from the sealed checkpoint-100 response. The explanation reads an earlier decision without updating memory.

At this checkpoint, the selected policy has 12 observations and 11 successes in the applicable context. The explanation also returns earlier successful examples for the relation right.

The trace labels that connection as a statistical association. No causal label appears. Read-only access uses the sealed state and does not update it.

These explanation records concern different inputs in a separate diagnostic stream. They are independent of ordinals 17, 58, and 133.

Snapshot State version Retained records Nodes Edges Hyperedges
Cold 1 0 0 0 0
Checkpoint 50 97 96 459 3,955 96
Checkpoint 75 150 149 654 6,109 149
Checkpoint 100 200 199 737 8,157 199
Post 100 201 200 747 8,206 200

The requests read earlier sealed decisions after the diagnostic stream reached 100 completed rows. Decision identifiers matched, read semantics remained frozen, and memory state did not change.

The guided reproduction uses two Domains for this observation sample. Adaptive state starts after cached perception. Results concern policy selection and measured outcome use.

Case 05 / Online transition estimation

Online transition estimation in a mini simulation

This mini sim tests next-step transition learning with generic signals and output series.

30task configurations
120steps per episode
0.051661early emitted-prediction MAE
0.006406late emitted-prediction MAE

For each task, the harness supplies four input signals and five returned output series. One retained sweep contains 30 task configurations. Each configuration uses a fresh Domain and one 120-step episode.

At each step, Adapt-1 makes a prediction before the environment returns the next transition. A controller can use the estimate. Next, the harness returns the actual transition.

Adapt-1 ingests that returned transition as a new sample. No current outcome can enter its own prediction.

Retained transition excerptStream 02 / initial abstention, first emitted estimate, and final step
Transition mini sim / stream 02selected fields · 120-step episode
TRANSITION MINI SIM / STREAM 02
selected fields from one 120-step episode

{
  "step": 0,
  "features": {"SIGNAL_1": 0.02, "SIGNAL_2": 0.07, "SIGNAL_3": -0.06, "SIGNAL_4": -0.09},
  "predicted_deltas": {},
  "selection_reason": "core_abstained",
  "action": {"buy": {}, "sell": {}},
  "actual_deltas": {"STOCK_A": 0.2799999999999976, "STOCK_B": -0.09000000000000341, "STOCK_C": 0.06999999999999318, "STOCK_D": 0.060000000000002274, "STOCK_E": -0.0799999999999983},
  "learner_eligibility": {"accepted": true, "model_version": 1, "sample_count": 1}
}

{
  "step": 1,
  "features": {"SIGNAL_1": -0.2, "SIGNAL_2": 0.1, "SIGNAL_3": 0.0, "SIGNAL_4": 0.02},
  "predicted_deltas": {},
  "selection_reason": "core_abstained",
  "action": {"buy": {}, "sell": {}},
  "actual_deltas": {"STOCK_A": -0.18999999999999773, "STOCK_B": 0.030000000000001137, "STOCK_C": 0.15000000000000568, "STOCK_D": 0.09999999999999432, "STOCK_E": 0.010000000000005116},
  "learner_eligibility": {"accepted": true, "model_version": 2, "sample_count": 2}
}

{
  "step": 2,
  "features": {"SIGNAL_1": -0.08, "SIGNAL_2": 0.12, "SIGNAL_3": 0.09, "SIGNAL_4": -0.01},
  "predicted_deltas": {"values.delta.STOCK_A": 0.03137437365783828, "values.delta.STOCK_B": -0.026521116678597998, "values.delta.STOCK_C": 0.1123192555476018, "values.delta.STOCK_D": 0.08115962777379905, "values.delta.STOCK_E": -0.03239083750894424},
  "selection_reason": "largest_predicted_relative_gain",
  "action": {"buy": {"STOCK_C": 617}, "sell": {}},
  "actual_deltas": {"STOCK_A": -0.16000000000000014, "STOCK_B": -0.23999999999999488, "STOCK_C": -0.020000000000010232, "STOCK_D": 0.0, "STOCK_E": 0.2600000000000051},
  "learner_eligibility": {"accepted": true, "model_version": 3, "sample_count": 3}
}

{
  "step": 119,
  "features": {"SIGNAL_1": 0.16, "SIGNAL_2": -0.07, "SIGNAL_3": -0.0, "SIGNAL_4": 0.12},
  "predicted_deltas": {"values.delta.STOCK_A": -0.05465165678216328, "values.delta.STOCK_B": 0.1679562498421532, "values.delta.STOCK_C": -0.03266424754039907, "values.delta.STOCK_D": 0.032153469264602316, "values.delta.STOCK_E": 0.20801036739769843},
  "selection_reason": "largest_predicted_relative_gain",
  "action": {"buy": {"STOCK_E": 1149}, "sell": {"STOCK_B": 1166}},
  "actual_deltas": {"STOCK_A": -0.07000000000000028, "STOCK_B": 0.1799999999999926, "STOCK_C": -0.030000000000001137, "STOCK_D": 0.04000000000000625, "STOCK_E": 0.20999999999999375},
  "learner_eligibility": {"accepted": true, "model_version": 120, "sample_count": 120}
}
  1. 1
    The first two queries abstain. Their returned transitions are accepted and raise the sample count to two.
  2. 2
    The third query drives an action. The controller buys STOCK_C before its returned delta is observed.
  3. 3
    The final estimate is close. STOCK_E is predicted near 0.208 and returns near 0.210 before ingest 120.
Selected retained fields use generic identifiers. Query and action precede the same-step return; model version and sample count are post-ingest. Steps 3 through 118 are omitted.

Both initial queries abstain because the transition model has no support. Their returned transitions become learning samples. Query three emits estimates for all five series.

At step 2, the estimate for series C is wrong. The returned transition enters the learner after the action. By step 119, the estimate for series E is close to the returned value.

The trace uses generic signal and stock identifiers. Predictions and controller actions precede the returned transition. Model versions and sample counts appear after ingest. Steps 3 through 118 are omitted.

Mean emitted-prediction absolute error, or MAE, is 0.051661 during steps 0 to 19. During steps 100 to 119, MAE falls to 0.006406.

Across the retained sweep, we measure an 87.60% fall in MAE. The calculation excludes the initial abstentions.

The mini sim measures online transition estimation. Delayed reward and long-horizon credit are outside this test. A later article will cover sequential learning with delayed rewards.

Implementation direction / Ontology and setup

Reducing ontology setup work

A Domain gives Adapt-1 a usable language for observations, actions, relations, and feedback. Live interaction fills that language with task-specific evidence.

Supplied structure changes with the task. A small contextual problem may need one input, two actions, and a reward. Symbolic Alchemy needs a richer description of stones, potion effects, topology, objectives, and delayed credit.

We can see that range in the tasks below. The table also includes two nearby retained evaluations.

Task Domain form What the ontology supplies What Adapt-1 learns
Symbolic Alchemy Rich symbolic schema for each run Stone fields; three perceived axes; native potion and cauldron actions; six signed effects; bijection and topology constraints; objective hypotheses; reset and delayed-credit rules Active potion mapping; hidden transition graph; objective evidence; useful action sequences; sequential policy evidence
CausaLab Reusable experiment schema across hidden causal systems Variables; interventions; measurements; equation family and ranges; legal actions; intervention budget Hidden graph; direct parents; equations and coefficients; held-out target prediction
ALFWorld One shared Domain across goals and layouts Rooms; objects; receptacles; command grammar; episode boundaries; feedback channels Reusable plans; object obligations; route preferences; compositional policy
BOP-Ask Two typed Domains by visual task family Input and output schemas; policy hypotheses; query templates; grouped action geometry; outcome feedback Policy choice; outcome associations; retained support across passes
PopGym HigherLowerHard Very thin retained Domain Current rank; guess_lower and guess_higher; binary reward; contextual relation schema Rank-conditioned action values
Transition mini sim Numeric temporal Domain for each task Public features; prediction target; transition timing; returned output series; accounting fields Hidden continuous dynamics; next-transition estimates
RoboSpatial Stateless fold evaluation No persistent Domain state; task-specific perception and relation handling before the query Frozen relation prediction on unseen images

Today, a person/LLM/agent still has to write or check this ontology. That setup creates friction even when Adapt-1 learns from very little task data.

Our next practical step is a Domain generator. A user should be able to provide an interface, a few sample observations, legal actions, and feedback. The generator should produce a valid Domain with useful defaults. It should ask only for missing information. Routine setup should approach zero.

Longer-term work is harder. An internal component must make the first ontology guess from the task stream. It must identify candidate entities, relations, actions, temporal boundaries, and feedback channels. Live transitions must test and revise that guess.

We see this as a direct extension of the current work. Adapt-1 now learns quickly inside a supplied ontology. A generator can remove most manual setup. Internal ontology inference would let the system form a useful task description before online adaptation begins.

Adapt-1app: https://app.reilabs.org/adapt-1

Documentation: https://docs.reilabs.org/docs/welcome

Behavioral Replication Guides: https://github.com/0xreisearch

References

1. Alchemy: https://arxiv.org/abs/2102.02926

2. Pinon et al.: https://arxiv.org/abs/2208.11535

3. AlKhamissi et al.: https://arxiv.org/abs/2112.08360

4. ALFWorld: https://arxiv.org/abs/2010.03768

5. MemHarness: https://arxiv.org/abs/2607.28272

6. CausaLab: https://arxiv.org/abs/2605.26029

7. BOP-Ask: https://arxiv.org/abs/2511.16857