
// The Dynamics of, Sequential, Transitional and Contextual Online Learning
Adapt-1 is available in open beta through the Adapt-1 app. Setup instructions and Domain guidance are in the documentation.
This article shows five learning procedures. Each section states the input, the persistent change, the later output, and the limit of the evidence. Retained run files supply the traces.
Earlier articles describe the architecture in more detail: Introducing Adapt-1 Preview and Adaptive State from Partially Observed Streams.
Adapt-1 was built around a specific objective: useful behavior as quickly as possible from the smallest amount of experience available, update from live outcomes as they arrive, and become queryable throughout that process. The emphasis is on sample-efficient adaptation, low mistake rates during learning, and retaining only the information that remains useful for the task.
Evaluations covered: Symbolic Alchemy, ALFWorld, CausaLab, BOP-Ask and Transition mini sim.
Case 01 / Rapid online adaptation
Rapid online adaptation in Symbolic Alchemy
For this release, we use Adapt-1 as a fast online learner. It must use limited interaction to improve later queries and reduce avoidable mistakes.
Symbolic Alchemy tests this role with a new hidden chemistry in every episode. Each avoidable action uses part of a fixed decision budget.
Later sections separate these behaviors. We start with Alchemy because one run brings several of them together.
Hidden chemistry defines potion effects and transition constraints. It also defines the stone value function. Each episode contains ten trials under one chemistry.
A trial provides new stones and potions and allows 20 decisions. One chemistry therefore remains active across 200 decisions.
Adapt-1 started with empty learned state and no Alchemy task pretraining. Reported results cover episodes 0 through 24 from the retained archive.
| Reported window | Value |
|---|---|
| Episodes | 0 through 24 |
| Live environment decisions | 5,000 |
| Total return | 6,290 |
| Mean return | 251.6 |
| Initial learned state | Empty |
| Alchemy task pretraining | None |
A twenty-sixth retained episode remains outside the reported mean, trace aggregates, and figures.
Adapt-1 kept learning throughout each chemistry. New transitions and rewards updated learned state as the run continued.
Stable structure and hidden chemistry
We supply stable task grammar and ontology through the Domain. It declares 39 public observation fields and 40 native slot actions.
Structural declarations include six signed-axis effects, a bijective assignment, reversible effects, connected topology constraints, and eight objective hypotheses.
Current chemistry stays hidden. Adapt-1 must infer the active potion mapping, transition graph, objective, and useful action sequences from live transitions.
For each live query, Core uses the current public observation and retained evidence to select a native action. The next state and source reward then update the learner.
Each episode begins a fresh session for local chemistry evidence. Persistent learner state continues through the selected 25-episode sequence.

Adaptation within a chemistry
In episode 0, we observe a rapid move from unresolved effects to supported predictions.
EPISODE 0 | TRIAL 1
STEP 0 potion slot 11 → stone slot 2
explore | support 0 | state changed
STEP 1 potion slot 7 → stone slot 0
explore | support 0 | state changed
STEP 2 potion slot 3 → stone slot 1
explore | support 0 | state changed
STEP 3 potion slot 0 → stone slot 0
explore | support 0 | state changed
STEP 4 potion slot 6 → stone slot 0
explore | support 1 | state changed
STEP 5 potion slot 1 → stone slot 0
predicted | exact_state | support 1
STEP 6 potion slot 8 → stone slot 0
predicted | transferred_transform | support 1
- 1Unresolved effects. Early selections record exploration and state changes.
- 2Exact-state support. Step 5 uses evidence from the current state.
- 3Transferred prediction. Step 6 applies an observed transform to another state.
We see five unresolved potion choices at the start. Step 5 then uses exact-state evidence.
At step 6, Core transfers an observed transform to another applicable state. Trial 2 then presents new stones and potions under the same chemistry.
EPISODE 0 | TRIAL 2
STEP 20 potion slot 7 → stone slot 1
predicted | transferred_transform | support 1
STEP 21 potion slot 3 → stone slot 1
predicted | constraint_elimination | support 3
STEP 22 potion slot 5 → stone slot 1
predicted | transferred_transform | support 1
STEP 23 stone slot 1 → cauldron
reward 1 | delayed-credit updates 23
STEP 24 potion slot 1 → stone slot 2
predicted | transferred_transform | support 1
STEP 25 potion slot 10 → stone slot 2
predicted | transferred_transform | support 2
STEP 26 stone slot 2 → cauldron
reward 15 | delayed-credit updates 2
- 1Constraint elimination. Three supporting observations narrow the potion effect.
- 2First delayed outcome. Reward 1 updates 23 earlier decisions.
- 3Later delayed outcome. Reward 15 produces two additional credit updates.
At step 21, Core uses constraint elimination after three supporting observations. Later cauldron rewards update earlier choices through sequential credit.
Reward 1 at step 23 produces 23 delayed-credit updates. Reward 15 at step 26 produces two more.
These updates let later outcomes revise decisions that created them.
Across 25 episodes, we see the same pattern. Trial 1 contains 148 explore rows among 500 decisions, or 29.6 percent.
Trial 5 contains five such rows, or 1.0 percent. Trials 6 through 10 contain six rows in 2,500 decisions.

Fast reduction matters because each trial allows only 20 decisions. Later queries can use supported predictions after uncertainty falls.
Across the selected window, 2,143 delayed-credit updates connect later outcomes to earlier choices.
Adapt-1 uses its native substrate and learner stack. Candidate learners can be neural or non-neural, and each starts this task with empty learned state. As evidence accumulates, Core trains the candidates, gates their contribution, keeps supported learners, and suppresses those that lack support. The same live substrate continues to serve queries at unchanged latency during these updates. Introducing Adapt-1 Preview describes the full architecture.
Early queries resolve potion effects. Later queries reuse learned transitions and inferred constraints while delayed outcomes continue to revise policy evidence.
Pinon et al.
Pinon et al. provide the strongest published comparison for task-data and inference efficiency.
Their method learns a Transformer dynamics model from one million Symbolic Alchemy episodes. Each training episode contains 200 environment steps.
Task-specific training therefore covers about 200 million environment steps before evaluation. Learned dynamics then support tree-search planning.
Their strongest configuration uses 10,000 tree expansions for every real environment action. It reaches 251.5 ± 4.5.
| Measure | Adapt-1 | Pinon et al. |
|---|---|---|
| Mean return | 251.6 | 251.5 ± 4.5 |
| Alchemy steps before evaluation | 0 | About 200,000,000 |
| Reported live steps | 5,000 | Reported separately from training |
| Training episodes | None | 1,000,000 |
| Learned state at start | Empty | Trained Transformer dynamics model |
| Planner | Bounded Domain transition planner | 10,000 tree expansions per action |
Adapt-1 reaches the same raw score region after 5,000 live decisions from empty learned state. Alchemy interaction counts differ by about 40,000 times.
This ratio covers task-specific environment interaction. System development and Domain engineering remain outside it.
Pinon reports tree expansions. Adapt records a bounded planner configuration, so inference budgets stay in their published units.
We can compare task data directly and report inference budgets separately. Pinon shows how much Alchemy training previously supported a score near 252.
AlKhamissi et al.
AlKhamissi et al. provide the strongest comparison for representation and behavioral scaffolding.
Their best system reaches 272.85 ± 1.78. Its modified observation counts remaining potions by hue.
A modified action selects a potion hue and a stone. A wrapper then chooses an available potion instance with that hue.
Structured episodic memory stores the prior stone state, potion color, reward, and next stone state. Trials use 15 decisions.
Additional penalties cover null transitions, invalid choices, used stones, and repeated colors on one stone. Their ablations show large changes.
| AlKhamissi et al. configuration | Mean return |
|---|---|
| Modified representation and penalties | 272.85 ± 1.78 |
| Modified representation, no penalties | 236.06 ± 2.18 |
| Canonical input, output, and memory | 158.91 ± 1.60 |
Adapt-1 keeps the public 39-field slot observation, 40 native slot actions, official 20-decision trials, and source reward.
Core selects a specific potion slot. Alchemy learning begins at episode 0.
Domain design gives Adapt-1 more explicit declarative structure. Its action interface provides less behavioral abstraction than AlKhamissi's modified setup.
From their ablations, we observe that representation and reward design can move performance by more than 100 points.
This comparison clarifies how explicit ontology, native action selection, and online learning shape the Adapt result.
Result scope
One retained sequence over 25 official chemistry episodes supports this result. A repeated-seed interval is unavailable.
Published protocols differ in task interface, training, planner, trial length, and reward signal. A matched leaderboard requires a shared protocol.
This run starts from a supplied ontology. Adapt-1 discovers the active chemistry online and uses it during the episode.
Within that scope, Adapt-1 achieves a 251.6 mean after 5,000 live decisions from empty learned state.
We observe quick uncertainty reduction inside each chemistry and continued delayed-credit updates throughout use.
Repeated queries update the learner. Later choices use the resulting evidence before the interaction budget is exhausted.
Case 02 / Online extraction and frozen execution
Online extraction followed by frozen execution
ALFWorld gives an agent household goals, text observations, and legal actions. Some goals require two objects. Other goals require cleaning, heating, or cooling before placement.
Adapt-1 uses a clear teaching boundary. First, the harness supplies expert reference actions and observations. Adapt-1 extracts plans and policy parameters from this stream.
The acquisition stream contains 3,520 official training games and 80,635 decisions. Adapt-1 converts 2,797 successful episodes into reusable plans. Adapt-1 then trains the candidate policy for 45 epochs.
Evaluation starts after this work. Five seeds use the frozen learned state. Each seed runs 140 valid_seen games and 134 valid_unseen games.
During evaluation, the evaluator receives the current observation and the admissible commands. Expert actions and all learner updates stay disabled during evaluation.
Evaluation measures the quality of the earlier extraction and tests retained task structure with the current room state.

One trace shows partial completion in a goal that requires two pillows in one sofa.
ALFWORLD / HELD-OUT PICK TWO SUCCESS
seed-0 | valid_unseen | episode 2
goal: put two pillow in sofa
STEP 0 go to sofa 1
STEP 1 go to armchair 1
STEP 2 take pillow 2 from armchair 1
STEP 3 go to sofa 1
STEP 4 move pillow 2 to sofa 1
STEP 5 go to sidetable 1
STEP 6 look
STEP 7 go to shelf 9
STEP 8 go to sofa 1
STEP 9 go to shelf 5
STEP 10 go to shelf 6
STEP 11 go to shelf 8
STEP 12 go to armchair 1
STEP 13 take pillow 1 from armchair 1
STEP 14 go to sofa 1
STEP 15 move pillow 1 to sofa 1
STEP 9 OBSERVATION
You arrive at sofa 1. On the sofa 1, you see a creditcard 1,
a pillow 2, and a remotecontrol 2.
STEP 13 OBSERVATION
You arrive at armchair 1. On the armchair 1, you see a pillow 1.
STEP 15 OUTCOME
reward 1.0 | terminal true | won true
- 1First placement remains visible. The later sofa observation records pillow 2 at the destination.
- 2The second object is selected. The current observation identifies pillow 1 at the armchair.
- 3The second placement completes the goal. The terminal fields record reward 1 and a win.
Placing pillow 2 changes the room state. Here, we can observe how the evaluator uses that change. A later observation confirms that pillow 2 is already on the sofa. Next, the evaluator selects pillow 1 and completes the task.
Evidence from this trace supports current-state reuse during frozen execution. No retrieved plan identifier or internal completed-object list appears in the trace.
A second trace shows a change in object identity during one task.
ALFWORLD / OBJECT SWITCH AND DESTINATION REUSE
seed-0 | valid_unseen | episode 86
goal: heat some apple and put it in fridge
STEP 4 take apple 3 from sinkbasin 1
STEP 5 clean apple 3 with sinkbasin 1
STEP 6 go to microwave 1
STEP 7 go to fridge 1
STEP 8 open fridge 1
STEP 9 move apple 3 to fridge 1
STEP 10 take apple 1 from fridge 1
STEP 11 go to microwave 1
STEP 12 heat apple 1 with microwave 1
STEP 13 go to fridge 1
STEP 14 move apple 1 to fridge 1
STEP 9 OBSERVATION
You open the fridge 1. The fridge 1 is open. In it, you see a apple 1,
[remaining listed objects omitted]
STEP 10 OBSERVATION
You move the apple 3 to the fridge 1.
STEP 14 OUTCOME
reward 1.0 | terminal true | won true
- 1Apple 3 reaches the fridge. The destination is used before the task finishes.
- 2Execution changes object identity. The evaluator selects apple 1 from the updated fridge state.
- 3The heated apple completes the task. The terminal fields record reward 1 and a win.
Apple 3 reaches the destination. Task execution continues. Next, the evaluator selects apple 1 from the same fridge. After heating apple 1, the evaluator returns it to the fridge.
All five retained seeds contain this exact selected-action sequence.
A failure trace shows a remaining limit in a task that requires a cooled potato in a microwave.
ALFWORLD / COOLING FAILURE BOUNDARY
seed-0 | valid_unseen | episode 15
goal: cool some potato and put it in microwave
STEP 1 take potato 2 from sinkbasin 1
STEP 2 clean potato 2 with sinkbasin 1
STEP 3 go to fridge 1
STEP 4 go to microwave 1
observation: The fridge 1 is closed.
STEP 5 open microwave 1
STEP 6 move potato 2 to microwave 1
STEP 7 go to garbagecan 1
reward 0.0 | terminal false | won false
STEP 4 RANKING
1 go to microwave 1
final 0.7172916666666667 | plan 0.8090087079216964 | milestone 0.95
2 open fridge 1
final 0.565210515069136 | plan 0.5022426592193358 | milestone 0.7270876968049386
3 go to countertop 1
final 0.5526616915422886 | plan 0.5718588831224252 | milestone 0.6686567164179105
4 go to countertop 2
final 0.5143283582089553 | plan 0.5718588831224252 | milestone 0.6686567164179105
5 go to garbagecan 1
final 0.5020398009950249 | plan 0.569521487673977 | milestone 0.6507462686567165
STEP 49 go to stoveburner 1
reward 0.0 | terminal true | won false
- 1The cooling path remains available. The observation says the fridge is closed before the ranking is produced.
- 2The ranking favors the destination. Movement to the microwave ranks above opening the fridge.
- 3The milestone order fails. The episode ends after 50 actions with reward 0.
Here, the evaluator ranks movement to the microwave above opening the fridge. Adapt-1 puts the potato in the destination without a cooling action. After 50 actions, the episode ends.
Frozen task knowledge can still produce the wrong milestone order.
Across five seeds, Adapt-1 reaches 93.46% on valid_seen and 90.87% on valid_unseen. Sample standard deviations are 1.13 and 0.97 percentage points.

MemHarness studies the same general problem. Prior experience must remain useful when the current state changes. MemHarness and the first Adapt article appeared on 31 July 2026.
MemHarness uses Qwen2.5-7B-Instruct as its policy backbone. MemHarness then uses two benchmark-specific training stages before evaluation.
Stage one uses a supervised cold start. MemHarness uses 400 examples for each benchmark. One subset contains 200 multi-turn interaction trajectories. Another subset contains 200 trajectory-to-memory examples.
AgentGym supplies the source trajectories. GPT-5.1 creates retrieval queries, memory guidance, and target summaries from those trajectories.
MemHarness trains the policy on both subsets for two epochs. GPT-5.1 is used offline for data construction and is absent from reinforcement learning and evaluation.
Stage two uses Group Relative Policy Optimization, or GRPO. A successful episode gives reward 10. A failed episode gives reward 0. A small format reward also applies.
During GRPO, the memory bank starts empty and grows from the policy's successful and failed trajectories. Each memory contains a natural-language principle and its source observation.
During training, MemHarness retains half of the generated trajectories for memory distillation. When possible, it balances successful and failed trajectories. MemHarness also tracks memory utility and removes low-utility entries.
BGE-M3 retrieves the top three memories. For each action, the policy compares each source state with the current history. Before action generation, the policy can retain, revise, or reject a memory.
ALFWorld ablations show the effect of each stage. A cold-start model scores 7.6%. GRPO without memory scores 76.4%. Raw memory replay scores 70.1%. Full MemHarness scores 85.2%.
Without test-time memory, the trained MemHarness policy scores 83.0%. With retrieval but no reconstruction, its score is 79.6%. A generic language-model reconstruction scores 77.7%.
The out-of-distribution evaluation uses the same trained policy. Room layouts and object placements are unseen during training. Full MemHarness scores 85.9%. Without memory, the score is 83.0%. Raw memory replay scores 76.3%.
MemHarness produces the 85.2% result after both teaching stages. Reinforcement learning supplies most of the task skill. Reconstruction training also improves the policy. Test-time memory adds a smaller gain in the reported ALFWorld ablation.
| System or ablation | Evaluation | Success rate | Evaluation path |
|---|---|---|---|
| Adapt-1 | valid_seen, five frozen seeds |
93.46% | Frozen typed evaluator |
| Adapt-1 | valid_unseen, five frozen seeds |
90.87% | Frozen typed evaluator |
| MemHarness | Main evaluation | 85.2% | Retrieved memory and policy reconstruction |
| MemHarness | OOD evaluation | 85.9% | Same trained policy on unseen layouts and placements |
| MemHarness without test-time memory | Main ablation | 83.0% | Trained policy without retrieval |
| MemHarness without reconstruction | Main ablation | 79.6% | Retrieved memory without learned reconstruction |
| MemHarness without reconstruction | OOD ablation | 82.4% | Retrieved memory without learned reconstruction |
| Generic language-model reconstruction | Main ablation | 77.7% | Generic reconstruction with the actor held fixed |
| GRPO without memory | Main ablation | 76.4% | Reinforcement-learned policy without test-time memory |
| Raw memory replay | Main ablation | 70.1% | Unreconstructed retrieved memory |
An ablation also tests the source observation stored with each memory. Removing that source state lowers ALFWorld success from 85.2% to 80.0%. A random source state raises the rejection rate from 8.7% to 13.3%.
Adapt-1 also receives teaching before its frozen ALFWorld score. Its teaching uses expert actions, plan construction, and candidate-policy training. Its held-out evaluator makes no language-model call.
These scores provide context only. Training and evaluation paths differ. During evaluation, both methods apply earlier experience to the current state.
Case 03 / Structural revision
Learning causal structure from interventions
CausaLab places the learner in a synthetic laboratory. In each trial, the learner changes one variable and observes the returned transition. After the interventions, the learner predicts a held-out resonance frequency.
CausaLab also scores the recovered causal graph. Scoring both outputs separates endpoint prediction from mechanism recovery.
Each Adapt-1 episode starts with fresh task state and fresh Domain state. Adapt-1 follows a fixed intervention schedule and uses the full legal budget.
After each intervention, Adapt-1 receives the measured transition and updates an explicit graph and a numerical equation. A final query reads the completed episode state without another update.
One trace shows a false target relation and its later correction.
CAUSALAB / FALSE TARGET PARENT CORRECTED
seed 1 | 6nodes_32 | contiguous trials 4 through 7
TRIAL 4
request pressure=10
returned pressure 6→10, moisture 70→74, conductivity 56→68,
radiation 32→44, resonanceFreq 297→353
TRIAL 5
request radiation=90
returned radiation 44→90 and resonanceFreq 353→445
TRIAL 6
request conductivity=30
returned conductivity 68→30 and resonanceFreq 445→369
TRIAL 7
request moisture=10
returned moisture 74→10, conductivity 30→-34,
resonanceFreq 369→241
FINAL
visible conductivity 260 | moisture 119 | ph 57 | pressure 57 | radiation 189
predicted 1087.9999999999993 | truth 1088.0 | task_success true
all-edge F1 1.0 | SHD 0
- 1An early intervention changes several variables. The returned transition supplies evidence for the projected model.
- 2A later transition conflicts with the early relation. The graph revision described below follows this measurement.
- 3The final endpoint and graph are correct. The prediction matches 1088 Hz with F1 1 and SHD 0.
After trial 4, the projected equation includes moisture as a direct parent of frequency. Later transitions conflict with that relation. An audit comparison of successive retained projections locates the revision after trial 6.
After trial 4, the comparison adds moisture with coefficient -3.999999999988816. After trial 5, it adds radiation → resonanceFreq and retains the moisture term at -1.0000000000002802.
Trial 6 removes the moisture edge and adds a pressure edge. The trial 7 comparison adds pressure → conductivity. At the end, all eight edges are correct for this episode.
The audit comparison supplies the graph-change terms. Candidate-model scores and intermediate readiness do not appear in the retained material.
A second trace separates numerical accuracy, graph accuracy, and environment completion.
CAUSALAB / EXACT NUMBER, STRUCTURAL ERROR, REACTOR CLAMP
seed 1 | 6nodes_2
FINAL MODEL
resonanceFreq = base + c_density*density + c_temperatureC*temperatureC
base 16.000000001304898
c_density 0.999999999993735
c_temperatureC 3.0000000000002593
all-edge precision 0.6666666666666666 | recall 1.0 | F1 0.8 | SHD 4
HELD-OUT CASE
visible conductivity 1497 | density 247 | moisture 62 |
pressure 490 | temperatureC 3789
predicted_frequency 11630.000000000047
true_frequency 11630.0 | absolute_error 4.729372449219227E-11
ACTION STEP 47
You said: {'value': 11630.000000000047}
Allowable range: 0 to 10,000 Hz
Current resonance frequency: 10000 Hz | Frequency set to 10000 Hz
task_success false
- 1The equation retains graph errors. The final model has F1 0.8 and SHD 4.
- 2The number is precise. The held-out error is below floating-point display precision.
- 3The public action is clamped. The environment caps the submitted frequency at 10,000 Hz and records failure.
The numerical prediction matches the target to floating-point precision. However, four directed errors remain in the recovered graph. CausaLab clamps the submitted value and records failure.
These results describe different parts of the episode. Here, the equation predicts the target. Four graph errors remain, and CausaLab rejects the submitted action.
A third retained sequence shows hypothesis growth on a seven-node task.
CAUSALAB / 7NODES_6 HYPOTHESIS GROWTH
seed 1 | requested interventions and returned transitions
TRIAL 1
request conductivity=90
returned conductivity 28→90; quantumSize 41→103; ph 150→336;
radiation 42→104; resonanceFreq 326→698
TRIAL 2
request moisture=70
returned moisture 18→70; temperatureC 53→157; resonanceFreq 698→802
TRIAL 3
request ph=70
returned ph 336→70; resonanceFreq 802→536
TRIAL 4
request quantumSize=30
returned quantumSize 103→30; resonanceFreq 536→317
TRIALS 5–6
request radiation=90; then temperatureC=90
returned radiation 104→90; resonanceFreq stays 317;
then temperatureC 157→90; resonanceFreq stays 317
TRIAL 7
request conductivity=70
returned conductivity 90→70; quantumSize 30→10; ph 70→10;
radiation 90→70; resonanceFreq 317→197
FINAL
visible conductivity 27 | moisture 62 | ph 138 | quantumSize 110 |
radiation 45 | temperatureC 178
predicted 608.999999999998 | truth 609.0 | task_success true
all-edge F1 1.0 | SHD 0
- 1The first intervention changes several variables. The returned transition begins the episode evidence sequence.
- 2A later intervention exposes upstream effects. Conductivity changes quantum size, pH, radiation, and resonance frequency.
- 3The final endpoint and graph are correct. The prediction matches 609 Hz with F1 1 and SHD 0.
Audit comparisons across its projected models show the target equation grow from a base term to moisture, pH, and quantum-size terms. After trial 7, the graph adds conductivity → ph, conductivity → quantumSize, and conductivity → radiation.
Trials 8 through 24 are omitted. Requested values, returned transitions, and final fields remain inside the trace. The graph-growth terms come from comparisons across successive retained projections.
Across 500 episodes, all 500 pre-clamp predictions are within 5 Hz. CausaLab accepts 489 submitted answers. Macro all-edge F1 is 0.930.
Adapt-1 recovers the exact full graph in 238 episodes. Mean directed structural Hamming distance, or SHD, is 1.354.
Graph size changes the structural result. Exact recovery is 100% at three nodes. It falls to 10% at seven nodes. Target prediction stays within tolerance across the retained set.

One external comparison is useful for the seven-node public tasks. With a fixed schedule, Adapt-1 reports 96% environment accuracy, 0.852 all-edge F1, and 3.620 SHD.
CausaLab reports GPT-5.2-high at 64% accuracy, 0.745 F1, and 4.761 SHD. By comparison, the GPT agent selects interventions and can stop early. Adapt-1 uses a fixed full-budget schedule.
These runs test online structural revision from measured interventions. Adapt-1 does not select its experiments.
Case 04 / Contextual reinforcement
Contextual reinforcement from measured outcomes
BOP-Ask tests object-interaction reasoning. Its tasks include paths, projected boxes, grasp points, poses, and spatial relations.
In this observation sample, perception stays fixed. For each case, the harness supplies cached detections and the task input. Adapt-1 selects a declared policy and emits a benchmark response.
After each response, the measured reward updates support for the selected policy in its context. Each response receives an immediate outcome. The sample uses no delayed trajectory credit.

Two clean-state executions form the sample. Both use the same prompts, ground truth, and cached perception. Both runs share exactly 566 cases.
Policy selection changes in 436 pairs. Reward stays the same in 323 of those pairs. Reward changes in 113 pairs. No pair has the same policy label with a different reward.
| Paired outcome | Reward same | Reward changed |
|---|---|---|
| Policy same | 130 | 0 |
| Policy changed | 323 | 113 |
Among coordinate-bearing pairs, 21 policy changes keep the same geometry and reward. Another 128 policy changes alter the geometry and keep the same reward.
Case one shows a policy-label change with the same output.
ORDINAL 17
A policy distinct_direct_path | selected_by core_exploration_tie_break
B policy attribute_direct_path | selected_by core_unique_argmax
both prediction:
(1169,945) → (1070,858) → (971,770) →
(872,683) → (773,595) → (674,508)
both distance_error 47.80231056374646 | both reward 1.0
- 1Policy labels and selection bases differ. One run uses an exploration tie-break and the other uses a unique argmax.
- 2The emitted route matches. Both policies emit the same six-point path.
- 3Error and reward also match. Both retained outcomes receive reward 1.
Policy labels differ. Both runs emit the same six-point path. Both runs also receive the same benchmark reward.
Case two shows an output change with the same reward.
ORDINAL 58
A policy attribute_direct_path | selected_by core_unique_argmax
A: (753,345) → (795,372) → (836,399) →
(878,427) → (920,454) → (962,481)
A distance_error 57.22142351080793 | reward 1.0
B policy attribute_clearance_path | selected_by core_unique_argmax
B: (753,345) → (755,434) → (776,492) →
(818,519) → (880,515) → (962,481)
B distance_error 85.89788172440417 | reward 1.0
- 1The selected policies differ. Both selections use a unique argmax.
- 2The emitted geometry changes. The paths share endpoints and follow different routes.
- 3Both routes pass the benchmark. The retained errors differ while both rewards remain 1.
Both paths have the same start and end points. One policy emits a direct route. Run B emits a curved route. Both outputs receive reward 1.
Case three shows two different failed outputs.
ORDINAL 133
A policy attribute_perspective_cuboid | action_transition predicted
A bbox [[714,990],[996,990],[996,840],[714,840],
[680,963],[962,963],[962,813],[680,813]]
B policy confidence_box_cuboid | action_transition predicted
B bbox [[375,423],[900,423],[900,111],[375,111],
[375,423],[900,423],[900,111],[375,111]]
A action_count 8 | reward 0.0
B action_count 8 | reward 0.0
- 1The selected cuboid policies differ. Both retained rows record a predicted action transition.
- 2The eight-point outputs differ. Each run emits a distinct projected box.
- 3Both outputs fail. Each response contains eight actions and receives reward 0.
Run A and run B project different boxes. Both receive reward 0. Different policies can still receive the same reward.

Here, we can compare policy, output, and reward as separate fields. A policy change does not always change the output or reward.
A read-only explanation shows the evidence state for an earlier decision. At the cold decision, all priors are equal.
COLD DECISION
state_version 1 | sealed true
has_learned_state false
feedback_policy samples 0 | structured_transition samples 0
retrieved_memories 0 | supporting_memories 0 | counterevidence_memories 0
novelty 1.000000 | out_of_distribution true
trajectory_direct | policy attribute_direct_path | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072
trajectory_distinct | policy distinct_direct_path | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072
trajectory_clearance | policy attribute_clearance_path | expected_reward 0.500000 | credible_interval [0,1] | selection_score 0.702072
mode ucb | tie_size 3
selected_policy attribute_direct_path | selected_hypothesis trajectory_direct
selected_by core_exploration_tie_break
posterior_alpha 1 | posterior_beta 1 | posterior_std 0.288675
calibration accumulating | calibration support 0
model_score null | model_weight 0
update_memory_state false
read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: no_matching_group
- 1No outcome evidence exists. The decision has no learned feedback or retrieved memory.
- 2The candidates are tied. Equal priors, intervals, and scores lead to the exploration tie-break.
- 3The explanation is read-only. The sealed decision is read without changing memory state.
No measured policy outcome exists at this point. All three policies have the same expected reward and selection score. Adapt-1 uses the exploration tie-break.
A checkpoint from the same diagnostic stream shows weak evidence after 49 measured outcomes.
CHECKPOINT 50
checkpoint_target 50 | completed_rows_observed_before_request 100
decision_id_matches true | state_version 97 | sealed true
has_learned_state true | status ready
feedback_policy samples 49 | structured_transition samples 47
active_regime regime-7 | regime_probability 0.938525 | novelty 0.134538
grasp_pca_minor | policy attribute_pca_minor | expected_reward 0.375705 | selection_score 0.519342 | observations 3 | successes 0
grasp_pca_major | policy attribute_pca_major | expected_reward 0.342406 | selection_score 0.497427 | observations 2 | successes 0
grasp_learned | policy learned_normalized_grasp | expected_reward 0.324961 | selection_score 0.475656 | observations 2 | successes 0
grasp_axis | policy confidence_axis_grasp | expected_reward 0.293606 | selection_score 0.438027 | observations 2 | successes 0
mode ucb | selected_policy attribute_pca_minor | selected_hypothesis grasp_pca_minor
selected_by core_exploration
warning all_candidates_below_neutral
credible_interval [0,0.777888] | posterior_std 0.205195
calibration accumulating | calibration support 3
supporting_memories 0 | counterevidence_memories 3
model_score null | model_weight 0
read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: insufficient_support
- 1The state contains measured feedback. The explanation reports 49 outcome samples and the active regime.
- 2Core selects under weak support. The selected policy preserves the below-neutral warning.
- 3Uncertainty and counterevidence remain attached. The response retains its interval, calibration state, and three counterevidence records.
The selected grasp policy has three observations and no successes. The explanation preserves the below-neutral warning and three counterevidence records.
A later checkpoint shows retained outcome support.
CHECKPOINT 100
checkpoint_target 100 | completed_rows_observed_before_request 100
decision_id_matches true | state_version 200 | sealed true
has_learned_state true | status ready
feedback_policy samples 99 | structured_transition samples 100
active_regime regime-6 | regime_probability 0.264833 | novelty 0.287247
spatial_attribute | policy attribute_centers | expected_reward 0.750000 | selection_score 0.795812 | observations 12 | successes 11
spatial_confidence | policy confidence_centers | expected_reward 0.606246 | selection_score 0.788030 | observations 1 | successes 1
spatial_distinct | policy distinct_centers | expected_reward 0.500000 | selection_score 0.702072 | observations 0 | successes 0
mode ucb | selected_policy attribute_centers | selected_hypothesis spatial_attribute
selected_by core_exploration
raw_expected_reward 0.652943 | calibrated_contextual_reward 0.750000
credible_interval [0.252910,1] | posterior_std 0.204099
calibration calibrated | calibration support 12
supporting evidence IDs 12 | counterevidence IDs 1
model_score null | model_weight 0
QUERY-MATCHED RETAINED FEEDBACK
pass01-handal-0002 | relation right | policy attribute_centers | outcome success
pass01-handal-0009 | relation right | policy attribute_centers | outcome success
pass01-handal-0019 | relation right | policy attribute_centers | outcome success
pass01-handal-0060 | relation right | policy attribute_centers | outcome success
pass01-handal-0081 | relation right | policy attribute_centers | outcome success
pass01-handal-0097 | relation right | policy attribute_centers | outcome success
ACTIVATED CONNECTION
source right | target attribute_centers | score 0.7179726538693342
structure_kind statistical_association | learned true | predictively_validated false
read_semantics frozen | memory_state_mutated false
induced_structure disabled: structure_learning_disabled
transition_plan abstained: planning_disabled
transition_prediction abstained: insufficient_support
- 1The selected policy has outcome support. Twelve observations contain eleven successes.
- 2The query activates retained feedback. Six successful right-relation outcomes match the selected policy.
- 3The connection keeps its status. The response labels the connection as an unvalidated statistical association.
At this checkpoint, the selected policy has 12 observations and 11 successes in the applicable context. The explanation also returns earlier successful examples for the relation right.
The trace labels that connection as a statistical association. No causal label appears. Read-only access uses the sealed state and does not update it.
These explanation records concern different inputs in a separate diagnostic stream. They are independent of ordinals 17, 58, and 133.
| Snapshot | State version | Retained records | Nodes | Edges | Hyperedges |
|---|---|---|---|---|---|
| Cold | 1 | 0 | 0 | 0 | 0 |
| Checkpoint 50 | 97 | 96 | 459 | 3,955 | 96 |
| Checkpoint 75 | 150 | 149 | 654 | 6,109 | 149 |
| Checkpoint 100 | 200 | 199 | 737 | 8,157 | 199 |
| Post 100 | 201 | 200 | 747 | 8,206 | 200 |
The requests read earlier sealed decisions after the diagnostic stream reached 100 completed rows. Decision identifiers matched, read semantics remained frozen, and memory state did not change.
The guided reproduction uses two Domains for this observation sample. Adaptive state starts after cached perception. Results concern policy selection and measured outcome use.
Case 05 / Online transition estimation
Online transition estimation in a mini simulation
This mini sim tests next-step transition learning with generic signals and output series.
For each task, the harness supplies four input signals and five returned output series. One retained sweep contains 30 task configurations. Each configuration uses a fresh Domain and one 120-step episode.
At each step, Adapt-1 makes a prediction before the environment returns the next transition. A controller can use the estimate. Next, the harness returns the actual transition.
Adapt-1 ingests that returned transition as a new sample. No current outcome can enter its own prediction.
TRANSITION MINI SIM / STREAM 02
selected fields from one 120-step episode
{
"step": 0,
"features": {"SIGNAL_1": 0.02, "SIGNAL_2": 0.07, "SIGNAL_3": -0.06, "SIGNAL_4": -0.09},
"predicted_deltas": {},
"selection_reason": "core_abstained",
"action": {"buy": {}, "sell": {}},
"actual_deltas": {"STOCK_A": 0.2799999999999976, "STOCK_B": -0.09000000000000341, "STOCK_C": 0.06999999999999318, "STOCK_D": 0.060000000000002274, "STOCK_E": -0.0799999999999983},
"learner_eligibility": {"accepted": true, "model_version": 1, "sample_count": 1}
}
{
"step": 1,
"features": {"SIGNAL_1": -0.2, "SIGNAL_2": 0.1, "SIGNAL_3": 0.0, "SIGNAL_4": 0.02},
"predicted_deltas": {},
"selection_reason": "core_abstained",
"action": {"buy": {}, "sell": {}},
"actual_deltas": {"STOCK_A": -0.18999999999999773, "STOCK_B": 0.030000000000001137, "STOCK_C": 0.15000000000000568, "STOCK_D": 0.09999999999999432, "STOCK_E": 0.010000000000005116},
"learner_eligibility": {"accepted": true, "model_version": 2, "sample_count": 2}
}
{
"step": 2,
"features": {"SIGNAL_1": -0.08, "SIGNAL_2": 0.12, "SIGNAL_3": 0.09, "SIGNAL_4": -0.01},
"predicted_deltas": {"values.delta.STOCK_A": 0.03137437365783828, "values.delta.STOCK_B": -0.026521116678597998, "values.delta.STOCK_C": 0.1123192555476018, "values.delta.STOCK_D": 0.08115962777379905, "values.delta.STOCK_E": -0.03239083750894424},
"selection_reason": "largest_predicted_relative_gain",
"action": {"buy": {"STOCK_C": 617}, "sell": {}},
"actual_deltas": {"STOCK_A": -0.16000000000000014, "STOCK_B": -0.23999999999999488, "STOCK_C": -0.020000000000010232, "STOCK_D": 0.0, "STOCK_E": 0.2600000000000051},
"learner_eligibility": {"accepted": true, "model_version": 3, "sample_count": 3}
}
{
"step": 119,
"features": {"SIGNAL_1": 0.16, "SIGNAL_2": -0.07, "SIGNAL_3": -0.0, "SIGNAL_4": 0.12},
"predicted_deltas": {"values.delta.STOCK_A": -0.05465165678216328, "values.delta.STOCK_B": 0.1679562498421532, "values.delta.STOCK_C": -0.03266424754039907, "values.delta.STOCK_D": 0.032153469264602316, "values.delta.STOCK_E": 0.20801036739769843},
"selection_reason": "largest_predicted_relative_gain",
"action": {"buy": {"STOCK_E": 1149}, "sell": {"STOCK_B": 1166}},
"actual_deltas": {"STOCK_A": -0.07000000000000028, "STOCK_B": 0.1799999999999926, "STOCK_C": -0.030000000000001137, "STOCK_D": 0.04000000000000625, "STOCK_E": 0.20999999999999375},
"learner_eligibility": {"accepted": true, "model_version": 120, "sample_count": 120}
}
- 1The first two queries abstain. Their returned transitions are accepted and raise the sample count to two.
- 2The third query drives an action. The controller buys STOCK_C before its returned delta is observed.
- 3The final estimate is close. STOCK_E is predicted near 0.208 and returns near 0.210 before ingest 120.
Both initial queries abstain because the transition model has no support. Their returned transitions become learning samples. Query three emits estimates for all five series.
At step 2, the estimate for series C is wrong. The returned transition enters the learner after the action. By step 119, the estimate for series E is close to the returned value.
The trace uses generic signal and stock identifiers. Predictions and controller actions precede the returned transition. Model versions and sample counts appear after ingest. Steps 3 through 118 are omitted.
Mean emitted-prediction absolute error, or MAE, is 0.051661 during steps 0 to 19. During steps 100 to 119, MAE falls to 0.006406.
Across the retained sweep, we measure an 87.60% fall in MAE. The calculation excludes the initial abstentions.
The mini sim measures online transition estimation. Delayed reward and long-horizon credit are outside this test. A later article will cover sequential learning with delayed rewards.
Implementation direction / Ontology and setup
Reducing ontology setup work
A Domain gives Adapt-1 a usable language for observations, actions, relations, and feedback. Live interaction fills that language with task-specific evidence.
Supplied structure changes with the task. A small contextual problem may need one input, two actions, and a reward. Symbolic Alchemy needs a richer description of stones, potion effects, topology, objectives, and delayed credit.
We can see that range in the tasks below. The table also includes two nearby retained evaluations.
| Task | Domain form | What the ontology supplies | What Adapt-1 learns |
|---|---|---|---|
| Symbolic Alchemy | Rich symbolic schema for each run | Stone fields; three perceived axes; native potion and cauldron actions; six signed effects; bijection and topology constraints; objective hypotheses; reset and delayed-credit rules | Active potion mapping; hidden transition graph; objective evidence; useful action sequences; sequential policy evidence |
| CausaLab | Reusable experiment schema across hidden causal systems | Variables; interventions; measurements; equation family and ranges; legal actions; intervention budget | Hidden graph; direct parents; equations and coefficients; held-out target prediction |
| ALFWorld | One shared Domain across goals and layouts | Rooms; objects; receptacles; command grammar; episode boundaries; feedback channels | Reusable plans; object obligations; route preferences; compositional policy |
| BOP-Ask | Two typed Domains by visual task family | Input and output schemas; policy hypotheses; query templates; grouped action geometry; outcome feedback | Policy choice; outcome associations; retained support across passes |
| PopGym HigherLowerHard | Very thin retained Domain | Current rank; guess_lower and guess_higher; binary reward; contextual relation schema |
Rank-conditioned action values |
| Transition mini sim | Numeric temporal Domain for each task | Public features; prediction target; transition timing; returned output series; accounting fields | Hidden continuous dynamics; next-transition estimates |
| RoboSpatial | Stateless fold evaluation | No persistent Domain state; task-specific perception and relation handling before the query | Frozen relation prediction on unseen images |
Today, a person/LLM/agent still has to write or check this ontology. That setup creates friction even when Adapt-1 learns from very little task data.
Our next practical step is a Domain generator. A user should be able to provide an interface, a few sample observations, legal actions, and feedback. The generator should produce a valid Domain with useful defaults. It should ask only for missing information. Routine setup should approach zero.
Longer-term work is harder. An internal component must make the first ontology guess from the task stream. It must identify candidate entities, relations, actions, temporal boundaries, and feedback channels. Live transitions must test and revise that guess.
We see this as a direct extension of the current work. Adapt-1 now learns quickly inside a supplied ontology. A generator can remove most manual setup. Internal ontology inference would let the system form a useful task description before online adaptation begins.
Adapt-1app: https://app.reilabs.org/adapt-1
Documentation: https://docs.reilabs.org/docs/welcome
Behavioral Replication Guides: https://github.com/0xreisearch
References
1. Alchemy: https://arxiv.org/abs/2102.02926
2. Pinon et al.: https://arxiv.org/abs/2208.11535
3. AlKhamissi et al.: https://arxiv.org/abs/2112.08360
4. ALFWorld: https://arxiv.org/abs/2010.03768
5. MemHarness: https://arxiv.org/abs/2607.28272
6. CausaLab: https://arxiv.org/abs/2605.26029
7. BOP-Ask: https://arxiv.org/abs/2511.16857