REMMRevisable Embodied Memory
for Manipulation

Revanth Krishna Senthilkumaran1,2,*Ruwan Wickramarachchi2Alessandro Oltramari1,2Jonathan Francis1,2

* Work performed during an internship at Bosch Center for AI.

1 Carnegie Mellon University
2 Bosch Center for AI

TL;DR REMM links visual evidence to revisable task state, allowing a robot to update its memory and schedule recovery around a frozen manipulation policy.

Video overview

A three-minute overview of REMM, physical task demonstrations, and evaluation. Subtitles are included in the video.

Abstract

History-dependent manipulation requires a robot to distinguish what happened earlier, what remains true, and what still needs to be done. A previously verified placement can become invalid after an object moves, while the event itself remains part of the task’s history. Retaining observations alone does not resolve these changes.

REMM is an external, evidence-linked memory and supervisory-control layer for a frozen manipulation policy. It separates historical events from revisable current assertions, propagates loss of support through explicit dependencies, and inserts targeted recovery skills into a task queue. The executor receives one bounded instruction at a time. Across 51 physical trials, including operator-assisted runs, REMM raises mean recorded progress over the best baseline from 57% to 91% on intervention recovery and from 43% to 70% on sequence recall. Matched generated-video diagnostics further separate remembering history from grounding it into an executable instruction.

Method overview

Observations and task instructions enter REMM, which sends an atomic instruction to the frozen G0.5 executor and Galaxea R1 Lite robot.
The execution loop. REMM selects the atomic instruction ; the frozen G0.5 policy executes it with native visual and robot-state inputs.

REMM maintains task state outside the policy. Visual evidence updates memory; the supervisor selects the next nominal or recovery skill. The active instruction stays fixed until a verified skill boundary.

REMM internals: task decomposition and skill contracts, source-specific evidence gates, event ledger, assertion records, support graph, dependency invalidation, and recovery scheduling.
Inside REMM. Task, evidence, and memory revision connect accepted observations to a revised skill queue. Select either diagram to enlarge it.
= (,,,)

Hover, focus, or tap a symbol for its meaning.

Evidence-linked state revision

Accepted evidence updates current object relations in while preserving the event history in . Assertion records retain their support links in . When a supporting fact becomes inactive, dependent conclusions are invalidated. An object’s inferred location, for example, loses support when the containment relation no longer holds.

Recovery and verification

A broken completed goal creates or updates a repair skill . A relevant disturbance schedules repair after the active skill; unrelated recovery can wait until the remaining nominal work in is complete. A pending repair may change destination as new evidence arrives. Visual outcome checks determine whether the supervisor advances, retries, or pauses for review.

Illustrative comparison: a disturbed completed condition remains stale without revision, while REMM schedules a repair.
Revising a completed condition. The paper’s teaser illustrates how a disturbance can invalidate earlier task progress and lead to a targeted repair. This is a qualitative example, not an aggregate result.

Physical experiments

We evaluate REMM in 51 physical trials on Galaxea R1 Lite. All three systems use the frozen G0.5 executor: Vanilla G0.5 without external memory, a memory-enabled adaptation of HiMe, and REMM. The comparison measures the complete deployed systems; the memory baseline is a controlled adaptation, not an official reproduction.

T1 and disturbed T2 probe revision; nominal T2 and T3 probe retention and grounding, rather than directly testing dependency propagation. The tasks separate three questions: can a controller revise a previously satisfied condition, retain information that is no longer visible, and turn that information into the correct next instruction?

Physical task sequences for T1a, T1b, T2, and T3. Orange borders identify robot execution and blue borders identify human motion.
Physical evaluation setups. T1a tests recovery with queued work; T1b tests a changed recovery destination. T2 tests order memory and disturbed execution; T3 tests hidden-object identity.

T1a · Recovery with queued work

After a relevant placement and full arm retraction, a human moves the designated object out of its container while unrelated work remains queued. The controller must revise the affected completed condition and schedule a bounded repair without interrupting the active skill.

T1b · Destination changes after pouring

Exactly one designated residual object remains in the bowl after pouring and retraction. Its required final destination is the green tray. Later accepted evidence must revise stale completion and resolve the pending recovery against the current destination.

T2 · Ordered recall and disturbed repeats

Three colored cues are introduced sequentially and then covered. The robot must match their remembered order to current objects and execute the placements in sequence. Predeclared disturbed repeats relocate a pending or already placed target, testing repair without losing the remaining order. A delayed-recall variant inserts two unrelated placements before ordered retrieval.

T3 · Hidden-object rearrangement

The robot observes an object–cup binding and must retain the cover’s identity through a shuffle before selecting the cup to lift. The controlled physical protocol moves one cup at a time with pauses; faster or fully overlapping shuffles are a different observability condition.

On the robot

Physical demonstrations on Galaxea R1 Lite. Choose a system to explore the same four task families.

Showing REMM

Selected demonstrations, not aggregate results. Complete source clips without audio; playback speed is labeled and consistent across systems within each task.

Physical results

REMM has 10 of 17 trials recorded as full completion, compared with 2 of 16 for HiMe and 0 of 18 for Vanilla G0.5. Seven REMM trials remain partial; none is labeled a complete failure. These pooled counts are descriptive because the task mix and sample sizes differ. They include operator-assisted trials and are not unassisted-autonomy success rates. Two REMM full completions have task assistance recorded.

Baseline adaptation and assistance accounting

The local HiMe adaptation retains planner/sentry roles, episodic and procedural stores, and add/update/delete operations. Its configured planner and sentry use Qwen3-VL-4B-Instruct, MiniLM-L6-v2 retrieval, and an eight-frame working context, with task prompts and instruction parsing adapted to G0.5. It is not a reproduction of the original HiMe model stack. Retention versus revision is not a distinction between HiMe and REMM: both can edit memory. REMM makes the supporting records and downstream withdrawal explicit.

The plotted ledger records the following assistance categories. Unrecorded does not mean unassisted. The two assisted REMM full completions are T1b (orange-block nudge) and T2 (help completing strawberry pickup). We preserve the recorded outcomes rather than silently reclassifying them.

Assistance fields in the 51-trial plotted ledger
SystemRecorded noneTask assistanceUnrecorded
Vanilla G0.51206
HiMe adaptation628
REMM1043

A trial-level evidence audit is still needed before reporting an unassisted success rate or attributing the pooled gains to dependency propagation.

Full successPartial completionFailure

T1aQueued-work recovery

Vanilla G0.5
41
n = 5
HiMe
23
n = 5
REMM
22
n = 4

T1bDestination revision

Vanilla G0.5
5
n = 5
HiMe
21
n = 3
REMM
32
n = 5

T2Sequence recall

Vanilla G0.5
23
n = 5
HiMe
5
n = 5
REMM
13
n = 4

T3Hidden-object tracking

Vanilla G0.5
12
n = 3
HiMe
3
n = 3
REMM
4
n = 4
Physical outcomes. Numbers inside bars are trial counts; segment widths show the fraction of each method’s trials. Sample sizes differ. Assisted runs are included. HiMe denotes our adaptation.

Progress and cost, together

T1 pools T1a and T1b. All axes use a 0–100 scale; higher is better.

Vanilla G0.5Memory baseline (HiMe)REMM
T1 performance and memory-system profile20406080100CheckpointprogressDecisionaccuracyNon-failurecompletionTimeefficiencyQueryefficiencyVanilla G0.5 · Checkpoint progress: 39.0%Vanilla G0.5 · Decision accuracy: 30.0%Vanilla G0.5 · Non-failure completion: 90.0%Vanilla G0.5 · Time efficiency: 100.0 pointsVanilla G0.5 · Query efficiency: 100.0 pointsHiMe · Checkpoint progress: 57.1%HiMe · Decision accuracy: 71.4%HiMe · Non-failure completion: 87.5%HiMe · Time efficiency: 83.7 pointsHiMe · Query efficiency: 90.3 pointsREMM · Checkpoint progress: 91.1%REMM · Decision accuracy: 100.0%REMM · Non-failure completion: 100.0%REMM · Time efficiency: 76.3 pointsREMM · Query efficiency: 81.4 points

Checkpoint progressRecorded task progress, including missing-value fallbacks; see accounting below.

Decision accuracyRecorded high-level choices; denominators vary by task and method.

Non-failure completionPartial or full completion; not strict full-task success.

Time efficiency100 × best task-specific median time / system median time.

Query efficiency100 × best task-specific median query count / system median G0.5 query count; supervisory VLM calls are excluded.

Performance comes with overhead. Time and query efficiencies are relative scores, not seconds or query counts. Progress, decision accuracy, and non-failure completion measure different outcomes. Values reproduce the paper’s radial plot.
Exact values and metric definitions
Physical outcomes · raw counts, separated by task
TaskSystemTrialsFull successPartialFailure
T1aVanilla G0.55041
T1aHiMe5230
T1aREMM4220
T1bVanilla G0.55050
T1bHiMe3021
T1bREMM5320
T2Vanilla G0.55023
T2HiMe5050
T2REMM4130
T3Vanilla G0.53012
T3HiMe3003
T3REMM4400

Progress is different from full success

Checkpoint progress is intended to measure the fraction of task checkpoints satisfied; full success requires all checkpoints. The plotted export uses recorded progress where available, with outcome-based fallback (and T3 decision fallback) for missing progress. Full/partial/failure bars use the recorded outcome labels, not a threshold applied to progress. Relative to the strongest baseline in each family, REMM increases mean progress from 57% to 91% for T1 and from 43% to 70% for T2. REMM completes all four physical T3 trials.

Mean checkpoint progress · T1 pools T1a and T1b
SystemT1T2T3
Vanilla G0.539.0%13.3%33.3%
HiMe57.1%43.3%0.0%
REMM91.1%70.0%100.0%
Decision accuracy / non-failure completion · eligible labeled trials
SystemT1T2T3
Vanilla G0.530.0% / 90.0%0.0% / 40.0%33.3% / 33.3%
HiMe71.4% / 87.5%0.0% / 100.0%0.0% / 0.0%
REMM100.0% / 100.0%25.0% / 100.0%100.0% / 100.0%

Decision accuracy uses recorded high-level-choice labels: recovery timing/destination for T1, complete order for T2, and target-cover choice for T3. Non-failure completion includes both partial and full outcomes, so it is less strict than full-task success. Each rate excludes blank labels from its own denominator; T1 pools T1a and T1b.

Decision correctness · numerator / labeled denominator
SystemT1T2T3
Vanilla G0.53/100/31/3
HiMe adaptation5/70/30/3
REMM9/91/44/4

Decision quality and system cost

A correct instruction can still fail during grasping or placement. We therefore keep decision correctness, checkpoint progress , and full-task success separate. Extra perception and verification also carry a cost: REMM is not the fastest or most query-efficient system on every task, and successful recovery can keep a run active longer than an early baseline stop.

The following efficiencies are relative scores, not seconds or query counts. For each task and cost, the best observed median receives 100; other scores are 100 × best median / system median. Higher means lower cost. Query counts include fresh G0.5 calls only, not planner, sentry, or verifier VLM calls; they are not total system inference cost.

Relative median-cost efficiency · time / G0.5 queries
SystemT1 time / queriesT2 time / queriesT3 time / queries
Vanilla G0.5100.0 / 100.0100.0 / 83.3100.0 / 100.0
HiMe83.7 / 90.378.1 / 83.846.1 / 52.2
REMM76.3 / 81.469.4 / 100.092.9 / 80.0
Available records for both time and G0.5-query medians
SystemT1T2T3
Vanilla G0.510/102/53/3
HiMe adaptation8/82/53/3
REMM9/94/44/4

In particular, T2 baseline cost medians each use only two of five trials. Early stopping, assistance, and missing records limit efficiency comparisons; these scores do not establish lower cost for equal task performance.

Outcome labels come from the saved trial ledger and operator annotations, with external and policy-camera evidence where available; a complete independent re-adjudication is not established by this export. A model’s completion claim, return-home movement, or controller stop does not establish success. Protocol and accounting details →

Generated-video experiments

The matched diagnostic contains 12 T2 sequence clips and 30 T3 cup-shuffle clips. Every method receives the same clips without manifest-derived answers. These experiments evaluate memory readout and high-level command generation, not physical manipulation success.

How the videos were created

T2 · LTX-Video 0.9.8 13B Distilled. Sequence clips use endpoint-conditioned video generation, with endpoint scene images produced using OpenAI image generation.

T3 · VET-Bench shell-game generator. Cup-shuffle clips are deterministic Three.js renders, rather than outputs of a generative video model. Generation prompts, seeds, trajectories, and labels are withheld from the evaluated methods.

Models used to evaluate the videos

The matched comparison includes a no-memory Qwen3-VL-4B-Instruct baseline, HiMe, REMM, and controlled MemER-style and RoboMME-style high-level adaptations. The latter emit canonical language subtasks compatible with a downstream π0.5 executor; they are not official reproductions, and no π0.5 or G0.5 motor inference is run on passive videos.

No-memoryQwen3-VL-4B-Instruct · final frame only

HiMe adaptationEpisodic and procedural memory

MemER-styleVisual-keyframe memory planner

RoboMME-styleSymbolic-subgoal memory planner

REMMEvent memory and persistent identity binding

T2 · Sequence memory · LTX-Video. Remember three cue colors after they are covered, then ground the first cue to a visible object and destination.
T3 · Cover identity · VET-Bench. Retain the hidden target’s cup binding through a deterministic rearrangement and select a lift instruction.
Cue LCS/36Exact sequence/12First object/12Destination/12Grounded command/12

T2 partial Equal-weight mean of five metrics

No-memory38.9%Counts: 19 / 1 / 3 / 11 / 2
HiMe54.4%Counts: 35 / 11 / 1 / 8 / 1
MemER-style55.0%Counts: 36 / 12 / 0 / 9 / 0
RoboMME-style73.9%Counts: 34 / 10 / 6 / 11 / 6
REMM86.1%Counts: 35 / 11 / 11 / 10 / 8

T2 fullAll four fields correct

No-memory0.0%
HiMe0.0%
MemER-style0.0%
RoboMME-style41.7%
REMM66.7%

T3 hidden targetCorrect target selection

No-memory36.7%
HiMe33.3%
MemER-style43.3%
RoboMME-style36.7%
REMM100.0%
Recall, grounding, and full decisions. T2 partial averages five normalized metrics; segment labels are raw counts, in legend order. Full T2 requires all four decision fields on the same clip. There are 12 T2 and 30 T3 clips. The π0.5-compatible adaptations evaluate high-level planning, not motor execution.
Exact values and scoring definitions
Matched T2 diagnostic · raw component counts and aggregate scores
SystemCue LCS /36Exact sequence /12First object /12Destination /12Grounded command /12T2 partialT2 full
No-memory191311238.9%0/12 (0.0%)
HiMe351118154.4%0/12 (0.0%)
MemER-style361209055.0%0/12 (0.0%)
RoboMME-style3410611673.9%5/12 (41.7%)
REMM35111110886.1%8/12 (66.7%)

Cue LCS counts matched elements across 36 demonstrated cues (three per clip). T2 partial is the equal-weight mean of five normalized metrics: 100 × (LCS/36 + exact/12 + first object/12 + destination/12 + command/12) / 5. It is a descriptive composite, not the fraction of fully successful clips; LCS and exact sequence are related recall measures. Full T2 requires the exact sequence, first object, destination, and grounded command to be correct on the same clip.

Matched T3 hidden-target diagnostic · 30 clips
SystemCorrect hidden target
No-memory11/30 (36.7%)
HiMe10/30 (33.3%)
MemER-style13/30 (43.3%)
RoboMME-style11/30 (36.7%)
REMM30/30 (100.0%)

Remembering is not the same as grounding

MemER-style records all 12 T2 sequences, but REMM leads the stricter memory-to-action outcome: 8/12 fully grounded decisions, versus 5/12 for RoboMME-style and 0/12 for the other baselines. Destination alone is visible in the final scene and is not the main memory bottleneck.

On T3, REMM identifies the hidden target in 30/30 clips, compared with 13/30 or fewer for the baselines. Its Wilson 95% interval is 88.6–100%; this deterministic diagnostic does not establish perfect tracking under unrestricted occlusion or appearance change.

Benchmark status and evaluation scope

The scored snapshot contains only these 42 clips. Additional intervention-revision and distractor-robust tracks are specified but are not part of the reported score. The legacy snapshot stores human approval verdicts without reviewer-count or adjudication records; a blinded second review remains required for a fully documented benchmark release.

Ground-truth answers stay scorer-only. The intended evaluation interface uses opaque clip IDs, causal frame access, open-vocabulary T2 grounding, and the same left/middle/right answer space for every T3 clip.

Technical appendix

Inspect JSON records, step through memory revision, and follow the references from visual evidence to a recovery instruction.

Scope and limitations

REMM uses compiler-supplied dependencies, conservative 2D perception, and boundary-only repair. The physical study uses one robot and controlled tabletop interventions. The generated clips isolate memory readout and grounding, not physical task completion. Pooled outcomes include assisted runs and incomplete labels; they do not establish unassisted autonomy or isolate dependency propagation. Offline T3 uses a specialized tracker. Broader testing is needed for unconstrained occlusion, appearance shift, new task semantics, and other robots.