Technical appendix · REMM

Inside REMM

What is stored, what changes, and what reaches the robot.

01 Memory in motion

Step through an accepted disturbance. Select any fact to inspect its JSON.

The object is inside the bowl. The bowl is at the tray. Together they support the inferred object location.

Support graph Select a fact

Arrows name required support. Losing support does not establish a new location.

f3 ProcessFact

{
  "fact_id": "f3",
  "predicate": "at",
  "subject_entity_id": "object_01",
  "object_entity_id": "tray_01",
  "value": "true",
  "confidence": 0.9,
  "evidence_ids": [
    "e_load",
    "e_transport"
  ],
  "depends_on_fact_ids": [
    "f1",
    "f2"
  ],
  "supersedes_fact_ids": [],
  "status": "active",
  "sequence": 3,
  "invalidated_by_fact_id": null,
  "invalidation_reason": null,
  "metadata": {}
}
QueueActive skill unchangedRemaining nominal work

History stays. e_load and e_transport remain. Only the current projection changes.

Illustrative transition following the implementation’s support rules. This example does not run perception or robot control.

Head-camera view with colored object-identity tracks overlaid across the tabletop.
Recorded object traces. Colored paths show tracked object identities in the head-camera image; gaps indicate disconnected tracks. This assisted, partial example uses recompute-on-demand and illustrates the recorded representation, not an autonomous full success. Select the image to enlarge.

02 What a record looks like

An assertion record connects a proposition to its evidence, confidence, support parents, and lifecycle status. Active records form the supported state ; the full record set retains inactive assertions too.

An accepted change and its evidence. Raw model proposals have not reached this stage.

{
  "schema_version": 1,
  "event_id": "e_remove",
  "trajectory_id": "example_01",
  "timestamp_ns": 3000000000,
  "event_type": "external_disturbance",
  "action_type": "remove_from_container",
  "actor": "human",
  "source": "visual_observer",
  "expected": false,
  "entity_ids": [
    "object_01",
    "bowl_01"
  ],
  "predicate_deltas": [
    {
      "fact": {
        "predicate": "inside",
        "arguments": [
          "object_01",
          "bowl_01"
        ]
      },
      "before": "true",
      "after": "false",
      "confidence": 0.9,
      "direct": true,
      "evidence_required": true
    }
  ],
  "evidence_refs": [
    "frame_031",
    "frame_032"
  ],
  "confidence": 0.9,
  "policy_training_eligible": false,
  "memory_training_eligible": true
}

Synthetic examples using implementation fields; some optional/default fields are omitted. These are not trial logs.

How the records join: evidence → event → fact → repair

Follow the references without duplicating the observation. These IDs belong only to the synthetic example above.

Evidenceframe_031 · frame_032Referenced by the accepted event
Evente_removeRecords the negative containment delta
Factf5 supersedes f1f3 and f4 lose their required support
Repairrepair_01Restores the placement obligation

evidence_refs connect events to observations; supersedes_fact_ids connect new assertions to replaced records; depends_on_fact_ids connect conclusions to their support. Recovery history retains the triggering event reference.

Truth ≠ status

value: false is a supported negative assertion. status: invalidated means support was withdrawn.

Identity ≠ appearance

object_01 persists across records. “Blue block” is a language label, not the database key.

Direct ≠ derived

Direct facts name evidence. Derived facts also name required parents in depends_on_fact_ids.

Field conventions

Fact values: true / false / unresolved. Event deltas: true / false / unknown. Fact status: active / superseded / invalidated. Confidence is in [0,1], not calibrated probability. A derived spatial fact inherits the minimum confidence of its parents. Each record has one conjunctive support set: every listed parent is required. There is no alternative-justification manager that keeps the same record active when one parent fails; separately accepted evidence can create a new record.

03 When evidence becomes memory

A candidate event enters the ledger only after its source-specific acceptance checks pass.

Observation→Typed proposal→Source gate→Accepted event→Projection

Scene relation

  • Grounded identities
  • Allowed, noncontradictory relation
  • Confidence ≥ 0.8
  • Two distinct consistent observations
  • Literal visual evidence

Skill completion

  • Command-bound before/after views
  • Visible target and postcondition
  • Successful verdict with evidence
  • No contradiction
  • Current rollout ownership

Rejected evidence

  • No queue advancement
  • Missing detection ≠ absence
  • Ambiguity ≠ success
  • Bounded retries, then review
Verifier and observation paths

Frozen Qwen3-VL-4B-Instruct supports constrained decomposition and verification; OWLv2 provides 2D proposals. Verification uses head/wrist views with temporally separated after frames. The grounded-relation path rejects duplicate/out-of-order observation IDs and requires repeated agreement with a minimum 0.25-second separation; geometric preconfirmation has its own gate. Confidence is a source score, not calibrated probability. Operator-grounded evidence has separate provenance. Ownership checks reject delayed responses for a different rollout, subgoal, or instruction.

04 Memory → instruction

T1 · Revise

  1. Changed relation
  2. Withdraw dependent support
  3. Identify broken obligation
  4. Schedule repair

T2 · Recall order

  1. Retained cue events
  2. Chronological color order
  3. Ground current objects
  4. Compile placement queue

T3 · Retain identity

  1. Target–cover binding
  2. Track cover identity
  3. Resolve current position
  4. Compile lift instruction

The supervisor revises with repair skills while preserving the active instruction until a verified boundary.

At the policy boundary

{
  "instruction": "Place the blue block in the bowl.",
  "other_inputs": [
    "native multi-view observations",
    "proprioception",
    "robot embodiment identifier"
  ]
}

Conceptual input summary, not an API payload. The frozen G0.5 policy receives one instruction with native inputs. The database and support graph stay outside it.

Queue scheduling and readout details

Repairs sharing entities with the active skill go immediately after it; otherwise they follow remaining nominal work. Equivalent repairs are deduplicated. Pending destinations can change, but the active instruction remains fixed until a verified stationary boundary.

Offline T2 uses eight sampled views and frozen-VLM detections; offline T3 initializes a pixel-based binding and uses MIL tracking. These differ from the physical observation paths. Active-fact retrieval includes support ancestors for audit, not learned attention or memory-token injection.

05 Storage and update contract

L · LedgerAppend accepted events
F · FactsPreserve records; revise status
D · DependenciesParent → consequence links
Q · QueueEdit pending work
memory_items (key, type, payload, updated_at)

Typed JSON in SQLite; indexes by type and rollout.

Consistency guarantees
  • Same event ID and payload: idempotent replay. Changed payload under that ID: reject.
  • Out-of-order timestamps: reject in the active projector.
  • Inactive parent: cannot support a new derived fact.
  • Projection failure: roll back database writes and restore process state.
  • Nested updates: savepoints within an immediate transaction.
  • Spatial-support cycle: reject.

Consistency does not establish perception accuracy. An atomic update can still store an incorrectly accepted observation.

06 Details behind the results

Physical protocol · 51 trials

Matched instances, cameras, and workspace; external reset checks. T1 interventions follow full retraction. T2 orders are counterbalanced. T3 moves one cup at a time with pauses. Reported labels are drawn from the saved ledger and operator annotations, with camera evidence where available. Assisted trials remain included; unrecorded assistance is not evidence of autonomy. Model completion claims alone do not establish success.

Progress, decision accuracy, partial completion, and full success are separate. Each rate excludes blank labels from its own denominator. Valid timeouts remain failures; wrong setup or missing required evidence may be unscorable.

Generated videos · 12 T2 + 30 T3 clips

T2: endpoint-conditioned LTX-Video 0.9.8 13B Distilled, using endpoint scene images from OpenAI image generation. T3: deterministic Three.js renders from the VET-Bench shell-game generator. VET-Bench supplies the rendering pipeline, not a generative video model.

Matched clips with method-native observation selection and compute. Labels stay scorer-only. T2 grounding is open-vocabulary; T3 always uses left/middle/right. No G0.5 execution occurs. Additional revision and distractor tracks are unscored. Retained approval records lack reviewer-count and adjudication metadata.

Dependency controls and costs

Explicit invalidation reads broken-goal records. Recompute-on-demand checks saved postconditions against current direct facts. Direct-update-only changes relations without downstream broken-goal detection. A saved deterministic software fixture selects the same repair with explicit invalidation and recomputation; direct-update-only misses it in the state-query selector. A separate direct-reconciliation path can also recover it. This is a software control, not a physical ablation, and does not establish that propagation caused the pooled robot gains. Explicit links provide inspectable support provenance; behavioral superiority over recomputation is not demonstrated.

Time excludes setup and later labeling. Queries count fresh G0.5 action chunks, excluding supervisory VLM inference. Medians use available records only; T2 baseline cost coverage is two of five trials for each baseline. Early failure may appear cheaper than successful recovery. Logical memory traffic counts serialized writes and retrievals separately from camera streams, audit videos, and model weights. Cost efficiencies normalize against the best task-specific median.

Scope and provenance

Examples use synthetic IDs and inspected release-candidate schema fields. Support rules are compiler-supplied, not learned causal discovery. Symbolic locations and 2D boxes are not metric 3D reconstruction. Complete identity loss, appearance shift, and unrestricted occlusion remain limitations. The plotted ledger retains missing fields and some unresolved review-status annotations. Per-trial assistance, scoring provenance, and historical controller settings require reconciliation before a clean autonomous-only comparison or mechanism ablation can be claimed.