Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/Where Memory Belongs: Ledger, an…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

DATE September 28, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS TANGUY DIEUDONNÉ, JACK B. JEDLICKI, HENG YANGARXIV 2609.34554
// KEY TAKEAWAYS4 ITEMS
  1. 01Splitting Memory Between Inside and Outside the Policy
  2. 02Runtime Memory Routing Without a Task-Level Router
  3. 03External Object Memory Dramatically Improves Permanence and Reference
  4. 04Step-Boundary Replanning via Proprioception
// SUMMARY

Authors: Tanguy Dieudonné, Jack B. Jedlicki, Heng Yang (Harvard / ETH Zürich)


1. Key Themes

Splitting Memory Between Inside and Outside the Policy

The paper's central thesis is that robot memory should not live in one place. Short-term perceptual memory (repetition counts, timing, motion patterns) belongs inside the VLA policy as history-conditioned features. Long-term object memory (where objects rest, what covers them, event history) belongs outside the policy as an explicit, readable record. This split is operationalized as LEDGER: a SAM3 tracker maintains per-object stationary intervals and occlusion events, a VLM generates demonstration transcripts, and an LLM planner reads the record at step boundaries to decide what the policy should act on. The result: 64.3% average success on RoboMME vs. 45.9% for the strongest prior method, using a single set of policy weights (Table 1, Section 4.1).

Runtime Memory Routing Without a Task-Level Router

Instead of pre-selecting a memory mechanism per task (as DIRECT does by training a router on per-task success labels), LEDGER lets the LLM planner decide at runtime — from the instruction and the ledger alone — whether to ground a coordinate target or defer to the policy's internal memory. The planner "defers on every PatternLock and RouteStick episode and on two thirds of StopCube episodes" where deferring succeeds more often, and grounds targets where the instruction names an object (Section 4.2). This eliminates the need for task metadata at deployment time.

External Object Memory Dramatically Improves Permanence and Reference

The largest gains come on tasks requiring persistent spatial state. On Permanence (objects hidden under containers), LEDGER reaches 86.7% vs. 56.2% for MemER. On Reference (identifying which object to act on based on past events), 60.7% vs. 40.3% (Table 1). The tracker-only ablation already reaches 87.2% on Permanence, showing that geometric tracking of object positions — not learned memory — is what resolves spatial queries. The ledger adds the event history layer that lifts Reference from 32.7% (tracker alone) to 60.7% (Section 4.2).

Step-Boundary Replanning via Proprioception

The planner is invoked only at step boundaries — when the gripper completes a grasp, release, or press — detected entirely from proprioception rather than visual trackers or VLM evaluators. This avoids mid-motion intervention instability while still allowing the system to react to events that occur during execution. Table 2 shows that re-deciding after each step lifts UnmaskSwap from 68% to 86%, and carrying the remaining plan across decisions restores Repick from 28% to 54% (preventing the planner from miscounting completed repetitions).


2. Contrarian Perspectives

Embedding All Memory Inside the Policy Is the Wrong Architecture

Most VLA research (HAMLET, Gated Memory Policy, EchoVLA, π0.6-MEM, MemER) tries to compress observation history into the policy's internal representations. The paper argues this is fundamentally limited: "none keeps, per object, where it rested and what came to cover it" (Section 2, "Memory inside the policy"). The evidence is stark — the best in-policy memory reaches only 44.5% on RoboMME while oracle-grounded subgoals reach 84.1% with the same backbone and training data (Section 1). The gap is not in low-level execution but in resolving what to act on, which external memory addresses directly.

A Learned Task Router Is Unnecessary and Fragile

DIRECT trains a router on per-task success labels to select among memory mechanisms. LEDGER argues the routing decision should emerge from the instruction and the record at runtime: "Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router" (Abstract). The planner correctly defers on procedural tasks and grounds on spatial/object tasks without any task identifier at deployment, losing only one task (MoveCube) where it incorrectly grounds a motion that should be reproduced from memory (Section 4.2).

Static Scene Graphs Are Insufficient for Manipulation

Existing scene graph systems (ConceptGraphs, THEA, Hydra) maintain present-tense snapshots. The paper argues this is inadequate for manipulation: "ConceptGraphs relocates a moved object by asking an LLM for plausible containers, not by consulting where it last rested" and "THEA... records neither where an object rested earlier nor what came to cover it" (Section 2, "Explicit records and harnesses"). LEDGER's per-object stationary intervals and cover timestamps capture the spatio-temporal history that static representations cannot, which is what enables the 30-point gain on object permanence.


3. Companies Identified

Physical Intelligence (π0.5, π0.7)

  • Description: Developer of the π0.5 and π0.7 vision-language-action foundation models used as the low-level policy backbone.
  • Why relevant: LEDGER is built on top of π0.5 as a single fine-tuned policy. The paper also references π0.6-MEM and π0.7 as combining short and long-term memory using "compressed video and text representations" (Section 2), positioning Physical Intelligence's own memory efforts as an alternative approach that keeps memory inside the policy.
  • Quote: "a single π0.5 policy acts either on a target grounded in the ledger or from its own frame memory" (Section 1, contributions).

Anthropic (Claude Sonnet 5)

  • Description: Provider of the Claude Sonnet 5 LLM used as the high-level planner.
  • Why relevant: The entire decision loop — reading the ledger, generating subgoals, deciding whether to ground or defer — runs on Claude Sonnet 5. The planner is called only 2.7 times per episode on average (Appendix B), making this a lightweight API dependency. This demonstrates that frontier LLMs can serve as real-time robot planners with minimal call frequency.
  • Quote: "The high-level policy is an LLM planner (Claude Sonnet 5) that operates over the ledger and one annotated frame per decision, never the video stream or the motor commands" (Section 3.4).

Meta (SAM3)

  • Description: Developer of SAM3 (Segment Anything with Concepts), used for per-object tracking.
  • Why relevant: SAM3 is the backbone of the external object memory — it tracks every object across both demonstration and execution, maintains stationary intervals, and detects occlusion events. The tracker's recovery gate distinguishes arm-caused vs. environmental occlusions, which is what enables the cover-timestamp mechanism. This makes SAM3 a critical infrastructure dependency for the approach.
  • Quote: "A SAM3 tracker then tracks each instance across both the demonstration video (when available) and the live execution rollout" (Section 3.2).

Alibaba/Qwen (Qwen3-VL-30B)

  • Description: Provider of the Qwen3-VL-30B video VLM used for demonstration transcription.
  • Why relevant: The action transcript — an ordered natural-language account of the robot's execution in the demonstration — is generated once per episode by Qwen3-VL-30B. It takes about one minute per demonstration on one A100 (Appendix B). The paper notes it "reliably captures event sequences, temporal ordering, and repetition counts, but is less accurate at disambiguating destinations among visually identical objects" (Section 3.2), which is why the ledger uses geometric co-location rather than the transcript for spatial resolution.

4. People Identified

Heng Yang

  • Lab/Institution: Harvard Computational Robotics Group
  • Why notable: Senior author and PI. Also authored Compose by Focus (referenced as [18]), a scene graph-based atomic skills system. His lab is building a coherent research program on explicit structured representations for manipulation — scene graphs, object ledgers, geometric reasoning — as an alternative to end-to-end memory inside VLAs.
  • Quote: Co-author of the core principle: "Persistent world and event state (what the world did) is kept outside the policy as an explicit spatio-temporal record" (Section 1, contributions).

Chelsea Finn

  • Lab/Institution: Stanford (referenced via multiple papers)
  • Why notable: Co-author on RoboMME [5], DIRECT [6], MemER [20], π0.5 [13], π0.7 [14], and Hi Robot [19]. She is arguably the most referenced researcher in this paper, with her work spanning both the benchmark being evaluated and multiple competing memory approaches. Her group's MemER is the strongest prior baseline that LEDGER beats.
  • Quote: MemER "falls on RoboMME to 38% and 21% on the swap tasks where oracle grounding reaches 99% and 80%" (Section 2, "Memory inside the policy").

Tanguy Dieudonné & Jack B. Jedlicki

  • Lab/Institution: Harvard / ETH Zürich (Dieudonné); Harvard (Jedlicki)
  • Why notable: Equal-contribution lead authors. Dieudonné's ETH affiliation connects this work to the broader European robotics ecosystem (Hutter's group at ETH is referenced via [11]). This is a Harvard-ETH collaboration bridging US and European Physical AI research.

5. Operating Insights

Don't Build One Memory System — Build Two and Route at Runtime

For teams deploying VLA-based robots on long-horizon tasks, the architecture choice matters more than the policy backbone. The paper shows that a single fine-tuned π0.5 with the right memory architecture (64.3%) massively outperforms the same backbone with the best single in-policy memory (45.9%). The practical recipe: keep your VLA's existing short-term history mechanism for procedural/imitation tasks, but add an external object tracker + text record for spatial/object tasks, and let an LLM planner choose which to use based on the instruction. The subgoal dropout curriculum (0.9 dropout on procedural tasks, 0.15 on spatial/object tasks during fine-tuning) is what makes a single policy work with both input modes (Section 3.3).

Use Proprioception for Step Detection, Not Vision

A subtle but operationally critical choice: step completion is detected from gripper aperture regimes (shut/hold/open) with temporal filtering, not from visual change detection or VLM evaluation. This avoids the instability of mid-motion replanning and the unreliability of VLM-based success judgment — prior work shows "frontier VLMs near chance (54.7% against 50%) at fine-grained success judgment" (Section 2, referencing Robot Critics [21]). For production systems, this means your execution loop's state machine can be driven by joint encoders, which are far more reliable than camera-based verification.

The Oracle Gap Tells You Where Your Bottleneck Is

The gap between oracle-grounded performance (84.1%) and best autonomous performance (64.3%) is 20 points — all in target resolution, not execution. If your robot can execute a task when told exactly what to pick up but fails when it has to figure that out itself, the bottleneck is perception/reasoning, not motor control. This suggests investment priority: object tracking, event logging, and spatial reasoning infrastructure will yield more ROI than further policy fine-tuning.


6. Overlooked Insights

The VLM Transcript Is Wrong, and It Doesn't Matter

Appendix D shows a concrete example where the Qwen3-VL-30B action transcript is factually incorrect — it describes the red cube being stacked on the green one, when the actual demonstration involved container swaps. Yet the episode succeeds because the ledger's geometric stationary intervals independently capture the true spatial history, and the planner resolves targets from coordinates, not from the transcript. This is a powerful architectural insight: you can use an unreliable VLM for semantic narration as long as a reliable geometric tracker provides the ground truth for spatial queries. The system degrades gracefully because the two representations are decoupled.

2D Image Coordinates Limit Real-World Transfer

The paper acknowledges that LEDGER "operates in simulation, grounding targets in 2D image coordinates" and that "extending this architecture to mobile manipulation will require 3D spatial grounding and camera pose tracking" (Section 5). This is a significant limitation for anyone considering deploying this approach on a mobile manipulator (e.g., a wheeled robot or humanoid). The co-location primitive (two positions are "the same place" when within one object width in 2D pixel space) will not survive camera motion. Teams evaluating this architecture for real deployment should budget for a 3D tracking and grounding layer, which is non-trivial and not addressed in this work.