Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/EyeRobot 2.0: Active Gaze for Pr…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

DATE October 2, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS KUSH HARI, ANGJOO KANAZAWA, ET AL. (ARXIV PHYSICAL AI)ARXIV 2610.03710
// KEY TAKEAWAYS5 ITEMS
  1. 01Active Gaze Replaces Wrist Cameras for Precision Tasks
  2. 02The Action Representation Matters as Much as the Camera Setup
  3. 03Gaze Is Trained with RL on Real-World Data, Not Simulated
  4. 04Physical Attention Buys Robustness to Clutter and Distractors
  5. 05Data Efficiency: 10–53 Minutes of Teleop Per Task
// SUMMARY

Bottom line for operators: This paper demonstrates that a robot with a single fixed stereo camera — augmented with physically swiveling "eyes" that actively fixate on task-relevant objects — can match or beat the industry-standard ego + wrist camera setup for fine-grained bimanual manipulation. If validated at scale, this challenges a core hardware assumption embedded in nearly every bimanual manipulation stack shipping today (ALOHA-style systems, π0-style VLA deployments): that you need cameras on the wrists for precision.


1. Key Themes

Active Gaze Replaces Wrist Cameras for Precision Tasks

The core contribution is "Active Visual Fixation" (AVF): two eye viewpoints physically swivel to converge on a 3D fixation point, and the system hierarchically decides where to look (a target selector) and how to servo gaze onto it (a low-level gaze policy). The payoff is quantified against the standard setup: "Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap without any wrist cameras... It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)" (Abstract). For a CTO, the headline is that precision tasks like boba straw insertion went from 8% (passive stereo) to 44% success, and marker capping from 4% to 68% — using the exact same camera stream (Section 5, Fig. 6).

The Action Representation Matters as Much as the Camera Setup

A quietly enormous finding: canonicalizing gripper proprioception and action chunks into a rotating, fixation-relative SE(3) frame — a pure software change requiring no new hardware — was the single largest ablation effect. "We test an ablation of EyeRobot 2.0 in simulation which uses SE(3) actions and proprioception in the world frame rather than a fixation-relative frame. This ablation performs on par with the gaze-free ego baseline (−21%)" (Section 5.1, Table 5). The paper's framing: fixation "compacts the size of the action distribution to learn" (Abstract), which reduces overfitting to specific trajectory positions in low-data regimes.

Gaze Is Trained with RL on Real-World Data, Not Simulated

Both the gaze servoing policy and the target selector are trained with reinforcement learning directly on real teleoperation data, avoiding the sim-to-real gap that plagues reward-shaped perception policies. The trick is counterfactual rendering: "rotating a camera about its optical center only requires warping its image, without any 3D reconstruction. This lets AVF use the same stereo camera for data collection, training, and deployment" (Section 3.1). The target selector's reward is derived from the co-trained behavior-cloning policy's action prediction accuracy — no human gaze supervision, no hand-labeled task stages.

Physical Attention Buys Robustness to Clutter and Distractors

In distractor trials with colorful objects scattered on the table, active gaze dramatically outperformed wrist-camera policies on first grasps: on the wrench task, EyeRobot 2.0 achieved 20/25 first grasps vs. 10/25 for ego+wrist and 3/25 for ego-only (Table 1). The mechanism: "The wrist camera policy often reaches for the largest or most distinct object in its view... In contrast, EyeRobot 2.0 reaches for the correct object when its gaze is correctly directed at it" (Section 5). Foveation also cuts compute: "allocating more visual tokens to the image centers, focusing computation on task-relevant features" (Abstract).

Data Efficiency: 10–53 Minutes of Teleop Per Task

All policies (baselines and EyeRobot 2.0) were trained from scratch on single-task data of 10–53 minutes per task (Section 4, Table 4), with over 1,000 physical and 1,800 simulated trials. This is a small-data regime — meaning these gains are about architecture and representation, not scale, which matters for anyone evaluating whether foundation-model approaches are the only path to precision.


2. Contrarian Perspectives

Wrist Cameras Are a Structural Liability, Not a Necessity

The paper directly attacks the consensus hardware configuration: "wrist cameras, which are used by nearly all existing bimanual manipulation systems" (Section 1). The indictment is multi-pronged — wrist cameras "tie precise visual input to the manipulator, which is limiting in tool use, whole-body manipulation, or when grasped objects occlude the camera... adds complexity, often occludes the ego camera, introduces motion blur, and prevents sleeker gripper designs" (Section 1). The evidence is strongest under occlusion: on boba and pot tasks where grasped objects block the wrist view, "ego+wrist performance collapses to near the gaze-free stereo baseline... overall a 2× improvement" for EyeRobot 2.0 (Section 5, Fig. 7). Most robotics companies would push back hard on this — wrist cameras are table stakes in current deployments — but the occlusion failure mode is real and any operator running wrist-camera systems has seen it.

More Sensors Isn't the Answer — Smarter Use of One Sensor Is

The passive stereo baseline and EyeRobot 2.0 receive "the exact same stereo stream" (Section 5), yet passive stereo scores 4% on marker capping while AVF scores 68%. The delta comes entirely from where the system points its eyes and how it represents actions. This pushes against the industry instinct to solve perception gaps by adding cameras, higher resolution, or bigger pretrained backbones. Notably, the paper also argues active gaze "aims to provide a route towards lessening this gap by leveraging human-like fixation" for learning from egocentric human video data (Section 2.1) — a strategic claim that human-video pretraining requires human-like sensing, not wrist-mounted sensing.

Sim-Only Validation Can Mislead You About What Actually Matters

A buried but important finding: "in simulation this ablation [removing foveation] performs the same as EyeRobot 2.0" (Section 5.1), and stereo ablations "have little effect in simulated tasks" (Table 5) — yet both contribute ~20% each in the real world. The authors' hypothesis: "robot execution in simulation has little to no variance, whereas real-world execution varies from trial to trial and so demands more closed-loop visual servoing" (Section 5.1). The implication for due diligence: teams validating perception innovations only in sim may be optimizing the wrong things entirely.


3. Companies Identified

Amazon (Amazon FAR)

  • Description: Amazon's Fundamental AI Research (FAR) lab; co-affiliation for multiple authors including Justin Kerr, Carmelo Sferrazza, Jitendra Malik, and Angjoo Kanazawa.
  • Why relevant: Signals Amazon's serious investment in next-gen manipulation perception for warehouse robotics. Sferrazza is also associated with MuJoCo Playground, which supplied the robot's physical constants (Section 4).
  • Quote: Author affiliations list "2Amazon FAR" (title page); "we use the public calibrated PD and mass-inertial constants in MJ-Playground" (Section 4).

The Robot Learning Company (TRLC)

  • Description: Maker of the TRLC-DK1, "An Open Source Dev Kit for AI-native Robotics" (reference [64]).
  • Why relevant: Their teleoperation hardware was modified for the custom leader arms used in data collection — evidence of their dev kit's presence in frontier manipulation research.
  • Quote: "We use a custom leader arm modified from GELLO [63] and TRLC [64] to teleop the robot" (Section 4).

Meta

  • Description: SAM3 (segmentation), DINOv3 (vision backbone), and Quest 3 (VR teleop interface) are all used in the pipeline.
  • Why relevant: The entire gaze training pipeline is scaffolded on Meta's open vision models — "we first detect candidate objects by querying SAM3 [61]... We extract depth from the calibrated stereo pair with FoundationStereo" and "a frozen DINOv3 [67] ViT-S/16 backbone" extracts image tokens (Section 3.1, Appendix B). Meta's open-model strategy is quietly foundational to academic robotics.
  • Quote: "teleoperation with GELLOs through a custom VR interface which streams visuals to the operator in a Quest 3 headset" (Section 4).

NVIDIA

  • Description: FoundationStereo, the zero-shot stereo depth model, is from NVIDIA researchers (reference [62]).
  • Why relevant: Provides the 3D localization ground truth for gaze RL rewards — a dependency for anyone reproducing this approach.
  • Quote: "We extract depth from the calibrated stereo pair with FoundationStereo [62] and set the ground truth position to be the 3D centroid of depth intersected with the object mask" (Section 3.1).

Physical Intelligence

  • Description: Referenced via π0, the vision-language-action flow model (reference [5]); also connected to the ALOHA lineage (Zhao et al., references [3], [6], [16]).
  • Why relevant: π0 and ALOHA-style systems are the archetypal wrist-camera bimanual stacks this paper benchmarks against. If active gaze wins, it affects the sensing assumptions of the dominant VLA paradigm.
  • Quote: "wrist cameras, which are used by nearly all existing bimanual manipulation systems [3, 4, 5, 6]" (Section 1); baselines were augmented "as in [5]" (Section 5).

Google DeepMind

  • Description: Referenced via RT-1 and Open X-Embodiment dataset efforts (references [13], [15]).
  • Why relevant: Part of the scaling-data school of manipulation the paper implicitly contrasts with — EyeRobot 2.0 achieves its gains in a small-data, single-task regime.
  • Quote: "several recent efforts have been directed towards scaling teleoperation datasets [13, 14, 5, 15, 16]" (Section 2.1).

I2RT (YAM manipulator)

  • Description: Hardware vendor of the bimanual 14-DoF YAM manipulator used in all experiments.
  • Why relevant: The physical platform underpinning the results; relevant to anyone evaluating low-cost bimanual hardware.
  • Quote: "EyeRobot 2.0 is implemented using a bimanual, 14-DoF I2RT YAM manipulator with parallel-jaw grippers" (Section 4).

4. People Identified

Kush Hari & Justin Kerr (co-first authors, UC Berkeley)

  • Why notable: Kerr is the lead author of the original EyeRobot ("Eye, robot: Learning to look to act with a bc-rl perception-action loop," reference [7]), and this duo is building a sustained research program on active vision for manipulation. Kerr's industry affiliation with Amazon FAR suggests this line is being watched closely by the largest robotics operator in the world.
  • Quote: "Building on this, EyeRobot 2.0 uses a similar co-training loop but tackles much more fine-grained manipulation tasks" (Section 2.2).

Jitendra Malik (UC Berkeley / Amazon FAR)

  • Why notable: One of the most influential computer vision researchers of the past three decades; his involvement signals that active/foveated vision is being taken seriously at the highest level of the field.
  • Quote: Co-author; the paper's intellectual foundation draws on his domain — "Inspired by human vision" (Abstract).

Ken Goldberg (UC Berkeley, AUTOLAB)

  • Why notable: Long-time robotics manipulation and telerobotics authority; senior advisor on this work.
  • Quote: Co-author with advising credit (†).

Angjoo Kanazawa (UC Berkeley / Amazon FAR)

  • Why notable: Expert in learning from visual data and 3D understanding; co-advisor on both EyeRobot papers.
  • Quote: Co-author with advising credit (†).

C. Karen Liu (Stanford)

  • Why notable: Leading figure in character animation and physics-based simulation; her presence bridges the graphics/simulation and robotics communities.
  • Quote: Affiliated with "3Stanford University" (title page).

Carmelo Sferrazza (Amazon FAR)

  • Why notable: Co-author on MuJoCo Playground (reference [66]), the simulation platform whose robot constants this paper uses — a key figure in Amazon's robotics simulation tooling.
  • Quote: Co-author; "we use the public calibrated PD and mass-inertial constants in MJ-Playground [66]" (Section 4).

Referenced researchers worth tracking: Chelsea Finn (π0, ALOHA, Mobile ALOHA — references [3], [5], [6]), Shuran Song (ACT, UMI, Diffusion Policy — references [3], [4], [11]), Sergey Levine, Pieter Abbeel (GELLO, MuJoCo Playground — references [63], [66]), and Russ Tedrake (UMI — reference [4]). These are the authors of the wrist-camera-dependent baselines this paper beats.


5. Operating Insights

Canonicalize Actions into a Task-Relevant, Moving Frame — It's Nearly Free

The fixation-relative SE(3) action representation delivered a 21-point swing in simulation (Table 5), and the paper notes that even gaze without it "seems to still afford overfitting to specific trajectory positions in this low-data regime" (Section 5.1). Before buying any new sensor, a manipulation team should audit its action space: representing end-effector poses relative to a task-relevant anchor (fixation point, target object, or similar) compacted the learned distribution enough to function like a 20-point data multiplier. Relatedly, the authors found EE-relative SE(3) beat all other baseline action spaces in a grid search (Fig. 12) — worth replicating in-house before committing to a policy architecture.

Wrist-Camera Systems Have a Known, Quantified Failure Mode — Plan for It

If you deploy wrist-camera bimanual systems, your success rate on any task where the grasped object occludes the wrist view will collapse toward your camera-less baseline — a 2× gap in this paper's trials (48% vs. 22%, Fig. 7). The per-stage Sankey diagrams (Figs. 9–11) show exactly where: ego+wrist fails "severely on task stages (pot on lid, boba insertion) where wrists become occluded by the grasped object" (Appendix A). For due diligence on manipulation companies, ask which SKUs/tasks involve grasped-object occlusion and how the system handles it.

Active Gaze Is Also a Compute-Efficiency Story

Foveation isn't just about acuity — it's about token allocation. The system processes "a pyramid of centered crops... allowing the model to attend to tokens from multiple resolutions" (Appendix B), concentrating compute where it matters and "reduc[ing] computation allocated to distractors" (Section 3). The trade-off to watch: end-to-end throughput is "around 40 Hz for baselines and 20 Hz for AVF policies" (Appendix C) — active gaze currently halves your control frequency. If your application needs high-rate reactive control, budget for that latency or for inference optimization.


6. Overlooked Insights

The Gaze Policy Generalizes Beyond Its Training Distribution — Trained on Stills, Tracks Moving Objects

The gaze policy was trained purely on "randomly selected stereo still-frames from our teleop demonstration set" (Section 3.1), yet "Despite being trained only on still frames, the policy can reliably track moving objects during robot execution" (Appendix A). Average 3D fixation error is just 3.5 cm (Appendix A, Table 3). This is a strong sign the gaze servo is a general-purpose, reusable capability module — potentially a component you train once and reuse across tasks, even though the current target selector is per-task.

Baseline Policies Overfit to Training Positions in Ways That Flatter Your Benchmarks

The tape task was collected on a regular 8×8 grid, and "Baseline policies are surprisingly sensitive to this positional distribution, with a large gap in performance between AVF and wrist policies on this grid-collected task. This gap is lessened by collecting randomly distributed initial states" (Appendix D). The paper also notes sim-real correlation holds for rankings but not absolute success rates, and that "sim may favor wrist cameras over physical trials" (Appendix A). Practical implication for anyone evaluating manipulation claims: test-position distribution and sim-vs-real evaluation choices can swing results by tens of points — insist on randomized test distributions and matched real-world trials. Also note the honest limitation: "EyeRobot 2.0 currently only trains task-specific gaze and gripper policies" and multi-task generalization "may necessitate different architectures for generating prompts such as a vision-language model" (Section 6) — this is a single-task result, not yet a generalist system.