Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/ARC: A Reasoning Recipe for Robo…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

ARC: A Reasoning Recipe for Robot Foundation Models

DATE October 8, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS GOKUL PUTHUMANAILLAM, JENAI XUNING YANG, ET AL. (ARXIV PHYSICAL AI)ARXIV 2610.12386
// SUMMARY

1. Key Themes

Reasoning as a Substitute for Brute-Force Scaling

The prevailing approach to improving robot foundation models (RFMs) relies on collecting more robot demonstrations and training larger models. This paper demonstrates that adding the right "reasoning recipe" (ARC) to existing models yields unprecedented gains without new robot data or foundation-scale training. As stated in the abstract: "We show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs." On real robots, this approach improved $\pi_{0.5}$'s task success by 82.2 percentage points (Abstract).

Action-Grounded Causal Reasoning

The paper finds that standard chain-of-thought or subtask supervision is insufficient because models can ignore the language and predict actions directly from vision (shortcut learning). ARC introduces "action-grounded causal traces" that explicitly connect the current state to the intended effect and the resulting action. The trace consists of seven semantic fields: State, Cause, Consequence, Effect, Action, Avoid, and Completion. As noted in Section 4.1: "An action is selected because of the change it is expected to produce in the world. Reasoning should make this relation explicit."

Automated Relabeling of Existing Datasets

Creating reasoning supervision at scale is typically costly. The authors developed ARC-Trace, a pipeline that uses vision-language models to automatically relabel existing robot demonstrations. They applied this to the DROID dataset to create ARC-Trace-DROID, yielding "approximately 75K relabeled episodes and 1.2M action-aligned frames" without collecting any new robot trajectories (Section 4.2).

Efficiency Gains in Training and Inference

Beyond improving task success, the reasoning recipe makes models more efficient. Cosmos3-Nano-Policy reached baseline success with 4.3× fewer training updates (Section 5.2, KF#4). At inference, $\pi_{0.5}$+ARC matched its 10-step base policy with a single Euler step (10× improvement), and Cosmos3-Nano+ARC matched its 4-step base with 2 steps (2× improvement) (Section 5.2, KF#3).

2. Contrarian Perspectives

More Robot Data Isn't Always the Answer

The dominant strategy in the robotics industry is brute-force scaling: more data, bigger models. This paper challenges that directly, showing that refining the reasoning capabilities of existing models is a high-leverage axis for improvement. The authors state: "How can we improve the performance of existing RFMs without modifying their architectures, adding new robot data, or training at the scale of foundation models? In this work, we study this problem." (Section 1).

Standard Chain-of-Thought Fails Due to Shortcut Learning

Most companies would assume that adding any intermediate language reasoning (like plans or subtask labels) would help a VLA. The paper shows this is false. When testing ECoT-style reasoning and subtask supervision, the action expert paid little attention to the language tokens, indicating "shortcut learning, where the model can satisfy the language objective while continuing to predict actions from vision" (Section 4.1). Only by grounding the reasoning in the causal structure of the action did the model actually use the language.

External VLMs Are Better at Generating Reasoning Than the RFM Itself

A natural assumption would be that a VLA should generate its own reasoning traces internally. The authors found the opposite: "We use external generation because the built-in language components are less reliable at generating these traces than at using them for control" (Section 4.3). They deploy an external VLM (Qwen3.6-35B) to generate traces at inference time.

3. Companies Identified

NVIDIA

Description: AI computing and robotics company. Why relevant: NVIDIA developed Cosmos3-Nano-Policy (the World Action Model used in the paper), and several authors are NVIDIA researchers. The hardware setup uses NVIDIA Isaac objects. Quotes: "Cosmos3-Nano-Policy (NVIDIA, 2026), a world action model (WAM)" (Section 1).

Physical Intelligence

Description: Robotics foundation model company (developers of $\pi_0$ and $\pi_{0.5}$). Why relevant: Their $\pi_{0.5}$ VLA model is one of the two base models fine-tuned using the ARC recipe, showing massive improvements (82.2 p.p. on real robots). Quotes: "state-of-the-art VLAs such as $\pi_{0.5}$" (Abstract).

Proception

Description: Robotics company founded by Ankit Goyal. Why relevant: Ankit Goyal is a co-author and co-advisor on the paper, affiliated with both NVIDIA and Proception. Quotes: "Ankit Goyal3,5,†... 5Proception" (Author affiliations).

Alibaba (Qwen)

Description: AI and cloud computing company. Why relevant: Their Qwen3.6-35B-A3B model is used as the default external VLM to generate reasoning traces at inference time. Quotes: "We use Qwen3.6-35B-A3B as the default external VLM" (Section 5).

Google (Gemini)

Description: Technology company. Why relevant: Gemini 2.5 Pro is used in the ARC-Trace pipeline to interpret episode videos and select important frames for labeling. Quotes: "Gemini 2.5 Pro interprets the episode video and selects important frames" (Appendix A.1).

Franka Emika and Robotiq

Description: Robot arm and gripper manufacturers. Why relevant: The hardware evaluation uses a Franka Emika Panda arm with a Robotiq 2F-85 gripper (the DROID embodiment). Quotes: "seven-degree-of-freedom Franka Emika Panda arm equipped with a Robotiq 2F-85 two-finger gripper" (Appendix C.2).

4. People Identified

Gokul Puthumanaillom

Lab/Institution: University of Illinois Urbana-Champaign / NVIDIA Why notable: Lead author of the paper. Quotes: N/A

Jenai Xuning Yang

Lab/Institution: NVIDIA Why notable: Co-advisor. Developed the RoboLab-120 and RoboLab-Reasoning-50 benchmarks used for evaluation. Quotes: N/A

Ankit Goyal

Lab/Institution: NVIDIA / Proception Why notable: Co-advisor. Key figure in the Physical AI space, now at Proception. Quotes: N/A

Fabio Ramos

Lab/Institution: NVIDIA / University of Sydney Why notable: Senior author, prominent researcher in probabilistic AI and robotics. Quotes: N/A

5. Operating Insights

Decouple Reasoning Generation from Control Execution

A major operational challenge in adding reasoning to robots is latency. The paper solves this by running an external VLM asynchronously to generate reasoning traces at ~1 Hz, while the robot control loop runs at 15 Hz. "Reasoning latency therefore affects how recently the scene was interpreted, without blocking subsequent action generation" (Appendix B.3.1). This allows the system to benefit from heavy reasoning models without slowing down the control loop.

Use Counterfactual Training to Prevent Shortcut Learning

If you simply fine-tune a VLA with reasoning traces, it may ignore them. To force the model to actually condition its actions on the reasoning, the authors replace ~5% of traces with counterfactuals that contradict the demonstrated action and penalize the model if its prediction doesn't change. "To discourage this shortcut, we replace approximately 5% of traces with counterfactuals $r^-t$ contradicting the demonstrated action and exclude these examples from $L{pred}$" (Section 4.3). This is a critical implementation detail for anyone building reasoning-guided policies.

Full Fine-Tuning is Required, Not Just Action-Head Training

It might be tempting to only fine-tune the action head of a VLA to save compute. The authors found this fails because "pretrained language features do not expose causal distinctions in a form the controller can readily use. Action-head-only training leaves these features fixed and expects the controller to recover the relevant distinctions from them" (Section 5.2). Full fine-tuning allows prediction errors to reshape the language features themselves.

6. Overlooked Insights

Instruction Robustness via Reasoning

A buried but highly practical finding is that ARC makes the robot's performance nearly invariant to how specific the user's instruction is. On RoboLab-120, base policies deteriorate significantly as instructions become vaguer. However, "policies fine-tuned using ARC maintains comparable success across vague, default, and specific instructions" (Section 5.2, KF#2). This means reasoning can effectively substitute for detailed human prompting, a major UX win for commercial robotics.

The Specific Structure of the Trace Matters Significantly

The authors ablated the components of the reasoning trace. Removing "Cause & Effect" dropped success from 45.3% to 39.25%. Providing "Action only" (without the causal context) yielded only 32.25%. Interestingly, "State + Action" (27.42%) performed worse than "Action only" (32.25%), suggesting that adding state descriptions without the causal link can actually hurt performance by adding noise (Figure 9b). The full causal account is necessary for the gains.