Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/SplineWAM: Adaptive Action Horiz…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

DATE September 30, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS JUN GUO, HUAPING LIU, ET AL. (ARXIV PHYSICAL AI)ARXIV 2609.39873
// SUMMARY

1. Key Themes

Adaptive Action Horizons via B-Spline Representations

SplineWAM replaces the standard fixed-length action chunk used in World Action Models (WAMs) with a fixed-size window of cubic B-spline parameters (knot times and control points). This allows the policy to adaptively compress the action trajectory, spending fewer parameters on smooth free-space motion and more on complex, contact-rich manipulation. As stated in the paper: "One parameter budget then decodes into chunks of varying temporal resolution and duration, and both the executed span and the interval until the next policy call follow from the prediction itself" (Abstract). This directly addresses the inefficiency of uniform compute allocation in current WAMs.

Video-Action Temporal Alignment

The paper introduces a novel way to align video supervision with action prediction. Instead of sampling video frames on a uniform temporal grid, SplineWAM aligns the video supervision to the fitted knot times of the demonstration. This concentrates supervised frames where the action trajectory is most complex. The authors note: "Knot-grid alignment makes the video supervision share the non-uniform temporal structure of the action supervision, so that frames are supervised most densely precisely where the action trajectory is complex" (Section 3.3).

Asynchronous Execution via Jacobian-Pullback Real-Time Chunking (JP-RTC)

To deploy these models on real robots, the system must handle asynchronous execution where the robot continues moving while the next action chunk is computed. Standard Real-Time Chunking (RTC) fails when applied directly to spline parameters because the mapping from parameters to actions is nonlinear. The paper introduces JP-RTC, which "imposes chunk continuity on decoded actions and pulls it back through the decoder’s local Jacobian" (Section 3.4). This ensures the executed prefix agrees with the actions already committed, making the spline representation viable under real-world latency.

Efficiency and Success Rate Gains

SplineWAM demonstrates that you can improve task success rates while simultaneously reducing the computational load. On the LIBERO-Plus and RoboCasa benchmarks, SplineWAM improves success rates by 8.2 and 4.4 points over an action chunking WAM, while cutting policy calls per episode by 22% and 26% (Section 4.2, Table 1). On real bimanual robot tasks, it decodes 1.2 to 1.6 times as much executed motion per call compared to the baseline (Section 4.4, Table 2).

2. Contrarian Perspectives

Fixed-Length Action Chunks Are Suboptimal for WAMs

The conventional wisdom in robot learning heavily relies on action chunking (e.g., ACT), which predicts a fixed number of actions on a uniform temporal grid. This paper argues that this approach is fundamentally flawed for large WAMs because it "ties three quantities together that need not be tied: the number of predicted parameters, the number of executed control steps, and the temporal resolution of the prediction" (Section 1). By decoupling these, SplineWAM allows the robot to act longer over free space and replan more frequently during contact.

Spline Parameterization Alone Is Insufficient

One might assume that simply representing actions as splines (which are smooth and continuous) would yield the benefits seen in SplineWAM. However, the paper shows this is false. They compare against BEAST, a fixed-length B-spline tokenization method. The authors find: "A spline parameterisation alone does not account for the gains, since BEAST supplies one and does not obtain them; what separates it from the other spline rows is that the duration it represents is fixed in advance" (Section 4.2). The key innovation is allowing the represented duration to vary adaptively.

Naive Real-Time Chunking in Spline Space Is Actively Harmful

For asynchronous deployment, applying standard RTC directly to spline parameters (Naive RTC) degrades performance below the action chunking baseline. The paper states: "SplineWAM with naive RTC lands 1.3 points below the action chunking baseline under the same asynchronous protocol: the representation that was the stronger of the two synchronously, by 4.4 points, becomes the weaker one asynchronously" (Appendix C). This highlights that the nonlinear decoding of splines requires a specialized correction mechanism (JP-RTC) to be usable in real-time systems.

3. Companies Identified

Xiaomi Robotics

  • Description: A robotics company and research lab.
  • Why relevant: Several authors of the paper are affiliated with Xiaomi Robotics, indicating active research in foundational Physical AI and world action models within the company.
  • Quotes: "3Xiaomi Robotics" (Author Affiliations).

Team Wan

  • Description: Developers of the Wan video generative models.
  • Why relevant: The SplineWAM architecture is built by fine-tuning the Wan2.2-5B video generation model. This highlights the reliance of WAMs on large-scale video foundation models.
  • Quotes: "SplineWAM is obtained by fine-tuning the Wan2.2-5B video generation model (Wan et al., 2025) under the two-expert arrangement of Fast-WAM" (Appendix D).

4. People Identified

Jun Guo, Xiaoshen Han, et al.

  • Lab/Institution: Tsinghua University, Shanghai Jiao Tong University, Xiaomi Robotics, Peking University, CASIA.
  • Why notable: The core authors of the paper, demonstrating a strong collaborative effort between top Chinese academic institutions and industry to push the boundaries of world action models and action representations.
  • Quotes: "Jun Guo1,3∗, Xiaoshen Han2,3∗... 1Tsinghua University, 2Shanghai Jiao Tong University, 3Xiaomi Robotics" (Title Page).

Kevin Black, Manuel Y. Galliker, and Sergey Levine

  • Lab/Institution: (Referenced from prior work on Real-Time Chunking).
  • Why notable: Their work on RTC is the foundation that SplineWAM builds upon for asynchronous execution. The paper generalizes their method to handle nonlinear spline decoders.
  • Quotes: "Real-time chunking (RTC) (Black et al., 2025) offers a training-free solution: it poses this as an inpainting problem and adds a pseudoinverse-guidance term to the sampler’s velocity" (Section 3.4).

Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

  • Lab/Institution: (Referenced from prior work on Action Chunking with Transformers - ACT).
  • Why notable: Creators of the action chunking paradigm that this paper directly challenges and improves upon for WAMs.
  • Quotes: "Almost all current WAMs inherit the action chunk of Zhao et al. (2023): a fixed number of actions sampled on a uniform temporal grid" (Section 1).

5. Operating Insights

Cloud Inference Throughput is a Fleet Bottleneck

For companies deploying large WAMs from shared cloud infrastructure, the number of policy calls per episode directly limits fleet size. SplineWAM's ability to cut policy calls by 22-26% while improving success rates is a major operational win. The authors explicitly state: "Once such a policy is served to a fleet from shared cloud infrastructure, these redundant calls become a throughput bottleneck" (Section 1). CTOs should evaluate action representations not just on accuracy, but on inference call efficiency.

Deployment-Time Tuning of Executed Length

The amount of a predicted action window to commit to execution (the row budget H) is a deployment-time choice that should be tuned per task suite. The paper finds that the optimal executed length varies: "On RoboCasa every curve is an inverted U, so committing too little of the prediction also hurts, whereas on LIBERO-Plus all three fall monotonically and the shortest budget is always best" (Section 4.3). Engineering teams must build tuning pipelines for this parameter rather than assuming a single global default.

Task Suitability Considerations

SplineWAM is not a universal drop-in replacement. It is best suited for tasks with a mix of free-space and contact-rich motion. For tasks that are contact-rich throughout (like precision insertions), the compression benefits vanish. The authors note: "The representation is therefore best suited to tasks with a mix of free-space and contact-rich motion; for a task that is contact-rich throughout, the tolerance must be set tight enough that little compression remains" (Appendix F). Operators should analyze the temporal profile of their target tasks before adopting this approach.

6. Overlooked Insights

Shared Knot Vector Limits Compression in High-Dimensional Action Spaces

SplineWAM uses a single knot vector to describe all action dimensions simultaneously. This means the compression rate is bottlenecked by the hardest dimension to approximate (e.g., a binary gripper command or a precise wrist rotation). The authors admit: "A single knot vector describes all action dimensions at once, so the compression a given tolerance achieves is set by the hardest dimension to approximate, and the achievable compression falls as the action space grows" (Appendix F). This is visible in their real-robot results, where the bimanual mobile manipulation task (Arrange Bookshelf) decodes the shortest chunks despite involving the least fine manipulation. Companies working with high-DoF or bimanual systems should be aware of this scaling limitation.

Continuous Spline Representation of Binary Gripper Commands

The binary gripper command is represented as a continuous cubic spline, meaning an open-close transition is a steep ramp rather than a step. While the authors note this is mostly harmless due to thresholding, it consumes knot budget: "a transition in the gripper channel is precisely the kind of local feature to which the ℓ∞ criterion of Equation 2 responds, and because all dimensions share one knot vector, knots inserted to resolve it are expended on the entire row" (Appendix D). This hidden cost can reduce the temporal compression achievable in tasks with frequent gripper actuation.