MoWAM: Explicit Future Motion Prediction
for Efficient World Action Models

Jiayu Wang1, Bin Zhu2, Yue Yu1, Jingjing Chen3*

1 College of Computer Science and Artificial Intelligence, Fudan University

2 Singapore Management University

3 Institute of Trustworthy Embodied AI, Fudan University

* Corresponding author

Video Overview

MoWAM predicts future robot motion together with actions, avoiding dense future video generation at inference.

Abstract

World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet generating future videos at inference introduces substantial computational overhead. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Structured end-effector motion provides a compact representation of how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing future video generation to be removed at inference. The compact motion representation also enables sampling multiple motion–action candidates and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate strong in-domain performance, improved out-of-distribution robustness, and efficient inference-time candidate exploration.

Method

MoWAM architecture showing structured gripper keypoints, the Action-Motion Transformer, and the Video Transformer. Future video generation is used only during training.
The Video Transformer learns future visual dynamics during training. At inference, its observation representation conditions the Action-Motion Transformer, which jointly predicts robot actions and future motion.

Structured motion. MoWAM represents the end effector using the gripper center, fingertip span, and relative palm offset. Their temporal changes and relative geometry describe expected future execution.

Motion and action prediction. The Action-Motion Transformer jointly predicts actions and future motion from the observation features, allowing predicted motion to inform action generation. Future frames support training but are not generated during policy execution.

Pick Banana

Pick up the banana and place it in the basket.

Franka Research 3 · Successful MoWAM executions

Trial 1
Trial 2

Stack Bowls

Pick up one bowl and stack it inside the other.

Franka Research 3 · Successful MoWAM executions

Trial 1
Trial 2

Close Drawer

Close the open drawer through controlled contact.

Franka Research 3 · Successful MoWAM executions

Trial 1
Trial 2

LIBERO

Task success across all four suites. All 13 methods reported in the paper are shown.

LIBERO
MethodRobo. P.T.SpatialObjectGoalLongAverage
OpenVLA✓84.788.479.253.776.5
OpenVLA-OFT✓97.297.8969696.75
UniVLA✓95.494.894.890.893.95
π₀✓96.898.895.885.294.1
π₀-FAST✓97.897.888.26186.2
π₀.₅✓98.898.29892.496.9
LingBot-VA✓98.599.697.298.598.5
Motus✓96.899.896.697.697.7
LaWAM✓999697.293.496.4
Fast-WAM✗96.699.294.695.696.5
IDM-WAM✗98.699.498.496.498.2
Joint-WAM✗9999.29997.898.75
MoWAM✗98.899.89997.898.9

Success rates (%). Robo. P.T. (robotic embodied pretraining): ✓ = used; ✗ = not used. The first six baselines are VLA-based; LingBot-VA through Joint-WAM are WAM-based.

LIBERO-Plus

All 11 methods reported in the paper, evaluated across seven distribution shifts with LIBERO-trained checkpoints.

No LIBERO-Plus training data. All methods are evaluated directly with checkpoints trained on LIBERO. No LIBERO-Plus data is used for training or fine-tuning, and no additional adaptation is performed before evaluation.

LIBERO-Plus
MethodRobo. P.T.ObjectsCameraInitialLightBackgroundSensorLanguageAverageLatency (ms)
OpenVLA✓532152454324432.00128.71
OpenVLA-OFT✓7365279489708772.14103.3
UniVLA✓768627790198058.86492.2
π₀✓8162409281757071.5762.9
π₀-FAST✓7362267276707865.29274.22
Motus✓8948888378578976.001621.9
LaWAM✓8437709295739878.43108.52
Fast-WAM✗8242748563727570.43260.1
IDM-WAM✗8459799068869580.14995.4
Joint-WAM✗8551919567789780.57766.3
MoWAM✗8751869668869681.43293.5

Success rates (%); latency is milliseconds per action chunk on one RTX 5090 with identical action horizons. Robo. P.T. (robotic embodied pretraining): ✓ = used; ✗ = not used.

Real-World Evaluation

Pick Banana, Stack Bowls, and Close Drawer on a Franka Research 3 robot.

Real-World Evaluation
MethodRobo. P.T.Pick BananaStack BowlsClose DrawerAverageLatency (ms)
Motus✓65658571.671621.9
Fast-WAM✗30852045260.1
MoWAM✗65958080293.5

100 demonstrations and 20 standard evaluation trials per task. Success rates (%); latency in milliseconds per action chunk. Robo. P.T. (robotic embodied pretraining): ✓ = used; ✗ = not used.

Inference-Time Scaling

Paper scaling experiment: Pick Banana success is 65, 65, 75, and 80 percent at K=1,2,4,8. Latency bars and additional-memory lines compare four WAM variants at K=1,2,4,8,16,32; IDM-WAM is out of memory at K=32.
Left: success with motion-aware candidate verification. Right: latency (bars, top axis) and additional GPU memory (lines, bottom axis). The original paper figure includes all candidate counts and all four WAM variants.
Pick Banana success by candidate count
Number of candidates K1248
Success rate (%)65657580

MoWAM scales close to Fast-WAM. Joint-WAM and IDM-WAM become substantially more expensive as K increases, and IDM-WAM runs out of memory at K = 32. MoWAM generates eight candidates in less time than either full-visual-future variant requires for one candidate.

Citation

If you find this work useful, please cite:

@article{wang2026mowam,
  title={MoWAM: Explicit Future Motion Prediction for Efficient World Action Models},
  author={Wang, Jiayu and Zhu, Bin and Yu, Yue and Chen, Jingjing},
  journal={arXiv preprint arXiv:2609.20709},
  year={2026}
}