A survey paper on World Action Models.
WAMs are moving robotics from reacting to the present toward predicting consequences before acting.
A model only counts as a WAM when its predicted future directly helps produce, score, verify, or train the action.
The trend is "dream less, act more": full video generation is often too slow and memory-heavy for real control loops.
Many newer systems skip rendered video and use latent features, geometry, affordance maps, or motion representations instead.
Photorealistic futures are not necessarily the most useful; flow, masks, tactile signals, and physically grounded latents may constrain action better.
There is no single winning architecture because every design trades predictive richness against latency, memory, action-label cost, and physical reliability.
The biggest open question is whether robots can spend heavy predictive compute only when uncertainty, contact, or irreversible error makes it necessary.
- arxiv. org/abs/2606.20781
Title: "World Action Models: A Survey"