# Beyond VLAs: How World Action Models Reshape Robot Manipulation

DevFeed: [Beyond VLAs: How World Action Models Reshape Robot Manipulation](<https://devfeed.tech/articles/beyond-vlas-how-world-action-models-reshape-robot-manipulation-6764.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/beyond-vlas-how-world-action-models-reshape-robot-manipulation/>)

Author: Michelle Horton

Published: 2026-08-04T16:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Robotics](<https://devfeed.tech/topics/robotics.md>), [World models](<https://devfeed.tech/topics/world-models.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [NVIDIA Research](<https://devfeed.tech/topics/nvidia-research.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [featured](<https://devfeed.tech/tags/featured.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-research](<https://devfeed.tech/tags/nvidia-research.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [robot-manipulation](<https://devfeed.tech/tags/robot-manipulation.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [thor](<https://devfeed.tech/tags/thor.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [world-model](<https://devfeed.tech/tags/world-model.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

## AI overview

The article explains how World Action Models (WAMs) use video world models as backbones for robot policies, addressing the physical-generalization limitations of vision-language-action models. It discusses post-training WAMs into specialized policies and presents NVIDIA Cosmos 3 as a foundation for building them.

## Source excerpt

A central challenge in robotics is building policies that generalize beyond the demonstrations they're trained on. A policy that succeeds in a training scene...