← Back to feed عربي
AIResearch

Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models

NVIDIA explores Vision-Language-Action (VLA) and World-Action Models (WAM) that leverage pretrained large-scale VLM backbones to enable robots to generate actions from visual and language inputs.

1 min read

NVIDIA presents a glossary for readers new to VLA/WAM terminology. A Vision-Language-Action (VLA) model is a robot policy that begins with a pretrained Vision-Language Model (VLM) backbone and adapts it to generate actions from visual observations and language instructions. Large-scale VLM pretraining is essential to this approach. Similarly, a World-Action Model (WAM) is a policy starting from a pretrained world-model or video model, fine-tuned to act in real-world scenarios. These models represent a new paradigm in robotics and AI, leveraging pretrained imagination to enable robots to perform complex tasks by understanding and acting upon multimodal inputs.

Read at original source ↗