One of the most notable entries in the new generation of physical AI systems, NVIDIA Cosmos 3 Edge is a compact 4‑billion‑parameter world model built on the Nemotron backbone and engineered for real‑time inference on edge GPUs instead of data‑center infrastructure. The Mixture-of-Transformers design unifies perception and control, enabling native vision reasoning and world-and-action generation for embodied systems. Self-regressive and diffusion components let the model interpret sensor streams, forecast future states, and maintain coherent world representations over time. As part of the broader Cosmos 3 family, it builds on the first fully open omnimodel for physical AI to extend these capabilities to edge deployments.
Unlike data-center-scale variants, Cosmos 3 Edge targets high-throughput, memory-efficient deployment where latency and power budgets dominate. This emphasis on efficient edge inference positions the model as an infrastructure component for physical AI rather than a general-purpose chatbot.
Cosmos 3 Edge is optimized for NVIDIA Jetson modules, including the Jetson T2000 and T3000, as well as RTX GPUs and DGX systems. A single GeForce RTX-class graphics card can host the full 4-billion-parameter model, keeping data and control loops fully local. Running directly on robots, cameras, and vehicles removes dependency on cloud round trips and associated network variability.
Optimized for Jetson, RTX, and DGX, Cosmos 3 Edge brings fully local, cloud-free control loops to devices
On edge chips such as Jetson Thor, the world model shrinks the perception–prediction–action loop to millisecond-scale latency. That responsiveness is critical for safety-critical domains where delayed decisions translate into collisions, downtime, or degraded user experience. Omnimodel deployment across embedded and professional GPUs lets developers standardize on one world model while tuning only hardware profiles.
At the representation level, Cosmos 3 Edge understands and generates text, images, video, ambient audio, and actions in a unified multimodal space. The model ingests RGB camera feeds and other sensors to build explicit world states for robots and autonomous vehicles. Native vision reasoning enables interpretation of dense road and factory scenes, including object interactions and probable intent.
Self-regressive prediction and diffusion-based rollout support closed-loop simulation, future planning, and synthetic data generation for training physical policies. This consolidation lowers integration complexity while enabling more consistent behavior under distribution shifts and rare edge cases.
In industrial and service robotics, Cosmos 3 Edge provides on-device vision reasoning and policy generation for manipulation, navigation, and human–robot collaboration. Factory robots can perceive materials, lighting, and layout conditions in real time, adapting motion and control strategies locally instead of relying on remote supervisors.
Vision AI agents for smart infrastructure apply the same world model to monitor intersections, warehouses, or transit hubs through continuous video streams. The closed-loop world action modeling supports coordinated fleets of robots and vehicles that share consistent semantics while running inference independently on-device.
Benchmark results, including leading performance on VANTAGE-Bench for image understanding in its parameter class, indicate that this compact edge variant maintains competitive accuracy despite its modest size. By bringing advanced physical AI capabilities directly onto local devices, Cosmos 3 Edge narrows the gap between perception and control, making world-model-based robotics and autonomy more practical at scale.








