NVIDIA Cosmos 3 — Fiziksel AI için Açık Dünya Temel Modeli: Robotik, Otonom Araçlar ve Endüstriyel Görü / The Open World Foundation Model for Physical AI: Robotics, Autonomous Vehicles & Industrial Vision


:türkiye: Genel Bakış

NVIDIA, Computex 2026’da Cosmos 3’ü tanıttı — fiziksel AI için tasarlanmış, açık ağırlıklı dünya temel modeli (World Foundation Model). Cosmos 3, robotik politika geliştirme, otonom araç eğitimi ve endüstriyel görü sistemleri için ortak bir platform sunuyor. Modeller, Linux Foundation’ın OpenMDW1.1 lisansıyla yayımlandı; Hugging Face üzerinden indirilebilir, GitHub üzerinden post-training scriptlerine erişilebilir.

Fiziksel AI nedir?

Fiziksel AI, kameralar, robotlar ve otonom araçlar gibi otonom sistemlerin fiziksel dünyada karmaşık eylemleri algılamasına, anlamasına, mantık yürütmesine ve bu eylemleri gerçekleştirmesine ya da koordine etmesine olanak tanır.

Cosmos 3 Nedir?

Cosmos 3, Mixture of Transformers (MoT) mimarisi üzerine inşa edilmiş bir omni-model. Önceki Cosmos sürümlerinin aksine, algılama ve üretimi ayrı modellerle değil tek model içinde ele alıyor. Metin, görüntü, video, ortam sesi ve eylem — beş modaliteyi aynı anda anlayıp üretebiliyor.

Model önce akıl yürütüyor, sonra üretiyor; bu mimari karar, fiziksel tutarlılık açısından kıyaslama tablolarında belirgin bir avantaj sağlıyor.

Dört Temel Yetenek

  • Görsel AI Akıl Yürütme: VLM (Görüntü-Dil Modeli) olarak nesne, etkileşim ve niyet üzerinde akıl yürütüyor. Kalite denetimi, trafik izleme, lojistik ve otonom sürüş için gerçek zamanlı uyarı ve yoğun altyazı üretimi.

  • Politika Modeli Geliştirme: Dünya Eylem Modelleri (WAM) için omurga olarak kullanılıyor. Genel dünya temel modeli, platforma özgü kamera ve robot verisiyle post-training ile uzmanlaştırılabiliyor. 1X Technologies, Figure AI, Agility Robotics, Toyota Research Institute gibi önde gelen robotik şirketleri ekosisteme dahil.

  • Dünya Simülasyonu: Fizik tabanlı kontrol edilebilir dünya simülatörü olarak çalışıyor. Birden fazla yaklaşımı tahmin ediyor, kapalı döngüde sonuçları değerlendiriyor, doğru davranışa gerçek dünya riski olmadan yakınsıyor.

  • Sentetik Video Verisi Üretimi: Metin, görüntü, video, ses ve eylem girdisinden sonsuz gelecek senaryosu üretiyor. Fiziksel olarak yakalanmamış verilerle robot eğitimini mümkün kılıyor.

Açık Geliştirme Çerçevesi

  • Cosmos Curator: Büyük hacimli sensör verisini filtreleme, etiketleme ve tekilleştirme aracı
  • Cosmos Evaluator: Üretilen video çıktılarını ölçekte değerlendirme ve puanlama aracı
  • Post-training scriptleri: GitHub üzerinden her modalite için açık kaynak

:united_kingdom: Overview

NVIDIA unveiled Cosmos 3 at Computex 2026 — an open-weight World Foundation Model platform designed for physical AI. Cosmos 3 provides a unified development framework for robot policy learning, autonomous vehicle training, and industrial vision systems. Models are released under the Linux Foundation’s OpenMDW1.1 license and are available on Hugging Face, with post-training scripts on GitHub.

What Is Physical AI?

Physical AI lets autonomous systems like cameras, robots, and self-driving cars perceive, understand, reason, and perform or orchestrate complex actions in the physical world.

What Is Cosmos 3?

Cosmos 3 is an omni-model built on Mixture of Transformers (MoT) architecture. Unlike previous Cosmos releases, it handles perception and generation in a single unified model rather than separate systems, and supports five modalities natively: text, image, video, ambient sound, and action.

The model reasons first, then generates — an architectural decision that gives it a meaningful edge in physics accuracy across benchmarks.

Four Core Capabilities

  • Vision AI Reasoning: Operates as a VLM to reason over objects, interactions, and intent across complex real-world scenes. Enables real-time alerts and dense captioning for quality inspection, traffic monitoring, logistics, and autonomous driving.

  • Policy Model Development: Serves as the backbone for World Action Models (WAMs). The generalized world foundation model can be post-trained on embodiment-specific camera data, environments, and policies. Ecosystem partners include 1X Technologies, Figure AI, Agility Robotics, Toyota Research Institute, General Motors, and Uber.

  • World Simulation: Runs as a controllable, physics-grounded simulator. Predicts multiple behavioral approaches, evaluates outcomes in a closed loop, and converges on the right behavior without real-world risk.

  • Synthetic Video Data Generation: Generates infinite plausible futures from text, image, video, sound, and action inputs — enabling robot training data creation beyond what can be physically captured.

Open Development Framework

  • Cosmos Curator: Filter, annotate, and deduplicate large sensor datasets at scale
  • Cosmos Evaluator: Review and score generative video outputs at scale
  • Post-training scripts: Open-source on GitHub for each modality and module
  • NVIDIA TAO 7: Agent skills and tools for fine-tuning Cosmos 3 with coding agents and natural language prompts

:open_book: Platform: NVIDIA Cosmos
:laptop: Models (Hugging Face): Cosmos3 - a nvidia Collection
:laptop: GitHub: GitHub - NVIDIA/cosmos: NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. · GitHub
:page_facing_up: Technical Report: https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf