Models
18 models across the world-model ecosystem.


Agora-1
Odyssey
Multi-agent world model for multiple human or AI participants in one real-time simulation.


AlayaWorld
Alaya Lab
Interactive autoregressive world model with real-time camera control, prompt switching, and long-horizon memory consistency. Open-source, sustains interactive play past the one-minute mark.


Cosmos 3
NVIDIA
Omnimodal world-model family for physical AI spanning language, image, video, audio and action.


Genie 3
Google DeepMind
General-purpose world model generating photorealistic environments from text with real-time exploration.


HY-World 1.5 (WorldPlay)
Tencent Hunyuan
Tencent Hunyuan's streaming video diffusion model for real-time, interactive world modeling with long-term geometric consistency, released as WorldPlay. Given a single image or text prompt, it generates a next-chunk (16-frame) video prediction conditioned on live user keyboard/mouse actions, dynamically reconstituting context memory from past frames to keep the world geometrically consistent over long sessions. Runs at up to 24fps in first- or third-person, across real-world and stylized scenes. Open-sourced (code, weights, and full training pipeline). Predecessor to, and released separately from, the later HY-World 2.0.
1 world on QuantWorlds


HY-World 2.0
Tencent Hunyuan
Multi-modal world model framework for world generation and world reconstruction — accepts text, single-view images, multi-view images, and video to produce editable, persistent 3D world representations (meshes / 3D Gaussian Splatting) compatible with game engines (Blender, Unity, Unreal).


LTX-2.5
Lightricks
Lightricks' 22-billion-parameter open-weights audio-video generation model, released August 11, 2026. Turns text, image, and video inputs into synchronized, high-fidelity video and audio in a single pass, including native multishot generation (connected scenes with consistent character/style across cuts). Positioned by Lightricks and press coverage as a "world model" for video, robotics and simulation, distinct from this catalogue's live-interactive world models (Marble, MIRA, PAN, WorldPlay) — see research notes for the classification nuance.

Lucy 2.5
Decart
Real-time world editing/transformation model for live immersive video experiences.
Marble 1.0 Draft
World Labs
Fast Marble variant for rapid exploration.
Marble 1.1
World Labs
World Labs world-generation model with improved quality at fixed generation cost.
1 world on QuantWorlds
Marble 1.1 Plus
World Labs
Advanced Marble model from World Labs for larger persistent 3D worlds.


Marionette
Alaya Lab
Marionette predicts an explicit, interpretable 3D world state, renders its geometry with a graphics operator that has no learnable parameters, and asks a video-diffusion model for one thing only: appearance. A world model for interactive games with articulated characters, trained on recordings from a commercial action game (drawn from the WildWorld corpus).


Matrix-Game 3.0
Skywork AI
Real-time and streaming interactive world model with long-horizon memory, for controllable game world generation.


MIRA
General Intuition
Multiplayer Interactive World Models with Representation Autoencoders — a 5-billion-parameter diffusion transformer (with a 600M-parameter video representation codec) that simulates 2v2 Rocket League matches in real time at 20 FPS, taking action streams from up to 4 agents at once. Trained on ~10,000 hours of bot-generated 2v2 matches, without an explicit physics engine or 3D representation.
1 world on QuantWorlds


Oasis 3
Decart
Interactive world model for physical AI with controllable multi-view simulation in real time.


Odyssey-2
Odyssey
General-purpose real-time world model generating interactive simulations from text or image prompts.


PAN
MBZUAI
General, interactable, long-horizon world simulation model — can be manipulated at intermediate steps and maintains consistency over long time horizons, evaluated as competitive with leading commercial world models.
1 world on QuantWorlds

Starchild-1
Odyssey
Odyssey world model exploring richer multimodal interaction beyond visual observation alone.