

About
Tencent Hunyuan's streaming video diffusion model for real-time, interactive world modeling with long-term geometric consistency, released as WorldPlay. Given a single image or text prompt, it generates a next-chunk (16-frame) video prediction conditioned on live user keyboard/mouse actions, dynamically reconstituting context memory from past frames to keep the world geometrically consistent over long sessions. Runs at up to 24fps in first- or third-person, across real-world and stylized scenes. Open-sourced (code, weights, and full training pipeline). Predecessor to, and released separately from, the later HY-World 2.0.
Input types
textimage
Output types
interactive_video_world
