Kandinsky WM 1.0

Kandinsky WM 1.0 is an open family of image‑to‑video models for Physical AI. The models take a starting frame and a text description, then generate a physically plausible continuation of the scene. The family is designed for autonomous driving, robotics, and complex physical interactions.

The family is built on Kandinsky 5.0 Video Lite, a 2‑billion‑parameter video diffusion model.

4PAI-Bench-GInternal validation
5Physics-IQ Verified
6RBench

Current models

Autonomous Vehicle

Designed for autonomous driving. It continues road scenes while accounting for vehicle direction, geometry, traffic signals, pedestrians, and complex road events. It was trained on 1.5 million video samples derived from multi‑camera recordings captured across different countries, cities, road conditions, and weather.

1.5M video samplesmulti‑camera footagediverse weather & roads

Robotics

Designed for robotics. It generates continuations of scenes with robots and manipulators, including grasping, object movement, household tasks, and long action sequences. The model was trained on 2.2 million videos spanning different robotic platforms, viewpoints, and task types.

2.2M training videosdiverse robot platforms

Complex Physics

Designed for complex physical interactions. It generates human motion, object interactions, fluids, deformation, destruction, and mechanical and industrial processes. It was trained on 1.6 million carefully filtered and manually curated videos.

human motion & fluidsdeformation & destructionindustrial processes