Showreel
TOP

Ranked #1 open‑source lab on the Text‑to‑Video Arena

See the leaderboard

Kandinsky Video 5.0

Text‑to‑video and image‑to‑video generation, in two sizes — a lightweight model built for speed, and a full‑scale model that ranks #1 among open‑source models on the Text‑to‑Video Arena. Both are trained for native Russian‑language and cultural‑concept understanding

Video Lite

Fast

A line‑up of lightweight (2B parameters), high‑speed models for text‑to‑video and image‑to‑video generation of up to 10‑second clips at up to 768×512 resolution — ranked top‑5 among open‑source models on the Text‑to‑Video Arena

  • SFT models — highest generation quality after fine‑tuning on curated video and image data
  • No CFG distilled models — 2× faster inference by removing classifier‑free guidance
  • Distilled 16‑step models (Flash) — 6× speedup via Trajectory Segmented Consistency Distillation
  • Pretrain checkpoints — for further fine‑tuning and research
text‑to‑video image‑to‑video ≤10 s 768×512 24 fps
See ranking on Arena

Video Pro

Smart

A line‑up of high‑capacity (19B parameters) models for text‑to‑video and image‑to‑video generation of up to 10‑second clips at high resolution. Delivers state‑of‑the‑art visual fidelity, cinematic motion dynamics, and precise prompt adherence — ranked #1 among open‑source models on the Text‑to‑Video Arena

  • SFT models — highest generation quality after fine‑tuning on curated video and image data
  • Distilled 16‑step models (Flash) — 6× speedup via Trajectory Segmented Consistency Distillation
  • Pretrain checkpoints — for further fine‑tuning and research
text‑to‑video image‑to‑video ≤10 s 24 fps 1280×768
See ranking on Arena
Generated with Kandinsky Video 5.0

Previous versions

May 2024

Kandinsky Video 1.1

17.5B parameters — generates 5.5‑second videos, trained on 4.6M text‑video pairs with captions generated by LLaVA‑1.5

Nov 2023

Kandinsky Video 1.0

17.5B parameters — generates 8‑second videos from text. Uses keyframe interpolation at 512×512 resolution