A line‑up of lightweight (2B parameters), high‑speed models for text‑to‑video and image‑to‑video generation of up to 10‑second clips at up to 768×512 resolution — ranked top‑5 among open‑source models on the Text‑to‑Video Arena
- SFT models — highest generation quality after fine‑tuning on curated video and image data
- No CFG distilled models — 2× faster inference by removing classifier‑free guidance
- Distilled 16‑step models (Flash) — 6× speedup via Trajectory Segmented Consistency Distillation
- Pretrain checkpoints — for further fine‑tuning and research
text‑to‑video image‑to‑video ≤10 s 768×512 24 fps
See ranking on Arena A line‑up of high‑capacity (19B parameters) models for text‑to‑video and image‑to‑video generation of up to 10‑second clips at high resolution. Delivers state‑of‑the‑art visual fidelity, cinematic motion dynamics, and precise prompt adherence — ranked #1 among open‑source models on the Text‑to‑Video Arena
- SFT models — highest generation quality after fine‑tuning on curated video and image data
- Distilled 16‑step models (Flash) — 6× speedup via Trajectory Segmented Consistency Distillation
- Pretrain checkpoints — for further fine‑tuning and research
text‑to‑video image‑to‑video ≤10 s 24 fps 1280×768
See ranking on Arena