Research.
Open Models.

We develop open neural networks for image and video generation, publish research, and share code and models with the AI community

Models

Kandinsky Image

Image 6 Pro · New
Current: Kandinsky 6 Image Pro · Previous: Kandinsky 5.0

A line-up of text-to-image and image editing models that generate HD images from English and Russian text prompts with precise details and accurate text rendering, powered by a 6 billion parameter Mixture‑of‑Experts architecture

text-to-imageeditingEN / RU1280×768

Kandinsky Video

Lite 2B · Pro 19B
Family 5.0 · Flow matching

Lite — text‑to‑video and image‑to‑video producing up to 10 seconds of SD video with strong consistency. Pro — HD video with rich motion dynamics and precise camera control, English and Russian prompt understanding, Diffusion Transformer scaled to 19 billion parameters

text-to-videoimage-to-video≤10 s24 fps1280×768

K-VAE

Audio · Image · Video
Open tokenizers · Continuous latents

A family of open‑source variational autoencoders for audio, image, and video generation — compressing raw signals into compact continuous latents built for downstream diffusion models

audioimagevideo

Kandinsky WM

WM 1.0 · New
World models · Physical AI

Three specialized image‑to‑video models for autonomous driving, robotics, and complex physical interactions — generating physically plausible continuations of a scene

autonomous drivingroboticscomplex physics
Release4 Aug 2026

Kandinsky WM 1.0

Three video generation models fine‑tuned on Kandinsky 5.0 Video Lite (2B parameters) for autonomous driving, robotics, and complex physics, with post‑training RL and SOAR. Ranks top‑4 on PAI‑Bench‑G, top‑5 on Physics‑IQ Verified, and top‑6 on RBench. Open source under MIT license.

Read →
Release28 Apr 2026

Kandinsky Lab unveils Kandinsky 6 Image Pro — flagship model for editing and generation

The new flagship, succeeding Kandinsky 5.0, works up to twice as fast, better understands complex queries, and creates more detailed images. Built on a new Mixture‑of‑Experts architecture, it brings free professional-level editing — object removal, restyling, retouching — with no generation limits.

Read →
Research29 Jun 2026

KVAE-Audio: a full-range audio tokenizer for generative tasks

A 48 kHz audio tokenizer built on a modified DAC architecture with continuous latent representation. At 166.9M parameters and 64 latent channels, it beats larger competing models on text‑to‑audio quality across AudioSet, music, and speech benchmarks. Code and weights under MIT license.

Read →
Event6–11 Jul 2026

Kandinsky Lab at ICML 2026, Seoul

Our research team presented at the International Conference on Machine Learning at COEX, Seoul — sharing posters and talks on the lab’s latest generative modelling work.

Read →
EventApr 2026

“Time-Correlated Video Bridge Matching” at ICLR 2026, Rio de Janeiro

Our paper was accepted at the DeLTa Workshop on Deep Generative Models — a Bridge Matching approach for temporally-correlated data such as video, outperforming baselines on frame interpolation, image‑to‑video and video super-resolution.

Read →
Research16 Apr 2026

K-VAE 2.0 — an upgraded video tokenizer

Two new models in 4×8×8 and 4×16×16 compression, outperforming Wan 2.2 and HunyuanVideo 1.5 on objective metrics. Code and weights open on GitHub.

Read →
Release4 Aug 2026

Kandinsky WM 1.0

Three video generation models fine‑tuned on Kandinsky 5.0 Video Lite (2B parameters) for autonomous driving, robotics, and complex physics, with post‑training RL and SOAR. Ranks top‑4 on PAI‑Bench‑G, top‑5 on Physics‑IQ Verified, and top‑6 on RBench. Open source under MIT license.

Read →
Release28 Apr 2026

Kandinsky Lab unveils Kandinsky 6 Image Pro — flagship model for editing and generation

The new flagship, succeeding Kandinsky 5.0, works up to twice as fast, better understands complex queries, and creates more detailed images. Built on a new Mixture‑of‑Experts architecture, it brings free professional-level editing — object removal, restyling, retouching — with no generation limits.

Read →
Research29 Jun 2026

KVAE-Audio: a full-range audio tokenizer for generative tasks

A 48 kHz audio tokenizer built on a modified DAC architecture with continuous latent representation. At 166.9M parameters and 64 latent channels, it beats larger competing models on text‑to‑audio quality across AudioSet, music, and speech benchmarks. Code and weights under MIT license.

Read →
Event6–11 Jul 2026

Kandinsky Lab at ICML 2026, Seoul

Our research team presented at the International Conference on Machine Learning at COEX, Seoul — sharing posters and talks on the lab’s latest generative modelling work.

Read →
EventApr 2026

“Time-Correlated Video Bridge Matching” at ICLR 2026, Rio de Janeiro

Our paper was accepted at the DeLTa Workshop on Deep Generative Models — a Bridge Matching approach for temporally-correlated data such as video, outperforming baselines on frame interpolation, image‑to‑video and video super-resolution.

Read →
Research16 Apr 2026

K-VAE 2.0 — an upgraded video tokenizer

Two new models in 4×8×8 and 4×16×16 compression, outperforming Wan 2.2 and HunyuanVideo 1.5 on objective metrics. Code and weights open on GitHub.

Read →
Release4 Aug 2026

Kandinsky WM 1.0

Three video generation models fine‑tuned on Kandinsky 5.0 Video Lite (2B parameters) for autonomous driving, robotics, and complex physics, with post‑training RL and SOAR. Ranks top‑4 on PAI‑Bench‑G, top‑5 on Physics‑IQ Verified, and top‑6 on RBench. Open source under MIT license.

Read →
Release28 Apr 2026

Kandinsky Lab unveils Kandinsky 6 Image Pro — flagship model for editing and generation

The new flagship, succeeding Kandinsky 5.0, works up to twice as fast, better understands complex queries, and creates more detailed images. Built on a new Mixture‑of‑Experts architecture, it brings free professional-level editing — object removal, restyling, retouching — with no generation limits.

Read →
Research29 Jun 2026

KVAE-Audio: a full-range audio tokenizer for generative tasks

A 48 kHz audio tokenizer built on a modified DAC architecture with continuous latent representation. At 166.9M parameters and 64 latent channels, it beats larger competing models on text‑to‑audio quality across AudioSet, music, and speech benchmarks. Code and weights under MIT license.

Read →
Event6–11 Jul 2026

Kandinsky Lab at ICML 2026, Seoul

Our research team presented at the International Conference on Machine Learning at COEX, Seoul — sharing posters and talks on the lab’s latest generative modelling work.

Read →
EventApr 2026

“Time-Correlated Video Bridge Matching” at ICLR 2026, Rio de Janeiro

Our paper was accepted at the DeLTa Workshop on Deep Generative Models — a Bridge Matching approach for temporally-correlated data such as video, outperforming baselines on frame interpolation, image‑to‑video and video super-resolution.

Read →
Research16 Apr 2026

K-VAE 2.0 — an upgraded video tokenizer

Two new models in 4×8×8 and 4×16×16 compression, outperforming Wan 2.2 and HunyuanVideo 1.5 on objective metrics. Code and weights open on GitHub.

Read →

Mission

Our mission is to move the industry forward by creating models and approaches that open up new formats of creativity, accelerate scientific discoveries, and unlock the potential of generative AI