Releases, research notes, and where to find us in person
Three specialized video generation models for Physical AI domains: autonomous driving, robotics, and complex physics. Models are based on Kandinsky 5.0 Video Lite, a 2B‑parameter model, and fine‑tuned on data from the corresponding domains. On top of the domain‑adapted models, post‑training techniques such as RL and SOAR were applied. According to internal validation, Kandinsky WM 1.0 ranks top‑4 on PAI‑Bench‑G, top‑5 on Physics‑IQ Verified, and top‑6 on RBench. Open‑source code under the MIT license.
Read →Our research team presented at the International Conference on Machine Learning at COEX, Seoul — sharing posters and talks on the lab’s latest generative modelling work.
A 48 kHz audio tokenizer built on a modified DAC architecture with continuous latent representation — Snake activations, 64 latent channels, self-attention in the bottleneck. Beats MMAudio, MovieGen and SAME-L baselines on text‑to‑audio quality. Code and weights under MIT license.
Read →The new flagship, succeeding Kandinsky 5.0, works up to twice as fast, better understands complex queries, and creates more detailed images. New Mixture‑of‑Experts architecture and Image RAG bring free professional-level editing with no generation limits.
Read →Our paper was accepted at the DeLTa Workshop on Deep Generative Models. It proposes a Bridge Matching approach for temporally-correlated data such as video, outperforming baselines on frame interpolation, image‑to‑video and video super‑resolution.
Two new models in 4×8×8 and 4×16×16 compression, replacing GroupNorm with frame-wise RMSNorm and rebalancing encoder/decoder capacity. Outperforms Wan 2.2 and HunyuanVideo 1.5 on objective metrics; code and weights open on GitHub.
Read →“A Brief History of Visual Generative AI: Kandinsky Models” — from Malevich and Kandinsky 1.0 in 2022 to the Kandinsky 5.0 family recognized as the open‑source leader on lmarena.
The team joined the 40th AAAI Conference on Artificial Intelligence at the Singapore EXPO to connect with the broader research community and follow the latest in generative AI.
The complete Kandinsky 5.0 family: Image Lite (6B) for HD generation and editing, plus Video Pro (19B) for video up to 10 seconds. Video Pro matches Google’s Veo 3 in visual quality and motion dynamics. Native English and Russian prompt support throughout.
Read →Public access via the GigaChat and Kandinsky Telegram bots — 10 free generations, then 10 a day after sign-in, plus a referral bonus. Text-to-video, image‑to‑video, camera moves, and pairing with other Kandinsky Lab generative tools.
Read →Trained on 10M+ images of Russian text rendered through real materials — carved wood, metal casting, embroidery, rose petals — Kandinsky Image now weaves Cyrillic lettering naturally into a scene instead of overlaying it.
Read →A 2B-parameter open-source video model matching or beating far larger competitors. Diffusion transformer with cross-attention, NABLA sparse attention, trained on 124M video scenes and 520M images — the model that opened the Kandinsky 5.0 era.
Read →