KVAE 2.0 Image
KVAE 2.0 Image is a continuous image tokenizer designed for diffusion-based text-to-image generation. It uses 8×8 spatial compression with 32 latent channels and has approximately 0.2B parameters. In our reconstruction benchmarks, KVAE 2.0 Image outperforms the VAEs used in leading open-source text-to-image models in PSNR and SSIM. Its latent space also translates into stronger downstream generation: KVAE 2.0 Image is preferred by human side-by-side evaluation under a fixed text-to-image setup, which assesses prompt following, visual quality, and semantic consistency.