Vector Quantization (VQ) underpins discrete visual tokenizers used in autoregressive and masked image generation models. Although shared-projection codebook methods have improved codebook utilization, training stability remains a critical, underexplored challenge. The authors attribute instability to the entanglement of encoder-decoder and codebook training, proposing practical guidelines for stable VQ tokenizer training.
Read original
huggingface/daily-papers