Image tokenizers define the visual language of unified multimodal models, yet current evaluations typically measure them in isolation or only for generation and understanding tasks. The authors construct a pure-autoregressive testbed to monitor task-specific validation losses during multimodal continual pretraining across text, image, text-to-image, and image-to-text prediction, analyzing how visual tokens interact with text tokens in a joint modeling framework.

Read original