FLAT proposes a joint multimodal learning approach that resamples image and text into flexible‑length, aligned transmodal tokens in a 1D sequence. This eliminates the two‑stage pipeline where frozen visual embeddings bottleneck downstream generation, allowing the tokens to be directly consumed by generative decoders. The method enables linearly interpolatable embeddings that improve both retrieval and generation performance.
Read original
huggingface/daily-papers