The authors introduce Iris-3B, a 3‑billion‑parameter text‑to‑image transformer trained in pixel space rather than through a variational autoencoder bottleneck. Pretraining proceeds via a curriculum that upsamples from 256×256 to 512×512 and finally to 1024×1024 pixels, after an ablation study at 256² determines the optimal prediction target and representation alignment for scaling. In parallel, they convert the latent‑based FLUX.2 Klein base model (4 B parameters) to pixel space. Both the scratch‑trained Iris‑3B and the converted FLUX.2 Klein are subsequently fine‑tuned on two downstream tasks: monocular depth estimation using a direct‑regression recipe and 4× image restoration/super‑resolution on the DIV2K benchmark. Experiments reveal no measurable advantage for the pixel‑space generative prior; depth‑fine‑tuned Iris‑3B performs on par with the latent FLUX.2 Klein, while the converted pixel FLUX.2 Klein lags slightly behind. On DIV2K, neither pixel model exceeds the latent FLUX.2 Klein fine‑tune, with the converted variant trailing marginally. The paper details the training recipes, observed failure modes, and confounding factors underlying these negative results. Nonetheless, Iris‑3B demonstrates that pixel‑space pretraining employing the PixelDiT PiT head can scale to 3 B parameters and achieve text‑to‑image fidelity comparable to latent models, matching Qwen‑Image on OneIG at 1024² under official evaluation. Model weights and training code are released to facilitate further pixel‑space generation research.
Read original
huggingface/daily-papers