Open Dreamer is a JAX/Flax reproduction of the Dreamer 4 world model pipeline that eliminates VAE, KL, and adversarial losses. It utilizes a shared block-causal transformer for both a causal video tokenizer and an action-conditioned dynamics model. The architecture employs a Masked Autoencoder for tokenization, using space layers for intra-frame information and causal time layers for inter-frame transitions.
Read original
reddit/r/machinelearningnews