Uranus introduces a data‑driven robot simulator that leverages a joint‑trajectory‑conditioned autoregressive diffusion model to generate visual observations on‑the‑fly. At each simulation step the model consumes a future joint‑position trajectory supplied online and autoregressively produces a single latent representation that is decoded into four synchronized RGB frames, enabling open‑ended rollouts without a predefined horizon. After inference‑time optimizations the system sustains a generation rate of 24 frames per second, satisfying real‑time latency requirements for embodied‑AI experiments. The simulator exposes a unified control interface that supports synchronized multi‑view rendering across heterogeneous robot morphologies and camera setups, facilitating scalable experimentation with diverse embodiments. Quantitative and qualitative assessments were performed on both in‑distribution and out‑of‑distribution trajectories, reporting metrics such as frame‑wise LPIPS, action‑conditioned prediction error, and throughput, while highlighting current limitations including drift in long‑horizon predictions and sensitivity to unseen dynamics. The authors make the training code, pretrained weights, and evaluation scripts publicly available to encourage community adoption and further development of simulation‑based robot learning pipelines.
Read original
huggingface/daily-papers