The paper introduces Multi-agent Egocentric World Model (ME-World) to extend egocentric world modeling from single-agent to multi-agent settings with fine‑grained embodied interactions. Existing approaches predict first‑person observations using coarse actions such as locomotion or discrete commands, which neglect detailed interaction dynamics. ME‑World formulates the problem as synchronized ego‑stream generation for all agents sharing a common environment, imposing three consistency requirements: cross‑view action consistency, shared‑environment consistency, and propagation of interaction‑induced state updates. The model jointly denoises multiple ego streams within a unified token sequence, conditions each stream on the target‑view poses of every agent, and grounds the generative process in a shared environment memory that captures the world state observable by all agents. Training and evaluation are performed on both real and synthetic multi‑agent datasets, and the authors propose new shared‑world consistency metrics—environment consistency, update consistency, and identity consistency—to quantify how well the generated ego streams respect the common world and interaction effects. Experimental results show that ME‑World outperforms prior multi‑agent world‑modeling baselines in terms of these consistency metrics, action control fidelity, identity preservation, and overall video quality, demonstrating its capacity to model detailed, coordinated embodied behavior among multiple agents.

Read original