LIFT presents a unified image‑to‑video generation framework that augments conventional camera‑trajectory conditioning with Layout‑In‑Future (LIF) control, allowing users to define the semantic content and spatial arrangement of a designated future frame, typically the last frame, in addition to camera motion. By treating the last‑frame layout as an explicit control signal, LIFT addresses the difficulty of generating coherent scenes under large viewpoint changes where newly revealed regions lack direct visual priors. Because learning from such sparse layout guidance is considerably harder than from dense per‑frame layouts, the authors introduce on‑policy self‑distillation (OPSD): a teacher model trained on full‑sequence layout annotations transfers its control capability to a student model that receives only the final‑frame layout, thereby distilling the teacher’s knowledge while maintaining the student’s policy during training. To support this setting, they curate LIFT‑Vista, a dataset comprising videos with substantial camera motion and temporally consistent layout annotations for both camera trajectories and future frames. Experimental results demonstrate that LIFT yields higher visual fidelity, improved control over the specified future layout, and better adherence to camera trajectories compared with baseline methods that rely solely on camera controls or text prompts.

Read original