This paper introduces the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task to address the challenge of editing long videos with multiple instructions. Existing methods struggle with entity fragmentation, editing hallucinations, and disrupted temporal continuity when naive chunking strategies are applied. The authors propose an agentic reasoning framework to achieve consistent multi-shot video editing across extended sequences.

Read original