The paper introduces EmbodiedSkills, a unified framework for orchestrating, training, and deploying vision‑language‑action (VLA) agents on long‑horizon robotic tasks. It highlights that VLA models alone cannot guarantee task validity or success, as they must also handle perception, planning, execution, progress verification, and recovery as the physical