HarnessVLN is a zero-shot, training-free framework for embodied navigation that employs an Agent Harness to reconcile multimodal large language model action proposals with spatial evidence, task progress, and execution feedback, overcoming generalization limits of training-based methods and the action‑spatial alignment gaps of existing training-free approaches.

Read original