Show-Harness is an embodied interface that enables foundation vision-language models (VLMs) to control robots through discrete semantic action units. Embodiment-specific interpreters deterministically translate these units into local robot actions, bridging high-level intent and physical control.
Read original
huggingface/daily-papers