Ambient's submission to the EgoLongQA track of the Wearable‑AI Challenge at ECCV 2026 won the ≤2B parameter division with a test score of 0.8279. The approach uses a single 2‑billion‑parameter vision‑language model that answers multiple‑choice questions about ten‑minute egocentric videos in one greedy forward pass. The model is created by distilling the junior perception module of a tool‑using agentic pipeline into a compact student, using filtered teacher traces.
Read original
huggingface/daily-papers