Where-OPD introduces a spatially guided on‑policy self‑distillation method for multimodal large language models (MLLMs) that leverages procedurally generated synthetic scenes providing automatic object identities and spatial coordinates. In this framework, a teacher model—either a frozen copy or an exponential moving average of the student—receives textual, spatially grounded guidance that pinpoints the visual regions relevant to a given query. The teacher uses this guidance to locate and fuse evidence from multiple image regions, producing a target response that the student must replicate using only the raw image and the question. Post‑training is performed exclusively on the synthetic data, eliminating the need for human‑annotated grounding or external teacher models. Experiments show consistent gains on counting, document, and chart understanding benchmarks across several MLLM architectures. Crucially, the synthetic‑only post‑training transfers to real‑world perception, delivering an average improvement of 3.23 points across CVBench, V*, ZoomBench, BLINK, HR‑Bench, and MME‑RealWorld. These results demonstrate that privileged spatial information can induce broader perceptual capabilities through on‑policy self‑distillation, enabling substantial synthetic‑to‑real generalization beyond the original training distribution.

Read original