Aphanta introduces an automated diagnostic framework for evaluating task-aligned image-edited intermediates in multimodal large language model (MLLM) pipelines. It systematically assesses three reasoning conditions: direct MLLM reasoning, reasoning with editor-generated visual intermediates, and reasoning with ground-truth intermediates. The framework enables closed-loop analysis of how faithfully image editors realize required transformations for spatial reasoning tasks.
Read
huggingface/daily-papers