OmniScope introduces a training-free token compression framework for omnimodal LLMs that addresses cross-modal salience mismatch, where audio and video relevance peaks at different moments for the same query. Unlike unidirectional
huggingface/daily-papers