The paper introduces a model‑agnostic framework that enables frozen large language models (LLMs) and vision‑language models (VLMs) to acquire new expertise from deployment experience without modifying their internal weights. The approach augments inference with three external modules: a Skill component that directs reasoning and invokes appropriate tools; a Knowledge Memory that accumulates reliable facts validated by prior cases or trusted external sources; and a Multimodal Knowledge Base that stores visual exemplars and aligns retrieved cases with the current input image to preserve spatial detail. Updates are accepted only when a dynamic validation strategy confirms improved performance on incoming cases while preserving accuracy on previously seen data, thereby avoiding catastrophic forgetting. Evaluated across six benchmarks spanning clinical diagnosis, workflow automation, medical reasoning, and both medical and non‑medical visual reasoning, the method was tested with four base models—two open‑weight and two closed‑source LLMs/VLMs. Online deployment yielded performance gains of up to 34.2% relative to the frozen baselines on medical tasks, demonstrated generalization to unseen cases, enabled zero‑shot transfer to alternative models, and retained effectiveness in non‑medical domains. These results show that external expertise modules can continuously enrich frozen multimodal AI systems in real‑world clinical settings.
Read original
huggingface/daily-papers