modlens is a vision plugin designed for DeepSeek Harness that provides vision capabilities to text-only coding agents like DeepSeek and GLM. By processing images, the tool generates structured JSON evidence including OCR, layout information, and semantic data. This extension acts as a vision bridge for models that otherwise lack native multimodal processing.

Read original