Ling-3.0-flash-VL extends the Ling-3.0-flash language model with native image and video understanding via a ViT visual encoder. It retains the original model’s language, reasoning, and long‑context abilities, featuring 124 B total parameters with 5.5 B activated per token and a context window of up to 256 K tokens. The architecture is geared toward real‑world reasoning and agentic workflows that require multimodal input.

Read original