Ling-3.0-flash-VL is inclusionAI's native multimodal instruct model, a 124B-parameter Mixture-of-Experts model with roughly 5.5B activated parameters per token. Built on Ling-3.0-flash, it adds native image and video understanding (up to 256K context) for visual reasoning, document and chart reading, and agentic GUI tasks. Distinct from the text-only inclusionai/ling-3.0-flash listing. Served free via the platform free pool, with BYOK as a fallback.
Share cards3 images


