MiMo V2.5 is a native omnimodal model by Xiaomi delivering Pro-level agentic performance at roughly half the inference cost. It surpasses MiMo V2 Omni in multimodal perception across image and video understanding tasks, with a 1M context window ideal for agent frameworks where reasoning, rich perception, and cost efficiency all matter.