Qwen
Models published by Qwen, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.
Qwen2.5 72B instruct — mature, broadly capable general chat model.
Alibaba's specialized coding model with 32B parameters, excelling at code generation, completion, and debugging across 92 programming languages.
Mid-size Qwen3 dense model balancing quality and latency.
Qwen3-235B-A22B-Instruct-2507 is Qwen's updated non-thinking MoE model for multilingual instruction following, coding, tool use, and long-context work. It activates 22B of 235B parameters and provides a native 262K context window.
Small Qwen3 dense model — fast, efficient, good for lightweight chat and tool use.
Qwen3 Embedding 4B is the 4B size in the Qwen3 embedding and reranking family, between the 0.6B and 8B cuts. It embeds text over a 32,768-token window at 2,560 dimensions, supports 100+ languages, is instruction-aware for task prefixes, and supports Matryoshka truncation from 32 to 2,560 dimensions. It ranks below the 8B cut on MTEB multilingual but well above the 0.6B cut, at a much lower cost.
Qwen3 8B embedding model for text embedding tasks with a 32,768 token context window.
Qwen3.5-27B is Alibaba's largest dense Qwen3.5 model, delivering near-frontier quality across reasoning, coding, and instruction following. Features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Best suited for production deployments and complex enterprise tasks requiring top-tier performance.
Alibaba's Qwen 3.5 is a 397B-parameter mixture-of-experts model with 17B active parameters, offering strong reasoning capabilities with efficient inference.
Qwen3.5-9B is a high-performance model from Alibaba's Qwen3.5 series with a hybrid Gated Delta Networks and sparse MoE architecture. Features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Excels at reasoning, coding, instruction following, and long-context tasks.
Qwen3.5 Plus is Alibaba's mid-tier "plus" model from the Qwen3.5 generation, with a 256K context window and strong reasoning, coding, and tool use.
Qwen3.6 Plus is Alibaba's "plus" model from the Qwen3.6 generation with multimodal (text + image) support, a 1 million token context window, and strong reasoning, coding, and tool use.
Qwen3.7 Flash is Alibaba's fast multimodal model with strong spatial understanding, real-world visual perception, and efficient inference. Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception. Available through BYOK with a user-owned API key (BYOK).
Qwen3.7 Max is Alibaba's flagship "max" tier from the Qwen3.7 generation with multimodal (text + image) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.
Qwen3.7 Plus is Alibaba's advanced multimodal modeltext and image input support, a 1 million token context window, and strong reasoning capabilities. Available
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen, the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total and a 1,010,000-token context window.
Qwen3.8-27B is Alibaba's dense Qwen3.8 model for agentic coding and complex reasoning, with route-dependent context up to 1M and native thinking mode.
Alibaba Qwen native vision-language model for coding, office, long-context reasoning, and agents. 1M context.
Qwen3.8 Max is Alibaba's flagship 2.4-trillion parameter MoE model from the Qwen3.8 generation with multimodal (text, image, video) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.
Qwen3.8 Omni Flash is Alibaba Cloud Qwen's native multimodal model for text, image, audio, and video input at 1M context. It is optimized for agentic programming, knowledge work, GUI operation, and audio/video-centered workflows. AIHubMix serves it as a distinct product from Qwen3.8 Flash.