Zhipu AI
Models published by Zhipu AI, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.
GLM-4.5 is Z-AI's foundational model for agent-oriented applications. A Mixture-of-Experts (MoE) model with 355B total parameters and 32B active parameters per forward pass, trained on 15 trillion tokens. Features a 128K context window (96K max output), hybrid reasoning modes (Thinking / Non-Thinking), and strong performance in tool invocation, web browsing, software engineering, and agent workflows. Supports 100+ languages.
GLM-4.6 achieves comprehensive enhancements in real-world coding, long-context processing, reasoning, searching, writing, and agentic applications. The context window has been expanded from 128K to 200K tokens with a 128K max output. Delivers stronger performance in tool use, search-based agents, and real-world coding benchmarks. Available via Z-AI API, DeepInfra, and Hugging Face.
GLM-4.6V is the vision variant of Z-AI's GLM-4.6 model. Supports text and image inputs with a 200K token context window, delivering enhanced visual understanding for document intelligence, code-screenshot interpretation, and multimodal reasoning tasks.
GLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages.
GLM-4.7 is Z-AI's multilingual text generation model with a 131,072 token context window. Supports chat, instruction-following, multi-turn tool calling, and reasoning across 100+ languages.
GLM-5.1 is Z-AI's flagship foundation model for long-horizon autonomous tasks. With a 200K context and 128K output window, it scores 58.4 on SWE-Bench Pro and supports thinking mode, function calling, and MCP integration. Overall capability aligns with Claude Opus 4.6 across reasoning, coding, and agentic benchmarks.
GLM-5.2 is Z-AI's flagship reasoning model with a 1M-token context and 128K output window. It supports controllable thinking effort, function calling, and MCP integration, and targets long-horizon autonomous coding and agentic tasks. Pair it with Claude Code over BYOK for a Claude-grade coding agent.
GLM-5.3-Flash is a native multimodal model from Z.ai (320B-A18B, MIT License, 1M-token context). It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
GLM-5.3-FlashX is Z.AI's high-speed inference variant for coding agents, real-time interactions, and long-running agentic workflows. AIHubMix lists it separately from GLM-5.3-Flash with its own rates and modality set.
GLM 5.3 Prime is the top reasoning tier of Z.ai's GLM 5.3 line, above the Flash and FlashX cuts, with a 1M-token context window. It sits alongside z-ai/glm-5.3 and z-ai/glm-5.3-flash as a separately priced official variant rather than a packaging of them, so it keeps its own listing row.
GLM-5.3 is Z.AI's reasoning model for coding and agentic workflows, built for complex software engineering, long-running agents, and vulnerability analysis. It uses the same base model as GLM-5.2 with scaled post-training for stronger coding performance, task execution, and token efficiency.
GLM-5 is Z-AI's general-purpose foundation model for reasoning, coding, and agentic workloads, with a 128K context window and 32K output, supporting thinking mode, function calling, and structured outputs.