Skip to content

GLM-5.3-Flash

Also accepted:zai-org/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai (320B-A18B, MIT License, 1M-token context). It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Providers
Capabilities
AIHubMix
aihubmix-byok
$0
$0
Unavailable
OpenCode Zen
opencode-zen-byok
$0
$0
Unavailable
CommandCode
commandcode-byok
$0
$0
Unavailable
OpenRouter
openrouter-byok
$0
$0
Unavailable
Venice AI
venice-byok
$0
$0
Unavailable
Cline
cline-byok
$0
$0
Unavailable
Usage analytics

Loading usage…

API & code
Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

Share cards
GLM-5.3-Flash share card
GLM-5.3-Flash
Hue upstream share card
Hue upstream
AIHubMix upstream share card
AIHubMix upstream
OpenCode Zen upstream share card
OpenCode Zen upstream
CommandCode upstream share card
CommandCode upstream
Nous Research upstream share card
Nous Research upstream
Nous Research upstream share card
Nous Research upstream
OpenRouter upstream share card
OpenRouter upstream
Venice AI upstream share card
Venice AI upstream
Cline upstream share card
Cline upstream
Credits
Use your own key

Run GLM-5.3-Flash on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Text generation
Context length1,048,576 tokens
Max output131,072 tokens
ArchitectureTransformer
Categorytext
ReleasedAug 26, 2026
Modalities
→
Capabilities