Changelog
Changelog
Changelog
- Added
Claude Sonnet 5.5
Anthropic's Claude Sonnet 5.5 joins the catalog — the direct upgrade to Claude Sonnet 5, at the same $2 / $10 per million input / output tokens, with adaptive thinking on by default, 1M-token context, 128K max output, and vision input.
Release note → - Added
Step 5 Preview
StepFun's next-generation flagship base model joins the catalog BYOK-only on AIHubMix — 1M-token context, 64K max output, text/image/video input, and adjustable reasoning at $1.08 / $3.0888 per million input / output tokens.
Release note → - Added
BGE-M3, BGE Multilingual Gemma 2, Mistral Small 3.2 — now on OVHcloud AI Endpoints
Two BAAI multilingual embedding models and Mistral's current 24B instruction model join the catalog, both on a new BYOK provider: OVHcloud AI Endpoints.
BGE-M3Multilingual embedding model (100+ languages, 8K context) returning a 1,024-dimension vector.BGE Multilingual Gemma 2Compact Gemma 2 based multilingual embedding model, 8K context, 3,584-dimension vector.Mistral Small 3.2 24B Instruct24B instruction model, 131K context, Apache-2.0.Release note → - Added
Kev 4B
Kev 4B is a small open-weight decision model — a LoRA adapter and pointer head on Qwen3.5-4B-Base. It answers typed questions about a state with calibrated probabilities instead of generated text, on the same native `/v1/decisions` contract as the rest of the family.
Release note → - Added
Decisions family + anyr decision
Fastino GLiNER2.5-Decide joins TypeSafe Jev and the anyrouter/decision preset, while the new `anyr decision` command gives the whole Decisions family a native terminal client instead of a chat-agent launcher.
GLiNER2.5-DecideFastino's 340M typed-decision classifier, BYOK + community pool on /v1/decisions.JevTypeSafe's noul / choice / score decision model on the same native contract.AnyRouter DecisionAuto-routed Decisions preset across the systemone models your keys can reach.Release note → - Added
LongCat-2.5-Preview
Meituan's LongCat-2.5-Preview — a ~1.6T-parameter MoE with ~48B active, 1M-token context, image understanding, and a thinking toggle — joins the catalog over the official LongCat API, BYOK.
Release note → - Added
Space Bunny Alpha
Space Bunny Alpha — an anonymous frontier-class stealth preview with blazing-fast inference, strong coding performance, and native multimodal input — joins the catalog as a free model. It always reasons with adjustable effort over a 1M-token context window, served $0 via OpenRouter, Nous, and OpenCode Zen platform hops.
Release note → - Added
Verified AIHubMix BYOK models
Five verified text-output models join the catalog through AIHubMix and other BYOK routes: Qwen3 235B A22B Instruct 2507, Mercury 2.5, L3.1 Euryale 70B, GPT-6 Sol, and GPT-6 Luna. Published provider prices and context windows are recorded per upstream; AnyRouter does not add platform markup to BYOK traffic.
Qwen3 235B A22B (Instruct 2507)Qwen's updated non-thinking MoE model, 262K context, AIHubMix and OpenRouter BYOK.Mercury 2.5Inception's diffusion language model, 260K context, AIHubMix and OpenRouter BYOK.L3.1 Euryale 70BSao10k's creative-roleplay text model, 131K context, AIHubMix and OpenRouter BYOK.GPT-6 SolOpenAI's high-end GPT-6 coding and agent model, 1.05M context, BYOK routes.GPT-6 LunaOpenAI's efficient GPT-6 model for focused, high-throughput tasks, 1.05M context, BYOK routes.Release note → - Disabled
Union Alpha disabled
Previously verified platform routes return 404 for stealth/union-alpha. Hidden from the catalog; configured BYOK fallbacks remain available for users who saved keys.
Release note → - Added
Claude Opus 5.5
Anthropic's Claude Opus 5.5 — the successor to Claude Opus 5 for long-running agentic coding and knowledge work — joins the catalog. It delivers Fable 5.1-level performance at 40% lower cost ($4 / $20 per million input / output tokens), with adaptive thinking always on, 1M-token context, vision input, and tool use.
Release note → - Added
Grok 4.7
xAI's Grok 4.7 joins the catalog — 500K context, adjustable reasoning, vision in, and six BYOK upstream routes.
Release note → - Disabled
Llama Nemotron Embed VL 1B v2 disabled
NVIDIA's hosted NIM API returns 404 for this catalog id. The model is downloadable self-host NIM only. The live NVIDIA embedding SKU is nvidia/nemotron-3-embed-1b.
Release note → - Added
Union Alpha now on AnyRouter
Stealth preview Union Alpha joins the catalog free — multimodal research, coding, and agentic workflows at $0.
Release note → - Added
Agnes 3.0 Flash and Nemotron Parse 2.0
Agnes AI's Agnes 3.0 Flash (512K, vision in, text out) and NVIDIA Nemotron Parse 2.0 (document images in, structured markdown out) join the catalog.
Agnes 3.0 FlashAIHubMix BYOK-style $0 billed; list $0.03 in / $0.15 out per 1MNemotron Parse 2.0Vision-in / text-out document parse; not image generationRelease note → - Added
AIHubMix BYOK hops on DeepSeek V4.1 Flash, Solar Pro 4, and GPT-6 Astra
Existing listings pick up an aihubmix-byok hop. Same catalog ids — not new SKUs. Platform aihubmix stays off (#3031).
DeepSeek V4.1 Flashaihubmix-byok list $0.155 / $0.62 per 1M; billed $0Solar Pro 4aihubmix-byok list $0.3 / $1.2 per 1M; billed $0GPT-6 Astraaihubmix-byok list $10 / $50 per 1M; billed $0Release note → - Added
Context-window floors on anyrouter/auto, free, byok, and the other virtual ids
Append [1m] or [500k] to any first-party virtual id to keep only members with at least that context window. Not a new listing — the same preset, with a token floor.
AnyRouter Autoanyrouter/auto[1m] / [500k]AnyRouter Freeanyrouter/free[1m] / [500k]AnyRouter BYOKanyrouter/byok[1m] / [500k]AnyRouter CodingAnyRouter AgentAnyRouter HermesAnyRouter CoworkAnyRouter LatestRelease note → - Upstream
Poolside platform route disabled
The Poolside platform key returns HTTP 429 usage limit exceeded. Listings stay: poolside/laguna-s-2.1, poolside/laguna-xs-2.1, poolside/laguna-m.1, poolside/laguna-xs.2 — not new SKUs. Platform poolside is off; hops parked. poolside-byok stays live.
Poolside (platform)HTTP 429 usage limit exceeded (#3297, #3159). Hops parked (#3298). poolside-byok stays on.Release note → - Added
Ling-3.0-flash-VL, Sante, Fin, and OpenRouter :free hops
inclusionAI Ling-3.0-flash-VL, Sante, and Fin join the catalog, plus Nex-N2.5 Mini/Pro. OpenRouter :free is the wire and an alias — Nemotron 3 Ultra and Inkling Small pick up hue hops so we can burn remaining free quota.
Ling-3.0-flash-VLVision SKU; OpenRouter :free is an alias, not the listing idLing-3.0-flash-SanteHealth/medicine SKU; Command Code and OpenRouter :free are aliases, not the listing idLing-3.0-flash-FinFinance SKU; OpenRouter :free is an alias, not the listing idNex-N2.5-MiniOpenRouter :free-only SKU; distinct from retired nex-n2-proNex-N2.5-ProOpenRouter :free-only SKU; distinct from retired nex-n2-proRelease note → - Added
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash joins the catalog as deepseek/deepseek-v4.1-flash, with official DeepSeek API and OpenRouter BYOK hops. Peak list $0.30 / $1.20 per 1M tokens.
Release note → - Disabled
2 models disabled
Caugiay backend disabled 2026-09-07 (#3146/#2871) — HTTP 401; this model's only platform upstream was caugiay. Plus 1 more.
Gemma 3 27B ITCaugiay backend disabled 2026-09-07 (#3146/#2871) — HTTP 401; this model's only platform upstream was caugiay.Qwen3.6-27BCaugiay backend disabled 2026-09-07 (#3146/#2871) — HTTP 401; this model's only platform upstream was caugiay.Release note → - Upstream
Caugiay and SambaNova platform routes disabled
The Caugiay platform key returns HTTP 401 Access denied, and the SambaNova platform account requires a payment method (HTTP 402). BYOK SambaNova stays live; Caugiay has no BYOK sibling. NVIDIA NIM hops that 404 for models NVIDIA does not serve were dropped per listing, not the whole nvidia backend.
Caugiay (platform)FPT Cloud 401 Invalid API Key (#3146, #2871).SambaNova (platform)HTTP 402 PAYMENT_METHOD_REQUIRED on google/gemma-4-31b (#3172). sambanova-byok stays on.Release note → - Added
Claude Fable 5.1 and GPT-6 Astra on Experiential
Experiential Cloud promotional hops for Claude Fable 5.1 and a new GPT-6 Astra listing (BYOK live; hosted platform hop pending #2842 recert).
GPT-6 AstraNew listing — Experiential Cloud promo slug gpt-6-astraClaude Fable 5.1Adds Experiential Cloud promo hop claude-fable-5.1Release note → - Disabled
Pareto disabled
Cloudflare AI Gateway Unified Billing is exhausted (#3121), so the platform cloudflare backend is disabled. Keep the verified unbiased/pareto model as a disabled catalog tombstone until the backend is funded and re-enabled.
Release note → - Added
Muse Spark 1.3
Meta's Muse Spark 1.3 joins the catalog over AIHubMix BYOK — a multimodal reasoning model for agentic tasks with a 1M-token context window.
Release note → - Added
Gemini 3.8 Flash and Mercury 2.5 Preview
Google Gemini 3.8 Flash and Inception Labs Mercury 2.5 Preview join the catalog. Claude Fable 5.1 and Qwen3.8 Max pick up AIHubMix hops on their existing ids.
Release note → - Disabled
Ox Alpha disabled
This model has been removed. Requests that still use stealth/ox-alpha or stealth/ox-alpha[1m] are routed to anyrouter/auto.
Release note → - Added
Claude Fable 5.1
Anthropic Claude Fable 5.1 joins the catalog as one listing, routed through Anthropic BYOK and Cloudflare Workers AI, with OpenRouter BYOK as a fallback.
Release note → - Added
Hy4 Preview on AIHubMix BYOK
Tencent Hunyuan Hy4 Preview joins the catalog as tencent/hy4-preview, routed through AIHubMix BYOK at the aggregator's published list rates.
Release note → - Added
Qwen3.8 Flash on AnyRouter
Alibaba Qwen3.8 Flash via AIHubMix BYOK. List $0.1126 / $0.38 per 1M.
Release note → - Disabled
2 models disabled
Every remaining platform host returns 404 for this catalog id. NVIDIA's current Nano omni/VL successor is nvidia/nemotron-3-nano-omni-30b-a3b-reasoning. Plus 1 more.
Nemotron Nano 12B V2 VLEvery remaining platform host returns 404 for this catalog id. NVIDIA's current Nano omni/VL successor is nvidia/nemotron-3-nano-omni-30b-a3b-reasoning.Nemotron Nano 9B V2Every remaining platform host returns 404 for this catalog id. NVIDIA's current Nano line is listed as nvidia/nemotron-3-nano-30b-a3b.Release note → - Disabled
Hy3 disabled
Hosted probes for this catalog id fail with a bad request. The model is unlisted rather than advertised as a live hosted chat model.
Release note → - Disabled
5 models disabled
The remaining platform host is out of credit, and the other platform route is already disabled. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model. Plus 4 more.
DeepSeek V3.2The remaining platform host is out of credit, and the other platform route is already disabled. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3 30B A3BThe platform host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3.8-Flash-NextThe only upstream for this catalog id is currently skipped by routing (community host unstable). The id is unlisted rather than advertised as a live hosted chat model.QwQ 32BThe platform host no longer serves this catalog id. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Kimi K2.6Wafer platform balance is empty (smoke 2026-08-27). Use moonshotai/kimi-k2.6.Release note → - Added
Catalog sweep — GLM-5.3, vision, Agnes, Sol Disc
Eight models from the post-closeout sweep: Z.AI GLM-5.3 (coding preview), DeepSeek vision + 0731-fast, Microsoft MAI-Thinking-1, OpenAI GPT-5.6 Sol Disc, and Agnes 2.5 Flash/Pro/Alpha via AIHubMix.
GLM-5.3DeepSeek-V4-Flash-Vision-ExpMAI-Thinking-1DeepSeek-V4-Flash-0731-FastGPT-5.6 Solwas openai/gpt-5.6-sol-discAgnes 2.5 FlashAgnes 2.5 ProAgnes 2.5 Pro AlphaRelease note → - Added
GLM-5.3-Flash — Ox Alpha, named
The Ox Alpha preview is GLM-5.3-Flash: native multimodal, 1M context, MIT-licensed 320B-A18B. Same AnyRouter routes, catalog id z-ai/glm-5.3-flash.
Release note → - Added
Qwen3.8-Flash-Next from Empero
Free community endpoint for Qwen3.8-Flash-Next — open by default, from Empero research lab. Catalog id qwen/qwen3.8-flash-next.
Release note → - Added
AIHubMix free tier — 15 more $0 routes
AIHubMix joins the platform free pool: MiniMax M2.7, Kimi K3, Gemini 3.x Flash, Nemotron 3, GPT-OSS 20B, and more, all routed at $0.
MiniMax M2.7via aihubmixKimi K3coding-kimi-k3 free idGemini 3.7 Flashfree id now platform-routedGemini 3.6 Flashfree id now platform-routedGemini 3.5 Flash Litefree id now platform-routedNemotron 3 Ultra 550Bfree id now platform-routedNemotron 3 Super 120Bfree id now platform-routedNemotron 3.5 Lightning 30Bfree id now platform-routedDots3-Note Previewfree id now platform-routedLing 3.0 Flashfree id now platform-routedLFM 2.5 2.6Bback — live $0 route againNorth Mini Codefree id now platform-routedLaguna XS 2.1free id now platform-routedOx Alphabare preview id on AIHubMixRelease note → - Disabled
2 models disabled
The only platform host (hoian) is balance-quarantined and no longer serves this catalog id reliably. Unlisted until a live hosted path returns. Plus 1 more.
Leanstral 1.5The only platform host (hoian) is balance-quarantined and no longer serves this catalog id reliably. Unlisted until a live hosted path returns.Qwen3-Coder-30B-A3BThe platform host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Release note → - Added
Ox Alpha
Ox Alpha joins the catalog free — a stealth reasoning model for coding, agentic work, and long-horizon engineering.
Release note → - Added
Qwen 27B dense SKUs and Llama 3.3 70B Instruct
Qwen3.6-27B, Qwen3.8-27B, Gemma 3 27B IT, and Meta Llama 3.3 70B Instruct join the catalog as paid chat models.
Release note → - Added
Qwen3.8-27B and DeepSeek V4 Flash, 50% off
Half off through 1 Sep. Same ids, same API — just a friendlier bill.
Release note → - Disabled
12 models disabled
The platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model. Plus 11 more.
Gemma SEA-LION v4 27B ITThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3 MaxThe partner-hosted platform path is no longer billable, and remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen 3.7 MaxCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.Qwen 3.7 PlusCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.Qwen 3.8 MaxCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.DeepSeek R1 Distill Qwen 32BThe platform Workers AI host no longer serves this catalog id. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.DiffusionGemma 26B A4BChat completions return empty content for this discrete-diffusion model under normal chat token budgets. It is unlisted rather than advertised as a live chat model.Gemma 3 12B ITThe platform Workers AI host no longer serves this catalog id. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Llama 3.2 1B InstructThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Llama 3.2 3B InstructThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3-32BThe platform host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3 Embedding 0.6BThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted embedding model.Release note → - Disabled
Mistral 7B Instruct v0.1 disabled
Cloudflare Workers AI retired this SKU on 2026-05-30; live smoke 2026-08-14/18 returns model_unavailable (no_route). It was this model's only upstream, so it was disabled rather than the route pruned.
Release note → - Upstream
GitHub Models platform and BYOK routes disabled
GitHub Models was fully retired on 2026-07-30. The inference API and BYOK endpoints return HTTP 410, so the github and github-byok backends are off.
GitHub Models (platform)Retired 2026-07-30 — HTTP 410 on every try (#2237).GitHub Models (BYOK)Retired with the platform API on 2026-07-30 (#2237).Release note → - Added
DeepSeek V4 Pro
DeepSeek V4 Pro (0813) joins the catalog, hosted on Cloudflare Workers AI with a 1M-token context window.
Release note → - Disabled
4 models disabled
Cloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0. Plus 3 more.
Qwen 3.5 397B A17BCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.DeepSeek V4 Flash 0731 (NIM)NVIDIA NIM smoke timeout 2026-08-17; 99% of prod errors are $0-balance with no serving BYOK. Use deepseek/deepseek-v4-flash.Qwen3.5 397B A17BWafer 404 on every prod generation and smoke 502 (2026-08-17). No other non-BYOK host.MiniMax M3Wafer smoke failed repeatedly (2026-08-14). Use minimax/m3.Release note → - Added
Gemini 3.7 Flash and latest frontier chat SKUs
Google Gemini 3.7 Flash joins the catalog, plus Qwen3.8 2.4T A95B, Muse Spark 1.2, Inkling Small, Solar Pro 4, Seed 2.1 Turbo, Seed 2.0 Code, Claude Opus 5 Fast, and Dots3-Note Preview.
Gemini 3.7 FlashQwen3.8 2.4T A95BMuse Spark 1.2Inkling SmallSolar Pro 4Seed 2.1 TurboSeed 2.0 CodeClaude Opus 5 FastDots3-Note PreviewRelease note → - Disabled
5 models disabled
CF Workers AI LoRA host 404 / model_unavailable on every smoke probe (2026-08-14). Canonical Gemma instruct SKUs remain (gemma-3 / gemma-4). Sibling gemma-2b-it-lora already disabled. Plus 4 more.
Gemma 7B IT LoRACF Workers AI LoRA host 404 / model_unavailable on every smoke probe (2026-08-14). Canonical Gemma instruct SKUs remain (gemma-3 / gemma-4). Sibling gemma-2b-it-lora already disabled.Llama 3.1 405B InstructGitHub Models retired 2026-07-30 (HTTP 410; smoke 2026-08-14). Only platform backend was github. Do not re-enable unless a non-GitHub free/platform route appears.Llama 2 7B Chat HF LoRAWorkers AI LoRA host — sibling Gemma LoRA SKUs already 404 on this account. Disabled so smoke does not pick it as the next Cloudflare cheapest model.Mistral 7B Instruct v0.2 LoRAWorkers AI LoRA host 404 / model_unavailable (smoke 2026-08-14). Same class as gemma-7b-it-lora and llama-2-7b-chat-hf-lora.Big Pickleopencode-zen 502 on every 2026-08-14 smoke probe. Catalog-only connectivity toy; disable so the backend is not counted as a failing smoke target.Release note → - Disabled
7 models disabled
SiliconFlow no longer lists this id; its catalogue carries gemma-4-12B-it, gemma-4-26B-A4B-it and gemma-4-31B-it instead. Plus 6 more.
Gemma-4-27B-itSiliconFlow no longer lists this id; its catalogue carries gemma-4-12B-it, gemma-4-26B-A4B-it and gemma-4-31B-it instead.Nemotron 3.5 NanoIts only upstream went offline and no other provider serves this model — the NVIDIA Nemotron 3.5 line is now published as Lightning and Content Safety only. Nemotron 3.5 Lightning is the closest replacement.GPT-5 ChatCloudflare Workers AI retired this SKU; the binding returns HTTP 410 (Gone) on every attempt. In the same 2026-07-25 smoke run, the sibling partner-hosted model openai/gpt-5.6-luna reached the same OpenAI-on-Workers-AI rail and got HTTP 400 (request rejected, not gone), showing the retirement is specific to this SKU rather than a rail-wide outage or billing issue. It was this model's only upstream, so it was disabled rather than the route pruned.GPT-5.1 ChatCloudflare Workers AI retired this SKU; the binding returns HTTP 410 (Gone) on every attempt. In the same 2026-07-25 smoke run, the sibling partner-hosted model openai/gpt-5.6-luna reached the same OpenAI-on-Workers-AI rail and got HTTP 400 (request rejected, not gone), showing the retirement is specific to this SKU rather than a rail-wide outage or billing issue. It was this model's only upstream, so it was disabled rather than the route pruned.Laguna M.1Poolside's live listing now carries only Laguna XS 2.1 and Laguna S 2.1 — this generation has been withdrawn, and every route to it (including the aggregator mirrors) now 404s. Laguna S 2.1 is the direct replacement.Laguna XS.2Poolside's live listing (inference.poolside.ai/v1/models) now carries only laguna-xs-2.1 and laguna-s-2.1 — this generation is gone, so every upstream here 404s.GLM-4.6V-FlashNo upstream currently serves this model — its only provider went offline, and Z-AI's own API publishes the text GLM line only, not the vision variants. GLM-4.6V is the closest available alternative.Release note → - Added
Nemotron Lightning, Muse Glimmer, Seedance 2.5, DeepSeek NIM Flash
NVIDIA NIM additions (Nemotron 3.5 Lightning 30B, Meta Muse Glimmer 30B, DeepSeek-V4-Flash 0731), ByteDance Seedance 2.5 on Workers AI, and free Requesty route for Ling-3.0-tiny.
Nemotron 3.5 Lightning 30B A3BSparse MoE chat model on NVIDIA NIM.Muse Glimmer 30BMeta Muse Glimmer via NVIDIA NIM.DeepSeek V4 Flash 0731Dated NIM packaging of DeepSeek-V4-Flash.Seedance 2.5Audio-video generation on Cloudflare Workers AI.Release note → - Added
Free-route expansion + Liquid LFM 2.5
Expanded free platform routes across NVIDIA Nemotron free models, GPT-OSS-20B free, Ling-3.0-tiny on Novita, and new Liquid LFM 2.5 2.6B free. CommandCode free/promo routes for Laguna S 2.1 and MiMo V2.5 Pro.
LFM 2.5 2.6BFree tier via platform free pool + BYOK free SKU.Nemotron 3.5 Lightning 30B A3BAdded free aggregator + free-pool routes.GPT-OSS 20BFree aggregator SKU added.Ling-3.0-tinyNovita platform/BYOK free routes.Release note → - Added
Grok 4.6
xAI Grok 4.6 joins the catalog via xAI BYOK and OpenRouter BYOK. 500K context, $2/$6 list on xAI.
Release note → - Disabled
3 models disabled
The catalogue is text/embedding-output only; video models stay on disk for future re-enable (same posture as seedance-2.0-mini). Unpriced CF video would also fail metered-catalog-billing. Plus 2 more.
Seedance 2.5The catalogue is text/embedding-output only; video models stay on disk for future re-enable (same posture as seedance-2.0-mini). Unpriced CF video would also fail metered-catalog-billing.EmbeddingGemma 300MSmoke reported model_unavailable (no upstream for platform keys). See #1799/#1806.Gemma 2B IT LoRAPlatform smoke and e2e hit model_unavailable — the Workers AI binding returns no usable upstream for this LoRA host model under the current account neuron/catalog posture. Canonical Gemma instruct SKUs remain available (gemma-3 / gemma-4 family). See #1799/#1806.Release note → - Added
EmbeddingGemma 300M and Ling-3.0-tiny
Google's EmbeddingGemma 300M open embedding model and inclusionAI's Ling-3.0-tiny free MoE chat model join the catalog.
EmbeddingGemma 300M300M multilingual text embeddings on Workers AI.Ling-3.0-tiny7.9B MoE (1.3B active) free instruct model via OpenRouter.Release note → - Added
Sakana Namazu
Sakana AI's Japanese-specialized LLM (built on Kimi K2.6) joins the catalog via BYOK.
Release note → - Added
Nemotron 3.5 Nano
NVIDIA's Nemotron 3.5 Nano joins the catalog via Blackbox AI BYOK.
Release note → - Added
Nemotron Nano, GLM-4.5/4.6, MiniMax M2, Qwen3.5-9B/27B
Eight new models discovered from upstream providers and added to the catalog — NVIDIA's free Nemotron Nano 12B V2 VL and 9B V2 via OpenRouter, Z-AI's GLM-4.5 and GLM-4.6 (with vision variant), MiniMax M2, and Alibaba's Qwen3.5-9B and Qwen3.5-27B.
Nemotron Nano 12B V2 VLFree multimodal reasoning model (text+image+video), 128K context.Nemotron Nano 9B V2Free reasoning model, 128K context, knowledge cutoff 2025-03-31.GLM-4.5355B MoE agent model, 128K context, 96K output.GLM-4.6200K context, stronger coding and reasoning than GLM-4.5.GLM-4.6VVision variant of GLM-4.6, 200K context.MiniMax M2Agentic coding and office productivity model.Qwen3.5-9B262K context, vision support, via DeepInfra and BYOK.Qwen3.5-27BDense frontier model, 262K context, reasoning and coding.Release note → - Added
Qwen3.7 Flash
Alibaba's Qwen3.7 Flash — a vision-language reasoning model for multimodal agents, visual coding, search, and computer interaction — joins the catalog via OpenRouter BYOK.
Release note → - Disabled
16 models disabled
The catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend). Plus 15 more.
FLUX.2 DevThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).FLUX.2 Klein 9BThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Seedream 5 ProThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Eleven Flash v2.5The catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Eleven Multilingual v2The catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 LargeThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 Medium TurboThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 MediumThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Llama 2 7B Chat FP16Cloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 2 7B Chat INT8Cloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3 8B Instruct AWQCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3 8B InstructCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3.1 8B Instruct AWQCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3.1 8B Instruct FastThe model is gone from Cloudflare's live account catalogue and the binding returns HTTP 404 (model does not exist). It was this model's only upstream, so it was disabled rather than the route pruned.Grok STTThe catalogue was narrowed to text- and embedding-output models. This model's output IS text, but its input is audio — it is a speech-to-text SKU, named explicitly in the same decision, so it goes out of service with the image / video / TTS models rather than being kept on the technicality of its output modality.Grok TTSThe catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Release note → - Added
Claude Opus 5
Anthropic's Claude Opus 5 — the flagship model for demanding reasoning, coding, and long-horizon agentic work — joins the catalog, BYOK-only like its Claude siblings.
Release note → - Added
Ling-3.0-flash
inclusionAI's Ling-3.0-flash — a 124B-parameter Mixture-of-Experts model (~5.1B active) tuned for token-efficient, production-scale agentic inference — joins the catalog, served free with a user-owned API key (BYOK).
Release note → - Upstream
Google unified-billing and Blackbox platform routes disabled
The Cloudflare AI Gateway unified-billing balance and the Blackbox platform budget both ran out. BYOK siblings stay live; platform traffic fails over or stops.
Google AI (platform)Unified billing 402 / insufficient balance (#1321). No recharge.Blackbox AI (platform)Provider budget cap (~$3). No traffic until the account is topped up (#1326).Release note → - Added
Gemini 3.6 Flash & Gemini 3.5 Flash Lite
Google's Gemini 3.6 Flash and Gemini 3.5 Flash Lite join the catalog as partner-hosted SKUs on Cloudflare Workers AI, with OpenRouter BYOK fallback.
Release note → - Added
Laguna S 2.1
Poolside's latest coding agent model Laguna S 2.1 (118B total, 8B active parameters, 1M context, 70.2% Terminal-Bench 2.1, 40.4% DeepSWE) joins the catalog — open-weight under OpenMDW-1.1.
Release note → - Added
Inkling + Nemotron 3 Embed 1B
Thinking Machines Lab's first open-weights foundation model Inkling (975B MoE, 41B active, 1M context, multimodal input with switchable reasoning) and NVIDIA's Nemotron-3-Embed-1B multilingual embedding model (2048 dims, 34 languages) join the catalog via NVIDIA NIM.
InklingNemotron 3 Embed 1BMultilingual text embeddings (2048 dimensions, 32k context) for search and RAG.Release note → - Added
Kimi K3
Moonshot AI's Kimi K3 is now available on AnyRouter — an open-source model with a 1M-token context window and native visual understanding (text, image, and video). Built for long-horizon software engineering and deep reasoning where code meets visual and spatial thinking, with thinking mode always on.
Kimi K3Muse Spark 1.1Meta's multimodal agentic reasoning model (text/image/video/audio/PDF in, 1M context) via OpenRouter BYOK. US-only.Release note → - Upstream
OpenAI and DeepInfra platform routes disabled
The platform OpenAI key has no quota (429 insufficient_quota) and DeepInfra has no payment method (402). BYOK users are unaffected.
OpenAI (platform)No quota — 5,179 attempts / 0 ok over 30d (#1283).DeepInfra (platform)No payment method — 1,240 attempts, 76.9% fail (#1284).Release note → - Upstream
Named OpenRouter platform backend disabled
The public `openrouter` backend is off. Platform traffic uses the masked hue route; user keys stay on openrouter-byok.
OpenRouter (named platform)Disabled #1152. Hue is the platform path; openrouter-byok is BYOK.Release note → - Added
Grok 4.5 and Laguna XS 2.1
Two frontier additions: xAI's Grok 4.5 joins the catalog and Poolside ships a refreshed Laguna XS 2.1 coding model.
Release note → - Added
Free-tier open models via SiliconFlow
A batch of open-weight models now has a free-tier upstream — Qwen3 across sizes, Gemma 4 27B, and Tencent Hy3.
Release note → - Added
Claude Sonnet 5
Anthropic's Claude Sonnet 5 is available through BYOK — bring your Anthropic key and route it behind the gateway.
Release note → - Added
Codex-class OpenAI models and GLM-5
The OpenAI Codex family lands — GPT-5.3 / 5.2 Codex and 5.1 Codex Max — alongside Z.ai's GLM-5 and Qwen3.7 Max.
Release note → - Added
LongCat-2.0
The LongCat provider joins the catalog with LongCat-2.0, available via BYOK.
Release note → - Added
Cerebras, Venice, and Wafer providers
Three new upstreams broaden coverage — Cerebras GPT-OSS, Venice's MiniMax M2.7, and Wafer's aggregated MiniMax M3 and GLM-5.2.
Release note → - Disabled
13 models disabled
Its BYOK aggregator upstream no longer lists this model id (verify:upstreams MISSING 2026-07-16), and that was its only upstream. See #1283. Plus 12 more.
Seedance 2.0 MiniMAI-Image-2.5Its BYOK aggregator upstream no longer lists this model id (verify:upstreams MISSING 2026-07-16), and that was its only upstream. See #1283.Nex-N2-ProCosmos 3 Nano ReasonerNVIDIA has not exposed this on the hosted API yet (integrate.api.nvidia.com 400s); it ships as self-host NIM + HF weights only.Ising Calibration 1.5 31BInkling 256KStaged, not served (#1512): Cloudflare publishes no per-token price for this SKU, and vendor guidance says it's intended for low-traffic testing/internal use, not high-throughput production.Kimi K2.7 CodeGLM-5.2Qwen3.6 35B A3BQwen3.7 MaxGLM-5.2Grok Imagine Video 1.5 (Preview)Grok Voice TTS 1.0Its BYOK aggregator upstream no longer lists this model id (verify:upstreams MISSING 2026-07-16), and that was its only upstream. See #1283.Release note →