Skip to content

Changelog

Changelog

Changelog

  1. Added

    Claude Sonnet 5.5

    Anthropic's Claude Sonnet 5.5 joins the catalog — the direct upgrade to Claude Sonnet 5, at the same $2 / $10 per million input / output tokens, with adaptive thinking on by default, 1M-token context, 128K max output, and vision input.

    Release note →
  2. Added

    Step 5 Preview

    StepFun's next-generation flagship base model joins the catalog BYOK-only on AIHubMix — 1M-token context, 64K max output, text/image/video input, and adjustable reasoning at $1.08 / $3.0888 per million input / output tokens.

    Release note →
  3. Added

    BGE-M3, BGE Multilingual Gemma 2, Mistral Small 3.2 — now on OVHcloud AI Endpoints

    Two BAAI multilingual embedding models and Mistral's current 24B instruction model join the catalog, both on a new BYOK provider: OVHcloud AI Endpoints.

    Release note →
  4. Added

    Kev 4B

    Kev 4B is a small open-weight decision model — a LoRA adapter and pointer head on Qwen3.5-4B-Base. It answers typed questions about a state with calibrated probabilities instead of generated text, on the same native `/v1/decisions` contract as the rest of the family.

    Release note →
  5. Added

    Decisions family + anyr decision

    Fastino GLiNER2.5-Decide joins TypeSafe Jev and the anyrouter/decision preset, while the new `anyr decision` command gives the whole Decisions family a native terminal client instead of a chat-agent launcher.

    Release note →
  6. Added

    LongCat-2.5-Preview

    Meituan's LongCat-2.5-Preview — a ~1.6T-parameter MoE with ~48B active, 1M-token context, image understanding, and a thinking toggle — joins the catalog over the official LongCat API, BYOK.

    Release note →
  7. Added

    Space Bunny Alpha

    Space Bunny Alpha — an anonymous frontier-class stealth preview with blazing-fast inference, strong coding performance, and native multimodal input — joins the catalog as a free model. It always reasons with adjustable effort over a 1M-token context window, served $0 via OpenRouter, Nous, and OpenCode Zen platform hops.

    Release note →
  8. Added

    Verified AIHubMix BYOK models

    Five verified text-output models join the catalog through AIHubMix and other BYOK routes: Qwen3 235B A22B Instruct 2507, Mercury 2.5, L3.1 Euryale 70B, GPT-6 Sol, and GPT-6 Luna. Published provider prices and context windows are recorded per upstream; AnyRouter does not add platform markup to BYOK traffic.

    Release note →
  9. Disabled

    Union Alpha disabled

    Previously verified platform routes return 404 for stealth/union-alpha. Hidden from the catalog; configured BYOK fallbacks remain available for users who saved keys.

    Release note →
  10. Added

    Claude Opus 5.5

    Anthropic's Claude Opus 5.5 — the successor to Claude Opus 5 for long-running agentic coding and knowledge work — joins the catalog. It delivers Fable 5.1-level performance at 40% lower cost ($4 / $20 per million input / output tokens), with adaptive thinking always on, 1M-token context, vision input, and tool use.

    Release note →
  11. Added

    Grok 4.7

    xAI's Grok 4.7 joins the catalog — 500K context, adjustable reasoning, vision in, and six BYOK upstream routes.

    Release note →
  12. Disabled

    Llama Nemotron Embed VL 1B v2 disabled

    NVIDIA's hosted NIM API returns 404 for this catalog id. The model is downloadable self-host NIM only. The live NVIDIA embedding SKU is nvidia/nemotron-3-embed-1b.

    Release note →
  13. Added

    Union Alpha now on AnyRouter

    Stealth preview Union Alpha joins the catalog free — multimodal research, coding, and agentic workflows at $0.

    Release note →
  14. Added

    Agnes 3.0 Flash and Nemotron Parse 2.0

    Agnes AI's Agnes 3.0 Flash (512K, vision in, text out) and NVIDIA Nemotron Parse 2.0 (document images in, structured markdown out) join the catalog.

    Release note →
  15. Added

    AIHubMix BYOK hops on DeepSeek V4.1 Flash, Solar Pro 4, and GPT-6 Astra

    Existing listings pick up an aihubmix-byok hop. Same catalog ids — not new SKUs. Platform aihubmix stays off (#3031).

    Release note →
  16. Added

    Context-window floors on anyrouter/auto, free, byok, and the other virtual ids

    Append [1m] or [500k] to any first-party virtual id to keep only members with at least that context window. Not a new listing — the same preset, with a token floor.

    Release note →
  17. Upstream

    Poolside platform route disabled

    The Poolside platform key returns HTTP 429 usage limit exceeded. Listings stay: poolside/laguna-s-2.1, poolside/laguna-xs-2.1, poolside/laguna-m.1, poolside/laguna-xs.2 — not new SKUs. Platform poolside is off; hops parked. poolside-byok stays live.

    Poolside (platform)HTTP 429 usage limit exceeded (#3297, #3159). Hops parked (#3298). poolside-byok stays on.
    Release note →
  18. Added

    Ling-3.0-flash-VL, Sante, Fin, and OpenRouter :free hops

    inclusionAI Ling-3.0-flash-VL, Sante, and Fin join the catalog, plus Nex-N2.5 Mini/Pro. OpenRouter :free is the wire and an alias — Nemotron 3 Ultra and Inkling Small pick up hue hops so we can burn remaining free quota.

    Release note →
  19. Added

    DeepSeek V4.1 Flash

    DeepSeek V4.1 Flash joins the catalog as deepseek/deepseek-v4.1-flash, with official DeepSeek API and OpenRouter BYOK hops. Peak list $0.30 / $1.20 per 1M tokens.

    Release note →
  20. Disabled

    2 models disabled

    Caugiay backend disabled 2026-09-07 (#3146/#2871) — HTTP 401; this model's only platform upstream was caugiay. Plus 1 more.

    Release note →
  21. Upstream

    Caugiay and SambaNova platform routes disabled

    The Caugiay platform key returns HTTP 401 Access denied, and the SambaNova platform account requires a payment method (HTTP 402). BYOK SambaNova stays live; Caugiay has no BYOK sibling. NVIDIA NIM hops that 404 for models NVIDIA does not serve were dropped per listing, not the whole nvidia backend.

    Caugiay (platform)FPT Cloud 401 Invalid API Key (#3146, #2871).SambaNova (platform)HTTP 402 PAYMENT_METHOD_REQUIRED on google/gemma-4-31b (#3172). sambanova-byok stays on.
    Release note →
  22. Added

    Claude Fable 5.1 and GPT-6 Astra on Experiential

    Experiential Cloud promotional hops for Claude Fable 5.1 and a new GPT-6 Astra listing (BYOK live; hosted platform hop pending #2842 recert).

    Release note →
  23. Disabled

    Pareto disabled

    Cloudflare AI Gateway Unified Billing is exhausted (#3121), so the platform cloudflare backend is disabled. Keep the verified unbiased/pareto model as a disabled catalog tombstone until the backend is funded and re-enabled.

    Release note →
  24. Added

    Muse Spark 1.3

    Meta's Muse Spark 1.3 joins the catalog over AIHubMix BYOK — a multimodal reasoning model for agentic tasks with a 1M-token context window.

    Release note →
  25. Added

    Gemini 3.8 Flash and Mercury 2.5 Preview

    Google Gemini 3.8 Flash and Inception Labs Mercury 2.5 Preview join the catalog. Claude Fable 5.1 and Qwen3.8 Max pick up AIHubMix hops on their existing ids.

    Release note →
  26. Disabled

    Ox Alpha disabled

    This model has been removed. Requests that still use stealth/ox-alpha or stealth/ox-alpha[1m] are routed to anyrouter/auto.

    Release note →
  27. Added

    Claude Fable 5.1

    Anthropic Claude Fable 5.1 joins the catalog as one listing, routed through Anthropic BYOK and Cloudflare Workers AI, with OpenRouter BYOK as a fallback.

    Release note →
  28. Added

    Hy4 Preview on AIHubMix BYOK

    Tencent Hunyuan Hy4 Preview joins the catalog as tencent/hy4-preview, routed through AIHubMix BYOK at the aggregator's published list rates.

    Release note →
  29. Added

    Qwen3.8 Flash on AnyRouter

    Alibaba Qwen3.8 Flash via AIHubMix BYOK. List $0.1126 / $0.38 per 1M.

    Release note →
  30. Disabled

    2 models disabled

    Every remaining platform host returns 404 for this catalog id. NVIDIA's current Nano omni/VL successor is nvidia/nemotron-3-nano-omni-30b-a3b-reasoning. Plus 1 more.

    Release note →
  31. Disabled

    Hy3 disabled

    Hosted probes for this catalog id fail with a bad request. The model is unlisted rather than advertised as a live hosted chat model.

    Release note →
  32. Disabled

    5 models disabled

    The remaining platform host is out of credit, and the other platform route is already disabled. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model. Plus 4 more.

    Release note →
  33. Added

    Catalog sweep — GLM-5.3, vision, Agnes, Sol Disc

    Eight models from the post-closeout sweep: Z.AI GLM-5.3 (coding preview), DeepSeek vision + 0731-fast, Microsoft MAI-Thinking-1, OpenAI GPT-5.6 Sol Disc, and Agnes 2.5 Flash/Pro/Alpha via AIHubMix.

    Release note →
  34. Added

    GLM-5.3-Flash — Ox Alpha, named

    The Ox Alpha preview is GLM-5.3-Flash: native multimodal, 1M context, MIT-licensed 320B-A18B. Same AnyRouter routes, catalog id z-ai/glm-5.3-flash.

    Release note →
  35. Added

    Qwen3.8-Flash-Next from Empero

    Free community endpoint for Qwen3.8-Flash-Next — open by default, from Empero research lab. Catalog id qwen/qwen3.8-flash-next.

    Release note →
  36. Added

    AIHubMix free tier — 15 more $0 routes

    AIHubMix joins the platform free pool: MiniMax M2.7, Kimi K3, Gemini 3.x Flash, Nemotron 3, GPT-OSS 20B, and more, all routed at $0.

    Release note →
  37. Disabled

    2 models disabled

    The only platform host (hoian) is balance-quarantined and no longer serves this catalog id reliably. Unlisted until a live hosted path returns. Plus 1 more.

    Release note →
  38. Added

    Ox Alpha

    Ox Alpha joins the catalog free — a stealth reasoning model for coding, agentic work, and long-horizon engineering.

    Release note →
  39. Added

    Qwen 27B dense SKUs and Llama 3.3 70B Instruct

    Qwen3.6-27B, Qwen3.8-27B, Gemma 3 27B IT, and Meta Llama 3.3 70B Instruct join the catalog as paid chat models.

    Release note →
  40. Added

    Qwen3.8-27B and DeepSeek V4 Flash, 50% off

    Half off through 1 Sep. Same ids, same API — just a friendlier bill.

    Release note →
  41. Disabled

    12 models disabled

    The platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model. Plus 11 more.

    Gemma SEA-LION v4 27B ITThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3 MaxThe partner-hosted platform path is no longer billable, and remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen 3.7 MaxCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.Qwen 3.7 PlusCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.Qwen 3.8 MaxCloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0.DeepSeek R1 Distill Qwen 32BThe platform Workers AI host no longer serves this catalog id. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.DiffusionGemma 26B A4BChat completions return empty content for this discrete-diffusion model under normal chat token budgets. It is unlisted rather than advertised as a live chat model.Gemma 3 12B ITThe platform Workers AI host no longer serves this catalog id. Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Llama 3.2 1B InstructThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Llama 3.2 3B InstructThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3-32BThe platform host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted chat model.Qwen3 Embedding 0.6BThe platform Workers AI host no longer serves this catalog id (model_unavailable). Remaining routes require a user-owned key, so the model is unlisted rather than advertised as a live hosted embedding model.
    Release note →
  42. Disabled

    Mistral 7B Instruct v0.1 disabled

    Cloudflare Workers AI retired this SKU on 2026-05-30; live smoke 2026-08-14/18 returns model_unavailable (no_route). It was this model's only upstream, so it was disabled rather than the route pruned.

    Release note →
  43. Upstream

    GitHub Models platform and BYOK routes disabled

    GitHub Models was fully retired on 2026-07-30. The inference API and BYOK endpoints return HTTP 410, so the github and github-byok backends are off.

    GitHub Models (platform)Retired 2026-07-30 — HTTP 410 on every try (#2237).GitHub Models (BYOK)Retired with the platform API on 2026-07-30 (#2237).
    Release note →
  44. Added

    DeepSeek V4 Pro

    DeepSeek V4 Pro (0813) joins the catalog, hosted on Cloudflare Workers AI with a 1M-token context window.

    Release note →
  45. Disabled

    4 models disabled

    Cloudflare has not published per-token pricing for this partner-hosted SKU (rate visible only in the CF dashboard once provisioned); shipping it live on the metered `cloudflare` backend without a real rate would bill $0. Plus 3 more.

    Release note →
  46. Added

    Gemini 3.7 Flash and latest frontier chat SKUs

    Google Gemini 3.7 Flash joins the catalog, plus Qwen3.8 2.4T A95B, Muse Spark 1.2, Inkling Small, Solar Pro 4, Seed 2.1 Turbo, Seed 2.0 Code, Claude Opus 5 Fast, and Dots3-Note Preview.

    Release note →
  47. Disabled

    5 models disabled

    CF Workers AI LoRA host 404 / model_unavailable on every smoke probe (2026-08-14). Canonical Gemma instruct SKUs remain (gemma-3 / gemma-4). Sibling gemma-2b-it-lora already disabled. Plus 4 more.

    Release note →
  48. Disabled

    7 models disabled

    SiliconFlow no longer lists this id; its catalogue carries gemma-4-12B-it, gemma-4-26B-A4B-it and gemma-4-31B-it instead. Plus 6 more.

    Gemma-4-27B-itSiliconFlow no longer lists this id; its catalogue carries gemma-4-12B-it, gemma-4-26B-A4B-it and gemma-4-31B-it instead.Nemotron 3.5 NanoIts only upstream went offline and no other provider serves this model — the NVIDIA Nemotron 3.5 line is now published as Lightning and Content Safety only. Nemotron 3.5 Lightning is the closest replacement.GPT-5 ChatCloudflare Workers AI retired this SKU; the binding returns HTTP 410 (Gone) on every attempt. In the same 2026-07-25 smoke run, the sibling partner-hosted model openai/gpt-5.6-luna reached the same OpenAI-on-Workers-AI rail and got HTTP 400 (request rejected, not gone), showing the retirement is specific to this SKU rather than a rail-wide outage or billing issue. It was this model's only upstream, so it was disabled rather than the route pruned.GPT-5.1 ChatCloudflare Workers AI retired this SKU; the binding returns HTTP 410 (Gone) on every attempt. In the same 2026-07-25 smoke run, the sibling partner-hosted model openai/gpt-5.6-luna reached the same OpenAI-on-Workers-AI rail and got HTTP 400 (request rejected, not gone), showing the retirement is specific to this SKU rather than a rail-wide outage or billing issue. It was this model's only upstream, so it was disabled rather than the route pruned.Laguna M.1Poolside's live listing now carries only Laguna XS 2.1 and Laguna S 2.1 — this generation has been withdrawn, and every route to it (including the aggregator mirrors) now 404s. Laguna S 2.1 is the direct replacement.Laguna XS.2Poolside's live listing (inference.poolside.ai/v1/models) now carries only laguna-xs-2.1 and laguna-s-2.1 — this generation is gone, so every upstream here 404s.GLM-4.6V-FlashNo upstream currently serves this model — its only provider went offline, and Z-AI's own API publishes the text GLM line only, not the vision variants. GLM-4.6V is the closest available alternative.
    Release note →
  49. Added

    Nemotron Lightning, Muse Glimmer, Seedance 2.5, DeepSeek NIM Flash

    NVIDIA NIM additions (Nemotron 3.5 Lightning 30B, Meta Muse Glimmer 30B, DeepSeek-V4-Flash 0731), ByteDance Seedance 2.5 on Workers AI, and free Requesty route for Ling-3.0-tiny.

    Release note →
  50. Added

    Free-route expansion + Liquid LFM 2.5

    Expanded free platform routes across NVIDIA Nemotron free models, GPT-OSS-20B free, Ling-3.0-tiny on Novita, and new Liquid LFM 2.5 2.6B free. CommandCode free/promo routes for Laguna S 2.1 and MiMo V2.5 Pro.

    Release note →
  51. Added

    Grok 4.6

    xAI Grok 4.6 joins the catalog via xAI BYOK and OpenRouter BYOK. 500K context, $2/$6 list on xAI.

    Release note →
  52. Disabled

    3 models disabled

    The catalogue is text/embedding-output only; video models stay on disk for future re-enable (same posture as seedance-2.0-mini). Unpriced CF video would also fail metered-catalog-billing. Plus 2 more.

    Release note →
  53. Added

    EmbeddingGemma 300M and Ling-3.0-tiny

    Google's EmbeddingGemma 300M open embedding model and inclusionAI's Ling-3.0-tiny free MoE chat model join the catalog.

    Release note →
  54. Added

    Sakana Namazu

    Sakana AI's Japanese-specialized LLM (built on Kimi K2.6) joins the catalog via BYOK.

    Release note →
  55. Added

    Nemotron 3.5 Nano

    NVIDIA's Nemotron 3.5 Nano joins the catalog via Blackbox AI BYOK.

    Release note →
  56. Added

    Nemotron Nano, GLM-4.5/4.6, MiniMax M2, Qwen3.5-9B/27B

    Eight new models discovered from upstream providers and added to the catalog — NVIDIA's free Nemotron Nano 12B V2 VL and 9B V2 via OpenRouter, Z-AI's GLM-4.5 and GLM-4.6 (with vision variant), MiniMax M2, and Alibaba's Qwen3.5-9B and Qwen3.5-27B.

    Release note →
  57. Added

    Qwen3.7 Flash

    Alibaba's Qwen3.7 Flash — a vision-language reasoning model for multimodal agents, visual coding, search, and computer interaction — joins the catalog via OpenRouter BYOK.

    Release note →
  58. Disabled

    16 models disabled

    The catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend). Plus 15 more.

    FLUX.2 DevThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).FLUX.2 Klein 9BThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Seedream 5 ProThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Eleven Flash v2.5The catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Eleven Multilingual v2The catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 LargeThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 Medium TurboThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Krea 2 MediumThe catalogue was narrowed to models whose output is text or embeddings; this model's output is an image, so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).Llama 2 7B Chat FP16Cloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 2 7B Chat INT8Cloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3 8B Instruct AWQCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3 8B InstructCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3.1 8B Instruct AWQCloudflare Workers AI retired this SKU on 2026-05-30; the binding returns HTTP 410 (deprecated) and the model is absent from the account's live catalogue. It was this model's only upstream, so it was disabled rather than the route pruned.Llama 3.1 8B Instruct FastThe model is gone from Cloudflare's live account catalogue and the binding returns HTTP 404 (model does not exist). It was this model's only upstream, so it was disabled rather than the route pruned.Grok STTThe catalogue was narrowed to text- and embedding-output models. This model's output IS text, but its input is audio — it is a speech-to-text SKU, named explicitly in the same decision, so it goes out of service with the image / video / TTS models rather than being kept on the technicality of its output modality.Grok TTSThe catalogue was narrowed to models whose output is text or embeddings; this model's output is audio (text-to-speech), so it was taken out of service. Related: #1551 (these SKUs bill $0 on a metered backend).
    Release note →
  59. Added

    Claude Opus 5

    Anthropic's Claude Opus 5 — the flagship model for demanding reasoning, coding, and long-horizon agentic work — joins the catalog, BYOK-only like its Claude siblings.

    Release note →
  60. Added

    Ling-3.0-flash

    inclusionAI's Ling-3.0-flash — a 124B-parameter Mixture-of-Experts model (~5.1B active) tuned for token-efficient, production-scale agentic inference — joins the catalog, served free with a user-owned API key (BYOK).

    Release note →
  61. Upstream

    Google unified-billing and Blackbox platform routes disabled

    The Cloudflare AI Gateway unified-billing balance and the Blackbox platform budget both ran out. BYOK siblings stay live; platform traffic fails over or stops.

    Google AI (platform)Unified billing 402 / insufficient balance (#1321). No recharge.Blackbox AI (platform)Provider budget cap (~$3). No traffic until the account is topped up (#1326).
    Release note →
  62. Added

    Gemini 3.6 Flash & Gemini 3.5 Flash Lite

    Google's Gemini 3.6 Flash and Gemini 3.5 Flash Lite join the catalog as partner-hosted SKUs on Cloudflare Workers AI, with OpenRouter BYOK fallback.

    Release note →
  63. Added

    Laguna S 2.1

    Poolside's latest coding agent model Laguna S 2.1 (118B total, 8B active parameters, 1M context, 70.2% Terminal-Bench 2.1, 40.4% DeepSWE) joins the catalog — open-weight under OpenMDW-1.1.

    Release note →
  64. Added

    Inkling + Nemotron 3 Embed 1B

    Thinking Machines Lab's first open-weights foundation model Inkling (975B MoE, 41B active, 1M context, multimodal input with switchable reasoning) and NVIDIA's Nemotron-3-Embed-1B multilingual embedding model (2048 dims, 34 languages) join the catalog via NVIDIA NIM.

    Release note →
  65. Added

    Kimi K3

    Moonshot AI's Kimi K3 is now available on AnyRouter — an open-source model with a 1M-token context window and native visual understanding (text, image, and video). Built for long-horizon software engineering and deep reasoning where code meets visual and spatial thinking, with thinking mode always on.

    Release note →
  66. Upstream

    OpenAI and DeepInfra platform routes disabled

    The platform OpenAI key has no quota (429 insufficient_quota) and DeepInfra has no payment method (402). BYOK users are unaffected.

    OpenAI (platform)No quota — 5,179 attempts / 0 ok over 30d (#1283).DeepInfra (platform)No payment method — 1,240 attempts, 76.9% fail (#1284).
    Release note →
  67. Upstream

    Named OpenRouter platform backend disabled

    The public `openrouter` backend is off. Platform traffic uses the masked hue route; user keys stay on openrouter-byok.

    OpenRouter (named platform)Disabled #1152. Hue is the platform path; openrouter-byok is BYOK.
    Release note →
  68. Added

    Grok 4.5 and Laguna XS 2.1

    Two frontier additions: xAI's Grok 4.5 joins the catalog and Poolside ships a refreshed Laguna XS 2.1 coding model.

    Release note →
  69. Added

    Free-tier open models via SiliconFlow

    A batch of open-weight models now has a free-tier upstream — Qwen3 across sizes, Gemma 4 27B, and Tencent Hy3.

    Release note →
  70. Added

    Claude Sonnet 5

    Anthropic's Claude Sonnet 5 is available through BYOK — bring your Anthropic key and route it behind the gateway.

    Release note →
  71. Added

    Codex-class OpenAI models and GLM-5

    The OpenAI Codex family lands — GPT-5.3 / 5.2 Codex and 5.1 Codex Max — alongside Z.ai's GLM-5 and Qwen3.7 Max.

    Release note →
  72. Added

    LongCat-2.0

    The LongCat provider joins the catalog with LongCat-2.0, available via BYOK.

    Release note →
  73. Added

    Cerebras, Venice, and Wafer providers

    Three new upstreams broaden coverage — Cerebras GPT-OSS, Venice's MiniMax M2.7, and Wafer's aggregated MiniMax M3 and GLM-5.2.

    Release note →
  74. Disabled

    13 models disabled

    Its BYOK aggregator upstream no longer lists this model id (verify:upstreams MISSING 2026-07-16), and that was its only upstream. See #1283. Plus 12 more.

    Release note →

Product updates

v1.7.0ImprovementSeptember 15, 2026

BYOK providers on the model page, newest key first

Ready BYOK keys sort newest-first on model pages, including /model/anyrouter/byok. Donated keys follow, then the rest. Not a new listing — same anyrouter/byok id (#3300).

  • Providers with a ready key appear first; newest created_at on top
  • Then donated keys, then providers with no key yet
  • UI on /model/anyrouter/byok — no new catalog id
v1.6.0FeatureSeptember 15, 2026

Virtual model ids accept [1m] and [500k] context floors

anyrouter/auto, anyrouter/free, anyrouter/byok, and the other first-party virtual ids keep only members at or above a context-window floor when you append [1m] or [500k]. Same listing, not a new SKU. Note at blog.anyrouter.dev/releases/2026-09-15-virtual-context

  • `anyrouter/auto[1m]` / `[500k]` set provider.min_context (1M / 500k tokens)
  • Same suffix on anyrouter/free, byok, coding, agent, hermes, cowork, latest
  • anyr claude peels the floor into extra body; model pages preview the filtered chain
  • Release note → blog.anyrouter.dev/releases/2026-09-15-virtual-context
v1.5.1FeatureAugust 20, 2026

Qwen3.8-27B and DeepSeek V4 Flash 50% off through 1 September

Qwen3.8-27B and DeepSeek-V4-Flash are 50% off list rates on AnyRouter through 1 Sep 2026 UTC. Sticker prices stay on the model page; billed rates are half. After that date the promo drops automatically — no catalog rewrite.

  • `qwen/qwen3.8-27b` list $0.30 / $3.25, billed $0.15 / $1.625 per 1M tokens
  • `deepseek/deepseek-v4-flash` list $0.14 / $0.28, billed $0.07 / $0.14
  • Try via AI SDK, anyr CLI, or the playground — note at blog.anyrouter.dev/releases/2026-08-20-qwen3-8-27b-sale
v1.5.0FeatureAugust 19, 2026

AnyRouter Agent — free on-site support copilot

Ask AnyRouter on the homepage and at /agent. A dedicated anyrouter-support Worker answers with free text models, docs search, and optional account tools after you opt in. Writes wait for an explicit Confirm card.

  • Ask AnyRouter in the public header (closed by default — no data until opened)
  • Full page at /agent plus agent.anyrouter.dev 301 to the apex page
  • Docs, catalog, and status tools for everyone; plan/usage/key prefix after opt-in
v1.4.1FixAugust 18, 2026

Docs and free-tier membership match the live text catalog

Customer docs, homepage examples, and the anyrouter/free membership list now match what production actually serves: text models (and embeddings via *:free), not retired GitHub Models copy-paste. Image generation stays a 410 tombstone.

  • anyrouter/free members are the text models it routes to; embedding SKUs stay on *:free only
  • First-request docs and homepage snippets use live text models such as z-ai/glm-4.7-flash
  • Leftover GitHub Models rows on unpublished SKUs are dropped
v1.4.0FixAugust 18, 2026

Routing health and free-path failover

Free-tier routing now skips smoke-excluded backends so a healthy sibling can serve. Public health providers list routing backends with last-hour traffic instead of idle catalog owners. BYOK-only classification ignores disabled leftover route ids.

  • Free-tier failover honors smoke exclusion (failCount ≥ 10)
  • Health providers are routing backend ids, not catalog owners
  • Ghost disabled backends no longer hide the BYOK-only 4xx
v1.3.0FeatureMarch 28, 2026

Privacy controls

Per-account provider blocklists and zero-retention routing preferences for compliance-sensitive workloads.

  • Block specific upstream providers per account
  • Opt-in zero-retention routing from Privacy settings
  • Provider restrictions enforced before routing
v1.2.1ImprovementMarch 15, 2026

Routing improvements

Faster upstream failover and clearer routing diagnostics when a provider is unavailable.

  • Shorter failover detection on upstream errors
  • Clearer upstream error metadata in API responses
  • Improved fallback chain visibility in logs
v1.2.0FeatureMarch 1, 2026

BYOK Support

Bring Your Own Key — use your existing provider API keys through AnyRouter for direct billing with full routing benefits.

  • Support for OpenAI, Anthropic, Google keys
  • Load-balancing strategies across multiple keys
  • BYOK usage reported separately in Dashboard → Usage
v1.1.0FeatureFebruary 15, 2026

Request logs & retention

Dashboard request logs with per-plan retention windows and searchable metadata for every inference call.

  • Per-request metadata in Dashboard → Logs
  • Retention windows by plan (7 days Free/Go, 30 days Pro+)
  • Filter by model, status, and API key
v1.0.1FixFebruary 1, 2026

Rate Limit Fix

Fixed an issue where rate limits were not correctly applied to BYOK requests.

  • Correct rate limiting for BYOK
  • Improved error messages
v1.0.0FeatureJanuary 15, 2026

AnyRouter Launch

Initial release of AnyRouter — the edge-native AI gateway. Route to 180+ models from 17 providers with a single API.

  • 180+ AI models supported
  • 17 provider integrations
  • OpenAI-compatible API
  • Auto-fallback routing
  • Usage analytics dashboard