Skip to content

A/B Compare Models in the Playground

Send the same prompt to several models from one AnyRouter key and compare answers, latency, and cost side by side before you commit to one.

Before you wire a model into production, see how a few candidates answer the same prompt — and what each one costs and how fast it responds. With AnyRouter, one key reaches every model, so comparing them is a matter of swapping the model field.

The idea

Because every model lives behind one key and one base URL, a comparison is just the same request sent N times with a different model. Each response reports its own latency and token usage, so you can weigh quality against cost in one pass.

flowchart LR
  prompt["Same prompt"] --> ar["AnyRouter (one key)"]
  ar --> m1["openai/gpt-5.4-mini"]
  ar --> m2["anthropic/claude-haiku-4.5"]
  ar --> m3["z-ai/glm-4.7-flash"]
  m1 --> cmp["Compare answer · latency · tokens"]
  m2 --> cmp
  m3 --> cmp

How it works

Every run — from the Playground or from code — goes through your AnyRouter account, so the cost of each comparison shows up in your usage explorer alongside everything else. The usage block on each response reports prompt, completion, and total tokens.

Implementation

The Playground lets you chat against any catalog model with your account key already attached:

  1. Open the Playground — your account key is attached automatically.
  2. Pick a model and send your prompt.
  3. Change the model and send the same prompt again.
  4. Compare the answers, plus the per-response latency and token usage shown in the metrics footer.

Send the identical messages array to each candidate and collect the results. One key, one base URL:

import { AnyRouter } from "@anyr/sdk"

const client = new AnyRouter() // reads ANYROUTER_API_KEY

const prompt = "Rewrite this commit message to be clear and concise: 'fix stuff'"

const candidates = [
  "openai/gpt-5.4-mini",
  "anthropic/claude-haiku-4.5",
  "z-ai/glm-4.7-flash",
]

for (const model of candidates) {
  const started = Date.now()
  const res = await client.chat.completions.create({
    model,
    messages: [{ role: "user", content: prompt }],
  })
  console.log(`\n=== ${model} (${Date.now() - started}ms) ===`)
  console.log(res.choices[0].message.content)
  console.log("tokens:", res.usage?.total_tokens)
}

If you run comparisons often, pass a session_id so the requests group together in your logs and are easy to find later.

{
  "model": "z-ai/glm-4.7-flash",
  "messages": [{ "role": "user", "content": "..." }],
  "session_id": "model-bakeoff-2025-01"
}