Skip to content

Embeddings API

POST /api/v1/embeddings — generate text embeddings for semantic search, clustering, and classification. OpenAI-compatible.

Generate embedding vectors for one or more texts. Drop-in compatible with OpenAI's POST https://api.openai.com/v1/embeddings — swap the base URL and key.

POST/api/v1/embeddings

Authenticate with an LLM API key (sk-ar-v1-…) from /dashboard/keys, sent as Authorization: Bearer …. Management keys (ak_…) are not accepted.

Request

FieldTypeRequiredDescription
modelstringyesModel id to use for embeddings. See Models.
inputstring | string[]yesText or array of texts to embed.
encoding_format"float" | "base64"noFormat for returned embeddings. Defaults to float.
dimensionsinteger (≥ 1)noReduce the output to this many dimensions, when the upstream model supports truncation.
userstringnoStable end-user identifier for abuse tracking. Passed through to the upstream.
session_idstringnoSticky-session id (≤256 chars) that groups related requests in Request Logs. Also settable via the x-session-id header; the body field wins.

Unrecognized fields are forwarded to the upstream — unsupported ones may be ignored or rejected depending on the provider.

Routing controls

AnyRouter routes each request through a fallback chain. Override the chain with these headers (no request body changes):

HeaderDescription
X-AnyRouter-ProviderForce a specific backend (e.g. cloudflare, deepinfra). Comma-separated to restrict to a set.
X-Routing-StrategyRouting strategy override. Common values: cheapest, fastest, most-available.

Caching

Embeddings are deterministic for a given (model, input, dimensions, encoding_format) tuple. Opt into Cloudflare AI Gateway caching by setting either header on the request:

HeaderDescription
cf-aig-cache-ttlCache the upstream response for N seconds. Identical subsequent requests are served from the gateway without billing the upstream.
cf-aig-skip-cacheSet to true to bypass the cache for this request.
cf-aig-cache-keyOptional custom cache key. Defaults to a hash of the request body.

When a response comes from the cache, AnyRouter echoes a X-AnyRouter-Cache-Status: HIT response header.

Response

{
  "id": "embd_req_01HXYZ...",
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "index": 0,
      "embedding": [0.0023064255, -0.009327292, 0.015797347]
    }
  ],
  "model": "qwen/qwen3-embedding-8b",
  "usage": {
    "prompt_tokens": 9,
    "total_tokens": 9,
    "cost": 0.00000011
  }
}
FieldTypeDescription
idstringUnique identifier for the response. Mirrors X-Request-ID.
object"list"Always list.
dataobject[]One entry per input string, in request order.
data[].object"embedding"Always embedding.
data[].indexintegerZero-based index of the input that produced this vector.
data[].embeddingnumber[] | stringFloat vector, or base64-encoded bytes when encoding_format=base64.
modelstringThe model that served the request.
usage.prompt_tokensintegerToken count of the input.
usage.total_tokensintegerSame as prompt_tokens for embeddings.
usage.costnumberUSD cost of the request. Present for paid models; omitted when free or pricing is unknown.
usage.prompt_tokens_detailsobjectPer-modality token breakdown. Only present when the input contains 2+ modalities and the upstream returns modality-level counts. Sub-fields: text_tokens, image_tokens, audio_tokens, file_tokens, video_tokens.

Response headers

HeaderDescription
X-Request-IDStable request identifier — quote this when filing a support issue.
X-AnyRouter-ProviderWhich backend actually served the request after fallback.
X-AnyRouter-Cache-StatusPresent when AI Gateway caching is enabled. HIT / MISS / BYPASS.

Encoding format

Request encoding_format: "base64" to reduce JSON payload size for large batches:

// float (default)
{ "embedding": [0.0023, -0.0235, 0.0456] }

// base64
{ "embedding": "AAAACgkJMzdG7y..." }

Dimensions

When the upstream supports it, request a truncated vector. Unsupported dimensions values are passed through and may be rejected upstream:

{
  "model": "qwen/qwen3-embedding-8b",
  "input": "Your text here",
  "dimensions": 512
}

Models

ModelDimensionsContextNotes
qwen/qwen3-embedding-8b409632,768 tokensHigher quality, longer context.

Pricing and live availability are listed on the Models page.

Examples

curl https://anyrouter.dev/api/v1/embeddings \
  -H "Authorization: Bearer $ANYROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3-embedding-8b",
    "input": "The quick brown fox jumps over the lazy dog"
  }'
{
  "model": "qwen/qwen3-embedding-8b",
  "input": [
    "First text to embed",
    "Second text to embed",
    "Third text to embed"
  ]
}

The response contains one data[] entry per input, ordered by index.

const embed = (text: string | string[]) =>
  fetch("https://anyrouter.dev/api/v1/embeddings", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.ANYROUTER_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ model: "qwen/qwen3-embedding-8b", input: text }),
  }).then((r) => r.json())

const docs = await embed(documents) // batch
const query = await embed(userQuery) // single
const similarity = cosineSimilarity(query.data[0].embedding, docs.data[0].embedding)

Batch inputs — one request with N strings is cheaper and faster than N requests. Embeddings are deterministic per (model, input), so cache by content hash. Normalize vectors before cosine similarity, and pick the smallest model that meets your recall needs.

Errors

StatusCodeWhen it happens
400missing_required_fieldsmodel or input missing.
400invalid_request_errorMalformed JSON or unsupported field value.
401invalid_api_keyMissing or invalid bearer token.
402insufficient_balanceAccount balance below the request cost (metadata.type: "insufficient_credits").
404model_unavailableNo upstream serves the requested model.
429rate_limit_exceededPer-key rate limit hit.
500 / 502 / 503upstream_error, upstream_exhaustedAll upstream providers failed.

Error responses follow the standard error contract.

  • Models — list and inspect every supported model
  • Errors — status codes and fixes
  • Smart Routing — provider preferences and fallback