Embeddings API
POST /api/v1/embeddings — generate text embeddings for semantic search, clustering, and classification. OpenAI-compatible.
Generate embedding vectors for one or more texts. Drop-in compatible with OpenAI's POST https://api.openai.com/v1/embeddings — swap the base URL and key.
/api/v1/embeddingsAuthenticate with an LLM API key (sk-ar-v1-…) from /dashboard/keys, sent as Authorization: Bearer …. Management keys (ak_…) are not accepted.
Request
| Field | Type | Required | Description |
|---|---|---|---|
model | string | yes | Model id to use for embeddings. See Models. |
input | string | string[] | yes | Text or array of texts to embed. |
encoding_format | "float" | "base64" | no | Format for returned embeddings. Defaults to float. |
dimensions | integer (≥ 1) | no | Reduce the output to this many dimensions, when the upstream model supports truncation. |
user | string | no | Stable end-user identifier for abuse tracking. Passed through to the upstream. |
session_id | string | no | Sticky-session id (≤256 chars) that groups related requests in Request Logs. Also settable via the x-session-id header; the body field wins. |
Unrecognized fields are forwarded to the upstream — unsupported ones may be ignored or rejected depending on the provider.
Routing controls
AnyRouter routes each request through a fallback chain. Override the chain with these headers (no request body changes):
| Header | Description |
|---|---|
X-AnyRouter-Provider | Force a specific backend (e.g. cloudflare, deepinfra). Comma-separated to restrict to a set. |
X-Routing-Strategy | Routing strategy override. Common values: cheapest, fastest, most-available. |
Caching
Embeddings are deterministic for a given (model, input, dimensions, encoding_format) tuple. Opt into Cloudflare AI Gateway caching by setting either header on the request:
| Header | Description |
|---|---|
cf-aig-cache-ttl | Cache the upstream response for N seconds. Identical subsequent requests are served from the gateway without billing the upstream. |
cf-aig-skip-cache | Set to true to bypass the cache for this request. |
cf-aig-cache-key | Optional custom cache key. Defaults to a hash of the request body. |
When a response comes from the cache, AnyRouter echoes a X-AnyRouter-Cache-Status: HIT response header.
Response
{
"id": "embd_req_01HXYZ...",
"object": "list",
"data": [
{
"object": "embedding",
"index": 0,
"embedding": [0.0023064255, -0.009327292, 0.015797347]
}
],
"model": "qwen/qwen3-embedding-8b",
"usage": {
"prompt_tokens": 9,
"total_tokens": 9,
"cost": 0.00000011
}
}
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the response. Mirrors X-Request-ID. |
object | "list" | Always list. |
data | object[] | One entry per input string, in request order. |
data[].object | "embedding" | Always embedding. |
data[].index | integer | Zero-based index of the input that produced this vector. |
data[].embedding | number[] | string | Float vector, or base64-encoded bytes when encoding_format=base64. |
model | string | The model that served the request. |
usage.prompt_tokens | integer | Token count of the input. |
usage.total_tokens | integer | Same as prompt_tokens for embeddings. |
usage.cost | number | USD cost of the request. Present for paid models; omitted when free or pricing is unknown. |
usage.prompt_tokens_details | object | Per-modality token breakdown. Only present when the input contains 2+ modalities and the upstream returns modality-level counts. Sub-fields: text_tokens, image_tokens, audio_tokens, file_tokens, video_tokens. |
Response headers
| Header | Description |
|---|---|
X-Request-ID | Stable request identifier — quote this when filing a support issue. |
X-AnyRouter-Provider | Which backend actually served the request after fallback. |
X-AnyRouter-Cache-Status | Present when AI Gateway caching is enabled. HIT / MISS / BYPASS. |
Encoding format
Request encoding_format: "base64" to reduce JSON payload size for large batches:
// float (default)
{ "embedding": [0.0023, -0.0235, 0.0456] }
// base64
{ "embedding": "AAAACgkJMzdG7y..." }
Dimensions
When the upstream supports it, request a truncated vector. Unsupported dimensions values are passed through and may be rejected upstream:
{
"model": "qwen/qwen3-embedding-8b",
"input": "Your text here",
"dimensions": 512
}
Models
| Model | Dimensions | Context | Notes |
|---|---|---|---|
qwen/qwen3-embedding-8b | 4096 | 32,768 tokens | Higher quality, longer context. |
Pricing and live availability are listed on the Models page.
Examples
curl https://anyrouter.dev/api/v1/embeddings \
-H "Authorization: Bearer $ANYROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3-embedding-8b",
"input": "The quick brown fox jumps over the lazy dog"
}'
{
"model": "qwen/qwen3-embedding-8b",
"input": [
"First text to embed",
"Second text to embed",
"Third text to embed"
]
}
The response contains one data[] entry per input, ordered by index.
const embed = (text: string | string[]) =>
fetch("https://anyrouter.dev/api/v1/embeddings", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.ANYROUTER_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ model: "qwen/qwen3-embedding-8b", input: text }),
}).then((r) => r.json())
const docs = await embed(documents) // batch
const query = await embed(userQuery) // single
const similarity = cosineSimilarity(query.data[0].embedding, docs.data[0].embedding)
Batch inputs — one request with N strings is cheaper and faster than N requests. Embeddings are deterministic per (model, input), so cache by content hash. Normalize vectors before cosine similarity, and pick the smallest model that meets your recall needs.
Errors
| Status | Code | When it happens |
|---|---|---|
| 400 | missing_required_fields | model or input missing. |
| 400 | invalid_request_error | Malformed JSON or unsupported field value. |
| 401 | invalid_api_key | Missing or invalid bearer token. |
| 402 | insufficient_balance | Account balance below the request cost (metadata.type: "insufficient_credits"). |
| 404 | model_unavailable | No upstream serves the requested model. |
| 429 | rate_limit_exceeded | Per-key rate limit hit. |
| 500 / 502 / 503 | upstream_error, upstream_exhausted | All upstream providers failed. |
Error responses follow the standard error contract.
Related
- Models — list and inspect every supported model
- Errors — status codes and fixes
- Smart Routing — provider preferences and fallback