Responses API
/api/v1/responses — the OpenAI Responses-compatible lifecycle. Send input, stream or retrieve responses, continue state, and inspect input items.
/api/v1/responses implements the OpenAI Responses API lifecycle. Unlike Chat Completions, the response is a self-contained envelope with typed output items — text blocks, tool calls, and reasoning summaries — rather than a flat choices array.
/api/v1/responses/api/v1/responses/{response_id}/api/v1/responses/{response_id}/api/v1/responses/{response_id}/input_items/api/v1/responses/input_tokens/api/v1/responses/{response_id}/cancel/api/v1/responses/compactAuthenticate with an LLM API key (sk-ar-v1-…) from /dashboard/keys. Management keys (ak_…) are not accepted on inference routes.
Request
| Field | Type | Required | Description |
|---|---|---|---|
model | string | yes | provider/model id. See List models. |
input | string | array | yes | User input. String or array of input items (text, image, etc.). |
instructions | string | no | System-level instruction passed before the input. |
max_output_tokens | integer | no | Token budget for the response. |
temperature | number | no | 0–2. Sampling temperature. |
top_p | number | no | 0–1. Nucleus sampling. |
stream | boolean | no | If true, returns an SSE stream of Responses events. |
store | boolean | no | Defaults to true. Stored responses can be retrieved, deleted, continued with previous_response_id, and inspected with /input_items. |
previous_response_id | string | no | Continue from a stored response owned by the same workspace. |
background | boolean | no | If true, creates a queued stored response that can be retrieved or cancelled. |
tools | array | no | Function definitions the model can call. Each tool has type, name, description, parameters. |
tool_choice | string | object | no | "auto", "none", "required", or {"type":"function","name":"…"}. |
reasoning | object | no | {"effort": "low"|"medium"|"high", "summary": "auto"|"detailed"}. |
modalities | string[] | no | Output modalities requested (e.g. ["text"]). |
metadata | object | no | Arbitrary key-value pairs attached to the response. |
session_id | string | no | Sticky-session id (≤256 chars) that groups related requests in Request Logs. Also settable via the x-session-id header; the body field wins. |
provider | object | no | Provider routing preferences. Supports only, ignore, order, sort, allow_fallbacks, max_price, preferred_max_latency, and preferred_min_throughput. |
trace | object | no | Optional tracing metadata. Supports trace_id, trace_name, span_name, generation_name, and parent_span_id. |
For the full field list see /docs/api.
Stored response lifecycle
Responses are stored by default for later retrieval. Pass store: false when you do not need lifecycle operations. Stored responses are scoped to the API key workspace; another workspace cannot retrieve, delete, continue, or cancel them.
# Retrieve a stored response
curl https://anyrouter.dev/api/v1/responses/resp_abc123 \
-H "Authorization: Bearer sk-ar-your-key"
# List its input items
curl https://anyrouter.dev/api/v1/responses/resp_abc123/input_items \
-H "Authorization: Bearer sk-ar-your-key"
Response
{
"id": "resp_…",
"object": "response",
"created_at": 1714000000,
"completed_at": 1714000001,
"model": "z-ai/glm-4.7-flash",
"status": "completed",
"output": [ /* typed output items */ ],
"output_text": "Paris.",
"usage": {
"input_tokens": 14,
"output_tokens": 3,
"total_tokens": 17
},
"incomplete_details": null,
"error": null
}
status is one of "in_progress", "completed", "incomplete", or "failed". When status is "incomplete", incomplete_details.reason is "max_output_tokens" or "content_filter".
Output item types
| Type | Description |
|---|---|
message | Assistant turn with a content array of output_text blocks. |
function_call | Tool invocation with call_id, name, arguments. |
reasoning | Reasoning summary (provider-dependent; present when reasoning tokens are surfaced). |
Responses-native upstreams can also pass through server tool items such as web_search_call, file_search_call, code_interpreter_call, and computer_call. Chat Completions and Anthropic Messages upstreams currently normalize only text, reasoning, and function-call output items.
Streaming events
Set stream: true to receive server-sent events. The event sequence is:
event: response.created
data: {"type":"response.created","response":{...}}
event: response.in_progress
data: {"type":"response.in_progress","response":{...}}
event: response.output_item.added
data: {"type":"response.output_item.added","output_index":0,"item":{...}}
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_xyz","output_index":0,"content_index":0,"delta":"Par"}
event: response.output_text.done
data: {"type":"response.output_text.done","item_id":"msg_xyz","output_index":0,"content_index":0,"text":"Paris."}
event: response.output_item.done
data: {"type":"response.output_item.done","output_index":0,"item":{...}}
event: response.completed
data: {"type":"response.completed","response":{...}}
event: response.anyrouter.metadata
data: {"type":"response.anyrouter.metadata","anyrouter_metadata":{...}}
The response.anyrouter.metadata frame is emitted after response.completed and carries billing and upstream info. Standard Responses API clients safely ignore unknown event types.
Examples
curl https://anyrouter.dev/api/v1/responses \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-4.7-flash",
"input": "What is the capital of France?"
}'
curl https://anyrouter.dev/api/v1/responses \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-4.7-flash",
"input": "What is the weather in Paris?",
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": { "location": { "type": "string" } },
"required": ["location"]
}
}
],
"tool_choice": "auto"
}'
curl https://anyrouter.dev/api/v1/responses \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/o1-mini",
"input": "Solve: x^2 - 5x + 6 = 0",
"reasoning": { "effort": "medium", "summary": "auto" }
}'
curl https://anyrouter.dev/api/v1/responses \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.4-mini",
"input": "Summarize the deployment notes",
"provider": {
"only": ["OpenAI", "Groq"],
"sort": "latency",
"allow_fallbacks": true,
"preferred_max_latency": 250
},
"trace": {
"trace_id": "4f8c2b9670b44b49a2e71e7fd5c0b1d3",
"trace_name": "release-pipeline"
}
}'
The minimal request above returns a stored response envelope:
{
"id": "resp_abc123",
"object": "response",
"status": "completed",
"model": "z-ai/glm-4.7-flash",
"output": [
{
"type": "message",
"id": "msg_xyz",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Paris." }],
"status": "completed"
}
],
"usage": {
"input_tokens": 14,
"output_tokens": 3,
"total_tokens": 17,
"cost": 0.0000021,
"cost_details": {
"upstream_inference_input_cost": 0.0000014,
"upstream_inference_output_cost": 0.0000007
},
"is_byok": false
},
"anyrouter_metadata": {
"requestId": "resp_abc123",
"model": "z-ai/glm-4.7-flash",
"upstream": { "provider": "openai_direct", "backend": "openai_direct", "latencyMs": 412 }
}
}
When you send a reasoning request, the output array contains a reasoning item before the message item, provided the model exposes reasoning tokens.
Provider routing behavior
The provider object is applied before AnyRouter dispatches upstream requests:
onlyandignoreshape the candidate set.orderpins the preferred backend order.sortreorders candidates byprice,latency, orthroughput.max_price.promptandmax_price.completionremove candidates above your budget.preferred_max_latencyandpreferred_min_throughputapply performance cutoffs.allow_fallbacks: falsemeans AnyRouter tries only the top candidate.
See Smart Routing for the full behavior.
Errors
All errors follow the standard AnyRouter JSON envelope with a stable error.code.
| Status | Meaning | Fix |
|---|---|---|
| 400 | Malformed body or invalid parameters. | Check the request against the field table. |
| 401 | Missing or invalid API key. | Send a valid sk-ar-… key as Authorization: Bearer. |
| 402 | Out of credits or upstream billing failure. | Top up credits, or read metadata.upstream_message. |
| 404 | Unknown model or stored response. | Verify the model id, or confirm the response id is owned by your workspace. |
| 429 | Rate limit hit. | Back off; inspect X-RateLimit-* response headers. |
See Errors for the full status table and retry recipes.