Skip to content

Responses API

/api/v1/responses — the OpenAI Responses-compatible lifecycle. Send input, stream or retrieve responses, continue state, and inspect input items.

/api/v1/responses implements the OpenAI Responses API lifecycle. Unlike Chat Completions, the response is a self-contained envelope with typed output items — text blocks, tool calls, and reasoning summaries — rather than a flat choices array.

POST/api/v1/responses
GET/api/v1/responses/{response_id}
DELETE/api/v1/responses/{response_id}
GET/api/v1/responses/{response_id}/input_items
POST/api/v1/responses/input_tokens
POST/api/v1/responses/{response_id}/cancel
POST/api/v1/responses/compact

Authenticate with an LLM API key (sk-ar-v1-…) from /dashboard/keys. Management keys (ak_…) are not accepted on inference routes.

Request

FieldTypeRequiredDescription
modelstringyesprovider/model id. See List models.
inputstring | arrayyesUser input. String or array of input items (text, image, etc.).
instructionsstringnoSystem-level instruction passed before the input.
max_output_tokensintegernoToken budget for the response.
temperaturenumberno0–2. Sampling temperature.
top_pnumberno0–1. Nucleus sampling.
streambooleannoIf true, returns an SSE stream of Responses events.
storebooleannoDefaults to true. Stored responses can be retrieved, deleted, continued with previous_response_id, and inspected with /input_items.
previous_response_idstringnoContinue from a stored response owned by the same workspace.
backgroundbooleannoIf true, creates a queued stored response that can be retrieved or cancelled.
toolsarraynoFunction definitions the model can call. Each tool has type, name, description, parameters.
tool_choicestring | objectno"auto", "none", "required", or {"type":"function","name":"…"}.
reasoningobjectno{"effort": "low"|"medium"|"high", "summary": "auto"|"detailed"}.
modalitiesstring[]noOutput modalities requested (e.g. ["text"]).
metadataobjectnoArbitrary key-value pairs attached to the response.
session_idstringnoSticky-session id (≤256 chars) that groups related requests in Request Logs. Also settable via the x-session-id header; the body field wins.
providerobjectnoProvider routing preferences. Supports only, ignore, order, sort, allow_fallbacks, max_price, preferred_max_latency, and preferred_min_throughput.
traceobjectnoOptional tracing metadata. Supports trace_id, trace_name, span_name, generation_name, and parent_span_id.

For the full field list see /docs/api.

Stored response lifecycle

Responses are stored by default for later retrieval. Pass store: false when you do not need lifecycle operations. Stored responses are scoped to the API key workspace; another workspace cannot retrieve, delete, continue, or cancel them.

# Retrieve a stored response
curl https://anyrouter.dev/api/v1/responses/resp_abc123 \
  -H "Authorization: Bearer sk-ar-your-key"

# List its input items
curl https://anyrouter.dev/api/v1/responses/resp_abc123/input_items \
  -H "Authorization: Bearer sk-ar-your-key"

Response

{
  "id": "resp_…",
  "object": "response",
  "created_at": 1714000000,
  "completed_at": 1714000001,
  "model": "z-ai/glm-4.7-flash",
  "status": "completed",
  "output": [ /* typed output items */ ],
  "output_text": "Paris.",
  "usage": {
    "input_tokens": 14,
    "output_tokens": 3,
    "total_tokens": 17
  },
  "incomplete_details": null,
  "error": null
}

status is one of "in_progress", "completed", "incomplete", or "failed". When status is "incomplete", incomplete_details.reason is "max_output_tokens" or "content_filter".

Output item types

TypeDescription
messageAssistant turn with a content array of output_text blocks.
function_callTool invocation with call_id, name, arguments.
reasoningReasoning summary (provider-dependent; present when reasoning tokens are surfaced).

Responses-native upstreams can also pass through server tool items such as web_search_call, file_search_call, code_interpreter_call, and computer_call. Chat Completions and Anthropic Messages upstreams currently normalize only text, reasoning, and function-call output items.

Streaming events

Set stream: true to receive server-sent events. The event sequence is:

event: response.created
data: {"type":"response.created","response":{...}}

event: response.in_progress
data: {"type":"response.in_progress","response":{...}}

event: response.output_item.added
data: {"type":"response.output_item.added","output_index":0,"item":{...}}

event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_xyz","output_index":0,"content_index":0,"delta":"Par"}

event: response.output_text.done
data: {"type":"response.output_text.done","item_id":"msg_xyz","output_index":0,"content_index":0,"text":"Paris."}

event: response.output_item.done
data: {"type":"response.output_item.done","output_index":0,"item":{...}}

event: response.completed
data: {"type":"response.completed","response":{...}}

event: response.anyrouter.metadata
data: {"type":"response.anyrouter.metadata","anyrouter_metadata":{...}}

The response.anyrouter.metadata frame is emitted after response.completed and carries billing and upstream info. Standard Responses API clients safely ignore unknown event types.

Examples

curl https://anyrouter.dev/api/v1/responses \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-4.7-flash",
    "input": "What is the capital of France?"
  }'
curl https://anyrouter.dev/api/v1/responses \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-4.7-flash",
    "input": "What is the weather in Paris?",
    "tools": [
      {
        "type": "function",
        "name": "get_weather",
        "description": "Get current weather for a location",
        "parameters": {
          "type": "object",
          "properties": { "location": { "type": "string" } },
          "required": ["location"]
        }
      }
    ],
    "tool_choice": "auto"
  }'
curl https://anyrouter.dev/api/v1/responses \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/o1-mini",
    "input": "Solve: x^2 - 5x + 6 = 0",
    "reasoning": { "effort": "medium", "summary": "auto" }
  }'
curl https://anyrouter.dev/api/v1/responses \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "input": "Summarize the deployment notes",
    "provider": {
      "only": ["OpenAI", "Groq"],
      "sort": "latency",
      "allow_fallbacks": true,
      "preferred_max_latency": 250
    },
    "trace": {
      "trace_id": "4f8c2b9670b44b49a2e71e7fd5c0b1d3",
      "trace_name": "release-pipeline"
    }
  }'

The minimal request above returns a stored response envelope:

{
  "id": "resp_abc123",
  "object": "response",
  "status": "completed",
  "model": "z-ai/glm-4.7-flash",
  "output": [
    {
      "type": "message",
      "id": "msg_xyz",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Paris." }],
      "status": "completed"
    }
  ],
  "usage": {
    "input_tokens": 14,
    "output_tokens": 3,
    "total_tokens": 17,
    "cost": 0.0000021,
    "cost_details": {
      "upstream_inference_input_cost": 0.0000014,
      "upstream_inference_output_cost": 0.0000007
    },
    "is_byok": false
  },
  "anyrouter_metadata": {
    "requestId": "resp_abc123",
    "model": "z-ai/glm-4.7-flash",
    "upstream": { "provider": "openai_direct", "backend": "openai_direct", "latencyMs": 412 }
  }
}

When you send a reasoning request, the output array contains a reasoning item before the message item, provided the model exposes reasoning tokens.

Provider routing behavior

The provider object is applied before AnyRouter dispatches upstream requests:

  • only and ignore shape the candidate set.
  • order pins the preferred backend order.
  • sort reorders candidates by price, latency, or throughput.
  • max_price.prompt and max_price.completion remove candidates above your budget.
  • preferred_max_latency and preferred_min_throughput apply performance cutoffs.
  • allow_fallbacks: false means AnyRouter tries only the top candidate.

See Smart Routing for the full behavior.

Errors

All errors follow the standard AnyRouter JSON envelope with a stable error.code.

StatusMeaningFix
400Malformed body or invalid parameters.Check the request against the field table.
401Missing or invalid API key.Send a valid sk-ar-… key as Authorization: Bearer.
402Out of credits or upstream billing failure.Top up credits, or read metadata.upstream_message.
404Unknown model or stored response.Verify the model id, or confirm the response id is owned by your workspace.
429Rate limit hit.Back off; inspect X-RateLimit-* response headers.

See Errors for the full status table and retry recipes.

Related