A model that runs on your Mac
Apple ships an on-device Foundation Model as part of Apple Intelligence — a small language model that runs entirely on the local hardware, with no data leaving the device. It's now in the AnyRouter catalog as apple/foundation-model, reachable through the exact same OpenAI-compatible API you already use for every other model.
The shape is different from every other model we host, and it's worth being precise about it: AnyRouter does not run this model for you. Your own Mac does. AnyRouter is the connection between the public API and the model server running on your machine — nothing more. That's what makes it $0 and what keeps your prompts on your own hardware.
How the relay works
Apple's model has no public endpoint — it runs on a machine with no server address for us to call. So instead of AnyRouter reaching out to a provider, your Mac reaches out to AnyRouter. The anyr CLI opens an outbound WebSocket to AnyRouter and holds it open. When a request comes in for apple/foundation-model, AnyRouter pushes it down that connection to your Mac, your local model server answers, and the response streams back up the same connection to whoever made the call.
sequenceDiagram participant C as API caller participant GW as AnyRouter participant M as Your Mac (anyr relay) participant FM as On-device model M->>GW: open outbound WebSocket, stay connected C->>GW: POST /v1/chat/completions (apple/foundation-model) GW->>M: push request down the socket M->>FM: run on-device FM-->>M: tokens M-->>GW: stream response back up GW-->>C: streamed response
Because the origin is a device you own and control, this is the one place in AnyRouter where a request doesn't route through the usual gateway path — there's simply nothing on the public internet to route to. The model server on your Mac is the origin.
Set it up in two steps
You need a Mac that supports Apple Intelligence and a local server exposing the on-device model over an OpenAI-compatible endpoint (Apple's fm serve, listening on 127.0.0.1:1976). Once that's running, one command connects it to your AnyRouter account:
# 1. Start Apple's on-device model server on your Mac
fm serve
# 2. Connect it to AnyRouter (auto-detects fm serve on port 1976)
anyr relay startrelay start pairs the machine on first run — it uses an existing AnyRouter API key if it finds one, or walks you through a browser login otherwise — then holds the connection open and reconnects on its own if your network drops. Leave it running and your Mac is on call to serve requests for apple/foundation-model.
Call it like any other model
With the relay connected, apple/foundation-model behaves like every other id in the catalog. Point a standard chat-completions request at it:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer $ANYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "apple/foundation-model",
"messages": [{ "role": "user", "content": "Summarize this in one sentence." }]
}'Or pick apple/foundation-model from the model dropdown in the Playground at playground.anyrouter.dev and chat with it in the browser. Streaming works the same way it does for any other model. If no paired device is online when the request lands, the call fails fast so a fallback model can answer instead — nothing hangs waiting on an offline Mac.
What to expect
Two honest caveats. First, the on-device model is small and built for on-device work — Apple caps its context window at 4,096 tokens, shared across your system instructions, any tool schemas, and the conversation. It's a genuine framework limit, not a setting we can raise. Treat it as a fast local model for focused tasks, not a long-context workhorse.
Second, throughput depends entirely on your Mac — the model runs on your silicon, so your hardware sets the pace. We're not going to quote latency numbers here, because the only ones that matter are the ones you measure on your own machine.
- Private by design. Prompts and completions run on your device and are never sent to a third-party model vendor.
- $0 per token. You supply the hardware, so there's nothing to bill — input and output are both priced at zero.
- No vendor account. There's no API key to obtain from Apple; the model is part of Apple Intelligence on your Mac.
Route your first request in 2 minutes
Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.
Start free