Local Relay — Serve Models From Your Own Machine
Route requests to a model running on your own computer (Apple Foundation Models, Ollama, LM Studio) through your AnyRouter API key, and optionally donate idle capacity to the shared pool for credits.
Local relay turns a computer you own into an AnyRouter backend. A small CLI on your machine opens an outbound connection to AnyRouter; the gateway forwards matching chat-completion requests down that connection to a model server running locally — Apple's on-device Foundation Model via fm serve, Ollama, LM Studio, or anything else that speaks the OpenAI chat-completions API — and streams the response back. To your application it looks like any other model behind the same API key.
What it's for
- Apple Foundation Models. Route to Apple's on-device model straight from your Apple Silicon Mac, through the normal
apple/foundation-modelcatalog entry — no separate integration. - Any local model server. Ollama, LM Studio, llama.cpp, vLLM — if it exposes an OpenAI-compatible
/v1/chat/completionsand/v1/models, the relay can serve it. - $0 billing. The compute and the model are yours, so requests served by your own device are never billed against your credits — the same policy as BYOK.
- Donate idle capacity. Once your device is paired, you can also opt it into the shared pool: when your own capacity is idle, other users' matching requests can route to it, and you earn credits for every one served.
Quickstart
Install the CLI
anyr relay start
The first run signs you in (or reuses your existing AnyRouter key) and pairs this device automatically — no separate pairing step required.
Start your local model server
For Apple Foundation Models, run:
fm serve
This starts a loopback-only OpenAI-compatible server (http://127.0.0.1:1976). Ollama (http://localhost:11434) and LM Studio work the same way — just have them running before you start the relay.
Start the relay
anyr relay start
The relay auto-detects your local server (it checks fm serve, then the historical default port, then Ollama, in order) and opens an outbound connection to AnyRouter. Nothing needs to be exposed — your device dials out, like a browser; there's no port forwarding and no public IP.
Call it through the normal API
Use your regular sk-ar-... key and call apple/foundation-model (or your local model's id) exactly like any other model:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer $ANYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "apple/foundation-model", "messages": [{"role": "user", "content": "hi"}]}'
Donate to the shared pool
Run the relay with --pool to opt your device into the shared pool:
anyr relay start --pool
Once enabled, other users' matching requests can route to your device when your own capacity is idle. Every request it serves earns you credits, the same rebate model as donating a BYOK key. You can see how many client machines are currently online and which models they're serving on the local relay page — it's a public, aggregate-only view (no device names or user data).
Manage pool sharing anytime from Dashboard → Devices: toggle donation on or off per device, see which models it's advertising, and track requests served and credits earned.
Honest constraints
- 4096-token context. Apple's on-device model has a hard context ceiling — shared across system instructions, tool schemas, and the transcript. This is a framework limit, not a marketing number.
- Your machine has to be online. If your device is asleep or disconnected, AnyRouter detects it within seconds and fails over to your other configured upstreams for that model — requests don't hang waiting on an offline device.
- $0 for your own device. Requests served by your own paired device are never billed against your credits.
Security
- Pairing mints a one-time-shown token; only its hash is ever stored. Revoke it anytime from Dashboard → Devices — the device loses access immediately.
- The relay only forwards a single, fixed chat-completion path down the connection. It has no access to your files, shell, or anything else on your machine.
- Your device only ever dials out — there's nothing to expose to the internet and nothing for it to scan.
Frequently asked questions
Do I need a public IP or open ports?
No. The relay only makes an outbound connection, like a browser does. There's nothing to forward and nothing exposed.
Who can route requests to my device?
Only you, unless you opt into the shared pool with --pool. Even then, no one can see your device's name or identity — pool routing is anonymous on both sides, and you can turn it off instantly.
What happens if my device goes offline?
AnyRouter detects it within seconds and fails over to your other configured upstreams for that model. Your application sees a normal response from the fallback, not a timeout.
How is relay usage billed?
At $0 for your own device — the model and the compute are yours. If a pool donor's device serves your request instead (only relevant when your own device is offline and the pool has a match), normal model pricing for that model applies.
Related
- Bring Your Own Key (BYOK) — attach your own provider keys
- Share a Key to the Shared Pool — the same credit-earning model for donated provider keys
- Local relay landing page — live pool availability and quickstart