Usage Explorer
Build cost, token, and request charts grouped by model, key, or time. Filter by date range, export, and read percentile latency.
"Which model is burning the most spend this week?" "Did latency regress after I switched keys?" "How does weekend traffic compare to weekdays?" Answering questions like these normally means building your own analytics pipeline. Usage Explorer is that pipeline, already built — an ad-hoc charting surface over your AnyRouter activity where you group, filter, and compare in a few clicks.
Overview
Open Usage Explorer at /dashboard/usage. Every view is URL-shareable — paste a link in Slack and your teammate lands on the exact same chart, filters and all.
A chart is built from four building blocks:
| Block | What it controls |
|---|---|
| Dimension | What rows are grouped by (one series per group) |
| Metric | What number goes on the y-axis |
| Granularity | The width of each bucket on the x-axis |
| Filters | Which requests are counted, applied before grouping |
How it works
Dimensions
A dimension controls how rows are grouped. Pick one at a time.
| Dimension | What it shows | Use it when |
|---|---|---|
time | One bar/line per bucket (hour, day, week) | You want a trend — "is my traffic growing" |
model | One series per model id (e.g. openai/gpt-5.4) | You want to attribute spend to specific models |
key | One series per API key | You're isolating cost or errors to a single key |
organization | One series per org or workspace | You run multi-team billing and need a P&L per team |
route | One series per API path (/chat/completions, /responses, /embeddings) | You want to see how Responses vs. Chat traffic splits |
Group-by dimensions stack: pick model while filtering to a single key, and you see model mix for that key only.
Metrics
The metric controls the y-axis. Switch metrics without losing your other filters.
| Metric | Unit | Notes |
|---|---|---|
| Cost | USD | Pass-through provider cost. No AnyRouter markup. |
| Prompt tokens | tokens | Input side — includes cached prompt tokens. |
| Completion tokens | tokens | Output side — what the model actually generated. |
| Requests | count | One row per request, regardless of size. |
| p50 latency | ms | Median end-to-end latency, server-measured. |
| p95 latency | ms | 95th percentile — how slow your slow requests are. |
| p99 latency | ms | 99th percentile — tail latency. Watch this for SLOs. |
| Error rate | percent | Non-2xx responses divided by total requests in the bucket. |
| Cache hit ratio | percent | Cached prompt tokens divided by total prompt tokens. |
Percentile latency is computed per bucket, not over the whole window. A p95 bar for "Monday 14
" is the 95th percentile of requests in that one hour, so it reacts quickly to incidents instead of being smoothed across a full day.Granularity
Granularity decides how wide each bucket is on the x-axis.
| Granularity | Range it covers well | Most useful for |
|---|---|---|
| Hour | Last 24–72 hours | Debugging a spike or incident — "what happened around 3pm" |
| Day | Last 7–60 days | Weekly cost reviews, day-over-day comparisons |
| Week | Last 60+ days | Quarterly trends, capacity planning, churn analysis |
Lower granularity (hour) is noisier but catches short events. Higher granularity (week) smooths noise out and is the right choice for a clean trend line.
Filters
Filters narrow the dataset before grouping. Every filter is AND-composed — a request must match all active filters to be counted.
- Date range — a start and end, or quick-presets (
Last 24h,Last 7d,Last 30d,Month to date). - Model — one or more model ids, for head-to-head comparisons.
- API key — one or more keys, by the
sk-ar-…prefix shown in the dashboard. - Organization / workspace — when you belong to multiple workspaces.
- Route —
/chat/completions,/responses,/embeddings,/messages. - Status — only
2xx, only errors, or both. - Cached — only cached or only uncached requests.
Every filter, dimension, metric, and granularity is encoded in the URL. A link like:
https://anyrouter.dev/dashboard/usage?granularity=day&dimension=cost&group_by=model&from=2026-04-01&to=2026-04-30&models=openai/gpt-5.4,anthropic/claude-sonnet-4.6
reproduces the exact chart. Bookmark it, drop it into an incident timeline, or send it to a teammate — no extra setup.
Comparing two keys or two models side-by-side. Set group_by=key (or group_by=model), then add both ids to the corresponding filter. The chart renders one series per group so you can read them off the same axis — the fastest way to A/B a prompt change across two keys, or to compare cost-efficiency between, say, openai/gpt-5.4 and anthropic/claude-sonnet-4.6 on identical traffic.
Reading the charts
- Bars vs. lines. Bars suit discrete buckets (cost per day); lines suit trends across many buckets (requests per hour over a week). The chart picker switches without rebuilding the query.
- Stacked vs. grouped. Grouping by
modelorkeystacks by default, so total height stays the overall metric. Switch to grouped (side-by-side) when comparing series matters more than reading the total. - Empty buckets. Zero-value buckets stay on the x-axis so traffic gaps are visible. A missing bar means no requests, not missing data.
Configure
Export to CSV
Chart-level CSV export — one row per (bucket, group) pair, matching exactly what's on screen — is coming soon.
Need the data today? The individual requests behind any chart can already be exported as CSV via the API: GET /api/v1/logs?format=csv, which accepts the same date range, model, and key filters. See Request Logs for the full filter reference.
Drill down to individual requests
Usage Explorer is the aggregate view. To see the individual requests behind a bar — "show me every failed request in this 1-hour bucket" — jump to Request Logs. Clicking a bar filters it straight into Logs.
Frequently asked questions
How is cost computed?
Cost is prompt_tokens × input_price + completion_tokens × output_price, using the per-model prices on each model detail page. Cache reads are billed at the upstream provider's cache-read rate. AnyRouter does not mark up cost — the dashboard number is what you actually pay.
What does cached mean here?
A request counts as cached when the upstream reports cached prompt tokens in usage.prompt_tokens_details.cached_tokens — either an explicit Anthropic-style breakpoint cache or OpenAI's automatic prefix caching. See Prompt Caching. The Cache hit ratio metric is cached_tokens / prompt_tokens; 0.85 means 85% of your input tokens were cache reads.
Why doesn't the chart total match my invoice?
The chart shows request-level cost billed at request time. Your invoice may bundle in adjustments (credits applied, free-tier usage, refunds) that don't appear per-request but are reflected in Credits.
My chart is empty — what's wrong?
Check three things in order: (1) Date range — the default is Last 30 days; widen it if you only started today. (2) Filters — they AND together, so a misclicked model filter is the usual culprit; clear and re-add one at a time. (3) Auth — viewing a workspace you can't access returns zero rows, not an error; check the workspace switcher.