FinOps dashboard
Understand and optimize token consumption across your Sentinel fleet
The FinOps dashboard shows how your fleet consumes tokens — where they go and how efficiently they're used — so you can spot waste and tune your AI spend. Find it in the SUPERWISE® app under Observability → Dashboards → FinOps.
Use the sentinel filter to focus on specific gateways, the time range (24H / 7D / 30D) to set the window, and the Live toggle to auto-refresh. Every metric compares against the previous period, and you can export the view as a PDF, PNG, or CSV.
Like every Sentinel dashboard, this is computed from anonymous metadata only — token counts and model names, never your prompts or responses. See Data privacy.
The dashboard has two sections.
Consumption
Total token usage for the period and the efficiency ratios behind it:
- Total consumption — all tokens consumed (prompt, cache read/write, and completion) across every provider and sentinel.
- Cache hit rate — the share of input tokens served from cache. Higher is cheaper.
- Reasoning ratio — reasoning tokens as a share of completion tokens, so you can see how much "thinking" your traffic pays for.
- Output-to-input ratio — completion tokens relative to input tokens.
- Consumption over time — token usage across the window.
Where the tokens go
The same consumption, broken down so you can find the heavy hitters:
- By provider — token usage per provider, including the cache and reasoning breakdown.
- Top models by tokens — which model versions consume the most.
- By sentinel — consumption across your gateways.
Updated about 18 hours ago
