Skip to main content
Two dashboard screens read the same TokenOps data. Overview (/) is the landing screen: fleet health, activity and spend on one page. TokenOps (/tokenops) is the cost view: what you spent, where it concentrated, and which agents or workflows drove it. Both screens take a period and a time range, compare every headline figure against the previous window, and export the attribution table as CSV.

Filters both screens share

Enterprise. The Org unit filter needs an organization hierarchy and manager-level access or above. On open-source clusters, or for users without that access, it stays disabled with a tooltip explaining why. See organizations.
Every KPI carries a movement chip: the change against the previous window of the same length. The arrow shows the direction; the color shows whether that direction is good news. Rising spend is an up arrow in a warning color.

Overview

The Overview screen answers “is the fleet healthy, and what is it costing?” in one place.

Filter bar

Alongside Period, Time range and Org unit, Overview adds a Scope control:
  • Org (default) counts every agent you can see.
  • Myself narrows the KPI strip and the attributions table to agents you own. The activity, performance and spend charts stay fleet-wide, and the screen says so under the filter bar while Myself is selected.
Myself means agents you own, not traffic you generated. If you mostly use agents other people deployed, Myself reads close to zero.

KPI strip

The gap between Active and Agents is often the most useful number on the screen: it counts agents that are deployed but idle.

Panels

  • Agent activity plots calls over time. Switch between All, Agent calls (one per trace) and Tool calls (tool-call spans inside those traces).
  • Performance plots p50, p95 and p99 latency over time. When the gap between p95 and p99 moves, a note under the chart reports it, for example “Latency tail widened 74% since Wednesday”.
  • Spend over time stacks each bucket’s heaviest agent against everyone else. Toggle $ for dollars or % for each bucket’s share of the window. The Agent and Provider selects in this panel filter this panel only.
  • Attributions lists every agent with Runs, Tokens, Spend, Cost/Op and Avg latency. Hover Runs for tool-call counts and Avg latency for p95/p99. Sort by Most run, Highest spend, Most tokens, Slowest (p50), Slowest (p95) or Name.
Buckets are hourly for short windows and daily for longer ones. With no agents deployed, Overview shows a first-run screen with Import an agent and Explore Artifact Library instead of empty charts.

TokenOps

The TokenOps screen is the cost view. Every filter on it reloads the whole screen.

Filter bar

Alongside Period, Time range and Org unit: Picking a month makes the month the window. Picking a time range after that returns to a range anchored on now.

KPI strip

A tile with no previous window to compare against shows no chip, rather than a misleading 0%.
An estimated figure is a real cost priced at a fallback rate, not a guess at usage. See how prices are determined for when a call is marked estimated.

Spend over time

Spend (left axis) and operation count (right axis) per bucket. Buckets are hourly for short windows and daily for longer ones.

Spend concentration

A day grid for the selected month sits above an hourly chart for one day. Click a day to drill into it.
  • Each hour’s bar is stacked by that day’s top four agents, plus Others.
  • The dashed line is the day’s average hourly spend, also written under the chart (“Averaging $3.20/hour on 2026-09-14”).
  • The legend lists the top agents with their spend for the day.
The day picker follows the Period select: a past month opens on its last day, the current month on today. It ignores Time range, but honors the Agent, Provider and Model filters. Future days are disabled.

Attributions

Toggle Agent or Workflow to attribute spend to deployed agents or to MAF workflows. Search by name, and sort by Highest spend, Most tokens, Most operations, Slowest, Most agent hours (agent view) or Name. A ~approx badge marks a very high-volume agent whose figures are approximate and may undercount. The split between Input and Output is the practical cost lever. A high input share usually means context is being resent; see reduce cost. A high output share usually means verbose answers, which a cheaper model or a brevity directive can address.

Export

Export report downloads the attributions table as CSV: overview.csv from Overview, tokenops.csv from TokenOps. The file follows the current filters, view and sort order. On TokenOps it includes every row in the window, not only rows matching the table’s search box.

Where the numbers come from

Every model call routed through the LLM router, every orchestrator call, and every reporting coding agent session records its token counts. Nothing is added to your agent code. Costs apply those counts to the price in effect when the call happened, so a later price change never re-costs last month.
  • Cost attribution explains the dimensions and why every call is attributed.
  • Pricing explains where rates come from.
The routing engine’s agent-selection call is metered as its own line, not folded into the agent it picked, so routing never inflates an agent’s numbers.
For a short written summary instead of raw figures, POST /api/observability/finops/insights returns up to three insight bullets over a KPI snapshot you send it. Useful for a weekly digest. See the API table.

Cost attribution

One cost schema across vendors, harnesses and teams.

Reduce cost

Compression, context budgets and cheaper models.

Sessions and traces

Drill from a cost figure into the traces behind it.

LLM router

Change which models your agents are routed to.