Route to a cheaper model
The largest lever is which model a call is routed to. Attach an LLM config with a smaller model, or set fallbacks so a failure retries cheaper rather than more expensive:Compress payloads
Nasiko can shrink tool results, JSON, logs, diffs and history before they are sent back into the model. Compression is structural (it keeps the shape of the data) and fails closed: if a compressor errors, the original payload is sent. Per agent. On the agent’s card in the dashboard, the token-optimization switch setscompress_enabled. It is off until you turn it on. The same field is on PUT /api/agents/{id}.
Cluster-wide defaults (self-hosting) are kill switches, not enablers. They default on so a dashboard toggle takes effect without a restart. Set one to false to stop that layer everywhere. Compression still does not run until the agent (or, for orchestrator history, every live agent you own) has compress_enabled.
A high Input share on the TokenOps attributions table usually means context is being resent. Compression and the context budget below are the two ways to cut that.
Context budget and selection
Each user has a PACMS budget tier and a context-selection strategy. They control how much chat history a request carries, not how much the cluster may spend overall.
API:
GET/PATCH /api/me/pacms-budget and GET/PATCH /api/me/context-strategy.
Provider prompt cache
When the outbound provider supports prompt caching, Nasiko records cache-read and cache-write tokens separately. Cached input is not billed as fresh input. You do not configure this; it follows the provider.What is not here
- Spend caps, alerts, chargeback, and kill switches are not a product surface.
- Flow-level token, depth, fan-out and timeout limits are — see Flow limits.
