Token Control
Hard-cap the token black hole before it caps your budget
Agent loops inflate context every iteration. Intercept at the tool layer with smart_context_gate, smart_dedup, and smart_budget_guard — then prove savings in Usage Reports.
The pain
FinOps for AI is the #1 forward-looking priority in 2026 — yet most teams price agent work like a single API call. Long tool outputs, overlapping search hits, and unbounded loops compound until the invoice arrives. Budget alerts are not budget enforcement.
1
Compress at the gateway
smart_context_gate — intelligent compression algorithm with optional purpose pre-filter. Gateway-local; not a substitute for your host LLM.
2
Dedup overlapping passages
smart_dedup — intelligent deduplication algorithm removes redundant chunks from multi-fetch or multi-search results.
3
Enforce team caps
smart_budget_guard: check before expensive jobs, record after. L1 monthly limit + L2 hard cap on Pro (β).
Before / after
| Scenario | Before | With SmartGate |
|---|---|---|
| Long fetch output | 12,000 tokens to host LLM | ~4,800 after ratio 0.4 compress |
| Overlapping search snippets | 5 redundant passages | 2 kept via intelligent dedup |
| Runaway agent loop | Soft alert, bill keeps growing | Hard cap blocks further tool calls |
| Visibility | No token_saved metric | Dashboard Est. USD saved from audit |
Configurable ratio — not a fixed 60% promise
Savings depend on input structure and ratio settings. Dashboard shows per-request compress metrics from audit — not marketing estimates.
MCP example
// Pre-flight budget check
{ "tool": "smart_budget_guard", "arguments": { "action": "check" } }
// Compress with purpose filter
{ "tool": "smart_context_gate", "arguments": {
"text": "…",
"ratio": 0.4,
"purpose": "API rate limits and pricing"
}}Start free — see Usage Reports
Playground + compress metrics on every plan
Start free — 2M tokens/month