SmartGate
Token Control

Hard-cap the token black hole before it caps your budget

Agent loops inflate context every iteration. Intercept at the tool layer with smart_context_gate, smart_dedup, and smart_budget_guard — then prove savings in Usage Reports.

The pain

FinOps for AI is the #1 forward-looking priority in 2026 — yet most teams price agent work like a single API call. Long tool outputs, overlapping search hits, and unbounded loops compound until the invoice arrives. Budget alerts are not budget enforcement.
1

Compress at the gateway

smart_context_gate — intelligent compression algorithm with optional purpose pre-filter. Gateway-local; not a substitute for your host LLM.

2

Dedup overlapping passages

smart_dedup — intelligent deduplication algorithm removes redundant chunks from multi-fetch or multi-search results.

3

Enforce team caps

smart_budget_guard: check before expensive jobs, record after. L1 monthly limit + L2 hard cap on Pro (β).

Before / after

ScenarioBeforeWith SmartGate
Long fetch output12,000 tokens to host LLM~4,800 after ratio 0.4 compress
Overlapping search snippets5 redundant passages2 kept via intelligent dedup
Runaway agent loopSoft alert, bill keeps growingHard cap blocks further tool calls
VisibilityNo token_saved metricDashboard Est. USD saved from audit

Configurable ratio — not a fixed 60% promise

Savings depend on input structure and ratio settings. Dashboard shows per-request compress metrics from audit — not marketing estimates.

MCP example

// Pre-flight budget check
{ "tool": "smart_budget_guard", "arguments": { "action": "check" } }

// Compress with purpose filter
{ "tool": "smart_context_gate", "arguments": {
  "text": "…",
  "ratio": 0.4,
  "purpose": "API rate limits and pricing"
}}

Start free — see Usage Reports

Playground + compress metrics on every plan

Start free — 2M tokens/month