SmartGateSmartGate

Industries

A topic library for MCP, agent tool traffic, context engineering, RAG, and agent governance — guides aligned with SmartGate as an MCP-native intelligence layer.

These pages are a topic library for builders and operators who work with Model Context Protocol traffic, agent tool loops, retrieval patterns, and gateway-side controls. They are not vertical compliance packs for banks or clinics, and they do not claim sector certifications. Use them to go deeper on the same themes as SmartGate Features (Token Control, Traffic Customs, Audit) and Docs — then wire the runtime over MCP when you are ready.

MCP

Canonical guides in this cluster — request indexing on these hubs first when GSC quota is limited.

What Is the Model Context Protocol? A First-Principles Guide

The Model Context Protocol is the wire format an AI host uses to reach tools it never shipped with. Here is the problem it solves, its three actors, one request end to end, and where it stops.

MCP Server Explained: Endpoint, Tool Registry, Sessions

An MCP server answers initialize, tools/list, and tools/call at one Streamable HTTP endpoint. What an MCP server is, when a session matters, and how to run one.

MCP Server List: A Buyer's Checklist Before You Install

A buyer's checklist for evaluating an MCP server before you install it: the auth model, the tool surface, the limits, the audit trail and where the counters live.

MCP Server Example: One Tool, from Config to First Call

A worked MCP server example end to end: the config entry a host reads, the tool list the server publishes, the first call, the record that proves it arrived, and the read path behind it.

What Is an MCP Gateway? Definition, Differences and Limits

An MCP gateway is the policy layer between an agent host and the MCP server: one upstream URL, per-key and per-team limits, plan-clamped budget caps and an audit row per tool call.

MCP Proxy vs Router vs Gateway: Which Layer Decides What

An MCP proxy moves bytes, a router picks an upstream, a gateway decides whether the call happens. Which layer owns each decision, and when one proxy is the whole stack.

MCP vs REST API: Choosing the Right Agent Interface

MCP and REST/OpenAPI answer different questions for agent tool access. Compare discovery, sessions, credentials and limits, and see what a gateway does for either surface.

MCP vs A2A Protocol: Vertical Tools, Horizontal Agents

MCP wires one agent to tools; A2A wires agents to each other. What each protocol carries, where the two boundaries meet, and where a gateway sits when both are in play.

Model Context Protocol (MCP) Message Format Explained

A Model Context Protocol message is JSON-RPC 2.0 with three shapes — requests, notifications, and results. Here is how initialize, tools/list, and tools/call actually travel over Streamable HTTP.

MCP Protocol Versions and Transports: HTTP, SSE, and stdio

How the MCP protocol negotiates versions and capabilities, which transport to use — Streamable HTTP, legacy SSE, or stdio — and what a gateway has to patch for older clients.

MCP Tools Reference: tools/list Schemas and Annotations

A working reference for the MCP tools list: the schema and annotations tools/list returns, what read-only hints mean for hosts, and how a gateway normalizes both before they reach a client.

MCP Specification Walkthrough: How to Read the Spec

A practitioner's walkthrough of the MCP specification: which chapters are normative, how initialize negotiates a version, the error shapes nobody defines, and the limits.

MCP Inspector Alternatives: Where to Debug a Live Server

The MCP Inspector is one of five places to look when a tool call fails: a session trace, server logs, an audit row, a health check and the client's own pane. How to pick.

MCP OAuth and Authorization: Key Hashes, Bearer, Roles

MCP OAuth defines the shape of an authorised request; the gateway owns the decisions. How a key is issued, hashed, validated, scoped to a team and capped.

MCP Resources, Prompts and Sampling Through a Gateway

MCP resources, prompts and sampling sit beside tools. What a gateway can serve, meter and govern for each primitive, and the one direction it cannot bill.

Python MCP Server Tutorial: Build It, Then Govern It

Build an MCP server in Python with the official SDK: one tool, one resource, a local smoke test, then Streamable HTTP — and the keys, limits and audit rows the traffic needs.

Build an MCP Client for AI Agents: Config, Auth, Transport

How to make an MCP client for AI agents actually connect: the per-platform config shapes each client demands, the two headers a gateway reads, the single upstream URL.

Anthropic MCP: Claude Clients, Servers, and the Gateway Between

Anthropic MCP is the Claude side of the Model Context Protocol: which clients speak it, how the transport moved to Streamable HTTP, and what a gateway adds in front.

Model Context Protocol Example: A JSON-RPC Round Trip

A Model Context Protocol example, message by message: the initialize handshake, a tools/call round trip, and the JSON-RPC error body a client has to read.

Enterprise AI Gateway Architecture Best Practices: MCP Gateway

An enterprise AI gateway is enforced by the layer around it: tenancy, per-request identity, canonicalized inputs, runtime config, scoped exclusions, validated parameters, and typed plan gates.

How to Deploy an AI Gateway in a Private Cloud: MCP and Audit

Deploy an AI gateway in your own cloud: the MCP endpoint you expose, the base URL that must fail loudly, the server list you version, and the audit window you operate.

RAG

Canonical guides in this cluster — request indexing on these hubs first when GSC quota is limited.

What Is Agentic RAG? Retrieval as a Loop, Not a Pipeline

Agentic RAG puts retrieval inside an agent's loop: the model searches, judges what came back, and searches again. What changes, what it costs, and when it is worth it.

Agentic RAG Survey: How the Research Splits the Field

An agentic RAG survey in practice: the taxonomies the papers use, where search-as-a-tool, self-reflection and evaluation benchmarks diverge, and what to take from each.

RAG vs Agentic RAG: Choosing per Workload, Not in General

RAG vs agentic RAG is a per-workload decision, not a preference: latency, cost, answer shape and the failure mode of each path, on one table you can decide from.

RAG vs Agentic AI: One Names a Pattern, the Other a System

RAG names a retrieval pattern with no side effects; agentic AI names a runtime that plans, calls tools and acts. Where the line sits, and what each term is used for.

Agentic Search vs RAG: A Tool Call or an Index You Own

An agent's search tool reads the live web and someone else owns the ranking; a RAG index reads a corpus you built. Freshness, cost per call, and how to pick the path.

What Is RAG Architecture? The Decisions You Cannot Defer

RAG architecture is a set of design decisions: which components you can swap, how each swap trades latency against answer quality, and the failure every choice invites.

RAG Architecture Diagram: the Boxes and What Each Arrow Costs

A RAG architecture diagram in words: the index, retriever, reranker, context builder and loop controller, and what every arrow between them costs in tokens and latency.

Advanced RAG Architecture for AI Agents: Four Stages to Own

An advanced RAG architecture for AI agents is four stages you own end to end: a rebuildable index, dedup before and after retrieval, compression with a reported ratio, and a search path with a floor.

Agentic Retrieval: The Loop That Decides When to Search

Agentic retrieval is the loop an agent runs when one search is not enough: decide whether to search, plan the query, call a tool, judge the result, and stop. How each move fails.

Traffic / Customs

Canonical guides in this cluster — request indexing on these hubs first when GSC quota is limited.

URL to Markdown: What a Model-Ready Page Has to Keep

URL to markdown is the extraction step between a web page and a model: choosing the main content, keeping tables and code, and fitting the result into a token budget.

Firecrawl Alternative: When to Replace a Crawl Pipeline

A Firecrawl alternative is a replacement decision, not a shortlist: five criteria — rendering, extraction quality, limits, record and data handling — decide when a crawl pipeline moves.

Agent Pipeline: What Runs, In What Order, and What Breaks

An agent pipeline is a fixed chain of steps that each receive one input and return one artefact. What belongs in a step, what a retry costs, and the four fields that make a failed run explainable.

What Is an Agentic Workflow? Loop, Pipeline and Orchestration

An agentic workflow is a loop with a policy: steps, state carried between them, a stopping rule and a record of what ran. Where pipelines beat agents, and what each costs.

Agentic Workflow Examples: Six Patterns You Can Run

Six agentic workflow examples with input, steps, failure handling and an acceptance signal each: research to decision, batch enrichment, review loop, fan-out, routing, monitoring.

AI Workflow Automation: Where the Judgement Step Belongs

AI workflow automation is a repeated process with the judgement step made explicit. Which processes are worth automating, what each shape costs, and the four records to keep.

AI Workflow Automation Tool: A Selection Scorecard

An AI workflow automation tool is chosen on evidence, not features: five gates a candidate must pass, six criteria to score, three trial scripts, and the disqualifiers.

AI Workflow Automation Platform: The Multi-Team Layer

An AI workflow automation platform is what a single-builder tool becomes once several teams share it: quotas, permissions, audit rows and cost allocation in one control plane.

AI Agent Orchestration: Routing, Handoff and Failure Modes

AI agent orchestration is the dispatcher layer: which worker runs next, what a handoff carries, when a run falls back, and the failure modes to design out before they bill you.

LLMOps: The Four Layers an Agent Stack Must Cover

LLMOps is the operating layer for agents in production: tracing, evaluation, orchestration and memory. Which tool class decides what, what to run yourself, and the order to buy in.

Token / Context

Canonical guides in this cluster — request indexing on these hubs first when GSC quota is limited.

How to Enforce a Token Quota Per Team in an AI Gateway

A per-team token quota takes four shipped pieces: the entitlement read, a pre-flight counter, a refusal your caller honours, and an alert at 80 and 100 percent.

Context Window Management Techniques for AI Agents

Context window management for agents in twelve real code paths: compression ratios, segmenting long input, pipeline truncation, semantic dedup with MMR reranking, team memory.

Context Compression: Four Methods and the Fidelity Each Costs

Context compression is four operations, not one: summarization, extraction, key-value structuring and de-duplication. What each loses in fidelity, and how to measure the ratio against task success.

Context Engineering: Four Mechanisms That Fill a Window

Context engineering decides what a model sees on every call. Four mechanisms fill the window — retrieval, compression, memory and tool return — and each one has its own failure mode.

LLM Memory: The Boundary Between Context and a Store

LLM memory is the boundary between a context window that dies with the call and a store that outlives it: what to persist, how long to keep it, and who may read it.

Memory Agent: When to Write, Read, and Forget

A memory agent is a write policy plus a read contract. When a write is triggered or refused, what one read returns, how conflicts and expiry settle, and where the audit boundary sits.

Agent Memory Architecture: What to Store and Retrieve

Agent memory architecture is a write path, a read path and a bill. What gets stored, when retrieval is allowed to run, and how a gateway exposes it as one tool.

LLM Gateways: What the Category Is and Where Its Edge Sits

An LLM gateway holds the provider credential, exposes one request surface, enforces the limit before the call and keeps one record per call. Where the category ends.

LiteLLM Alternative: A Migration Decision, Not a Shortlist

A LiteLLM alternative is a migration decision, not a shortlist: five criteria — protocol, credentials, record, limits, carry-over — decide when to switch and what you must move.

Cloudflare AI Gateway: A Managed Gateway, Dissected

Cloudflare AI Gateway is a managed control plane on Cloudflare's own network: it proxies model calls, caches responses, enforces limits and keeps one log per request. What it governs.

Databricks AI Gateway: Platform-Native or a Separate Layer?

Unity Gateway, the renamed Databricks AI Gateway, governs model and agent traffic through Unity Catalog. When the platform's permission model is the right control plane, and when to split.

Bifrost AI Gateway: The Self-Hosted Shape, and When It Pays

A Bifrost AI gateway is the lightweight, self-hosted form of the category. Four properties decide whether to run one: topology, fallback, key custody and observability.

Audit / Governance

Canonical guides in this cluster — request indexing on these hubs first when GSC quota is limited.

AI Compliance: The Evidence Trail an Auditor Expects

AI compliance is evidence work: the fields an audit row must carry, how to pick retention against the EU AI Act six-month floor, and how to hand the record over.

AI Compliance Framework: Regulation to Evidence Calendar

An AI compliance framework is a working map: each applicable rule turned into a role, an owner, an evidence artefact and a delivery date, plus the review that keeps the map honest.

AI Compliance Certification: Materials, Windows, Findings

AI compliance certification is an audit of a management system: the material a certification body asks for, how to schedule the audit window, and the findings that recur.

AI Agent Governance: Who Can Call What, Where Limits Live

AI agent governance is a control plane: which credential may call which tool, which layer enforces the rate limit and the token budget, and how over-reach is detected.

AI Governance Framework: Five Layers, Owners, Cadence

An AI governance framework is five layers of artefacts: policy and roles, risk tiering, technical controls, evidence and audit, and review — what each produces, who signs it, and when it is reviewed.

AI Governance Certification: The Evidence Chain Auditors Read

An AI governance certification is evidence about a management system, not a verdict on a model. What each claim rests on, which population an assessor samples, and how to keep the proof current.

AI Agent Security: The Runtime Surface Around Every Tool Call

AI agent security is a runtime problem: which tools an agent may call, which MCP servers it trusts, where credentials sit, and what the audit record can prove.

LLM Observability: What to Record on Every Tool Call

Agent observability needs more than prompt tracing: the record a tool call leaves behind, the queries that reconstruct a task, and where prompt-level LLM observability stops.

LLM Observability Tools: How to Evaluate and Shortlist

Shortlist LLM observability tools on four axes: sampling rate and the error tail, the cost model, the lookback window, and whether the tool can feed a per-team quota.

MCP Logging and Observability: Audit Rows, Retention

MCP logging means one audit row per tool call. What the gateway records, what it masks, how long each plan keeps it, and what those rows cannot tell you.

Other

AI Research Tool Selection: A Buy-Versus-Build Framework

Choosing an AI research tool is a class decision, not a feature comparison: the five tool classes, the eight evaluation dimensions, when to buy versus build, and the real total cost.

The AI Researcher Workflow: What to Automate, What to Keep

An ai researcher runs seven steps from question to citation. Here is which of them a model can own, which must stay human, and how a hallucinated citation slips into a draft.

AI Research Paper: The Five Stages and the Three Accidents

An AI research paper stands or falls on two human-checked artefacts: the reference list and the statistics. Where AI helps across five stages, and where fabricated citations creep in.

AI Literature Review: The Systematic Review Pipeline

An AI literature review is a systematic review with four artefacts: a protocol, a reproducible search string, a PRISMA flow and an evidence table. Here is what AI runs and what a human must own.

Deep Research AI: The Stages Inside the Loop

Deep research AI is a staged loop: plan, retrieve in rounds, read pages, synthesise, verify citations, then retrieve again. What each stage does, why it costs more than one-shot QA, and when it pays.

Deep Research Agent Architecture: Planner, Searcher, Verifier

A deep research agent runs one question through a planner, a searcher and a verifier. How to set stop conditions and budgets, and catch retrieval collapse, citation drift and self-confirmation.

Deep Research API: Wiring Research Into Your Own Product

How to integrate a deep research API into a product: asynchronous job design, task status and callbacks, quotas and rate limits, caching, retries, degradation, cost control and data retention.

What Is Deep Research? Definition, Scope, and Trust

Deep research is a multi-step AI workflow that plans, reads live sources and cites what it used before it writes. The definition, the three boundaries, and how to judge a report.

AI Research Agent: Where Automation Should Stop

An ai research agent can search, fetch and draft unattended, but a conclusion a reader will repeat still needs a human. Where to put the review gate and what to log.

Claude MCP Server Setup: Config, Scopes, and Transport

How to add and operate an MCP server in Claude: where the config lives, how local, project and user scopes differ, stdio versus HTTP, and the order to debug a failure.

AI Deep Research: Commission, Check, and Keep a Run

AI deep research turns one question into a cited report after several rounds. How to commission a run with a brief, accept the report against a check, and keep the record.

Deep Research on GitHub: Open-Source Implementations

Open-source deep research lives on GitHub. The repo families worth reading, how to vet one before you clone it, and the licence, model and self-host trade-offs that decide whether it fits.

What Is an MCP Gateway? Roles, Boundaries, and When You Need One

An MCP gateway is one authenticated address in front of your MCP servers: the four jobs it takes over in the agent request chain, its boundary with the servers, and when one is worth adding.

AWS MCP Gateway: Three Shapes AWS Ships for MCP

An AWS MCP gateway is three different things: Amazon Bedrock AgentCore Gateway, the managed AWS MCP Server, and a self-hosted MCP server on Lambda, ECS or EC2. This page says which fits when.

Microsoft MCP Gateway: Four Surfaces in One Name

Microsoft MCP gateway is four surfaces, not one product: API Management MCP servers, the AI Gateway tier, Foundry tool governance, and an open-source proxy on Kubernetes.

Composio MCP Gateway: Inside a Managed Tool-Router

A managed MCP tool-router hosts tool connections, credential custody and per-user authorization, so agents reach apps without per-app auth code. Here is the control you hand to the vendor.

Kong MCP Gateway: Reusing Your API Gateway for MCP Traffic

A Kong MCP gateway runs MCP traffic through the API gateway plugin chain you already operate. When reusing an existing gateway for MCP tool calls pays off, and when it does not.

TrueFoundry MCP Gateway: An Enterprise Control Plane View

An enterprise TrueFoundry MCP gateway is bought for its control plane, not its connector list: policy distribution, in-path enforcement, cross-environment consistency and audit.

LiteLLM MCP Gateway: What a Self-Hosted Build Owns

LiteLLM's proxy ships an MCP gateway you can self-host: one endpoint for MCP tools, access by key or team. Here is what deployment, quotas, records and upgrades then become yours to run.

AI Token Cost: From a Price per Million to a Real Bill

AI token cost is a unit rate plus a billing rule: input, output and cached-input tokens are priced separately, quoted per million, and converted into a per-request and per-session bill.

Inference Cost: GPU Hours, Batching and the API Crossover

Inference cost is a compute bill, not a rate card: GPU hours and utilization, batching and the prefill/decode split, and the point where a self-hosted model crosses a managed API.

AI Token Usage: How the Number Is Counted and Attributed

AI token usage is what a meter counts and who it is charged to: what belongs in the prompt, why the tokenizer shifts the count, and how to reconcile self-reported against provider-reported usage.

AI Model Pricing: How to Read a Rate Card, Not a Bill

AI model pricing is something you read, not compute: separate input, cached-input and output rates, batch and long-context tiers, and when two cards cannot be compared.

What Is Prompt Injection? Attack Surface and Defenses

Prompt injection is untrusted text a language model reads as an instruction. This page defines the problem, splits direct from indirect injection, and gives the four-layer defense framework.

Indirect Prompt Injection: the Path from Content to Action

Indirect prompt injection is untrusted content an agent reads as an instruction. This page follows the path from fetched page or email to a real action, and why least privilege is the defence.

AI Guardrails: Filters, Policy Engines and Classifiers

AI guardrails are input filters, output filters, a policy engine and classifier checks around a model. Where each runs, what it costs, and why detection is probabilistic.

AI Jailbreak and Prompt Injection: Attack Modes and Mitigation

An AI jailbreak talks a model out of its own policy; prompt injection hijacks the instruction channel. This page maps the main jailbreak families and why the risk only narrows.

AI Red Teaming: From Threat Model to Regression Case

AI red teaming is a method, not a scan: threat model the agent, grow an attack library, probe it automatically, score the rate, and turn each finding into a regression case.

LLM Security for Platform Teams: Keys, Logs, and Controls

LLM security from the platform side: key and credential custody, what logs may contain, gateway rate limits, quotas and audit trails, data boundaries, and dependency hygiene.

OWASP LLM Top 10: How to Read It and What to Fix First

The OWASP LLM Top 10 is a risk index, not a checklist. How to read all ten entries, which kind of control each one answers, and the three a small team should fix first.

Agent Sandboxing: Execution Isolation for Tool-Using Agents

Agent sandboxing is the execution boundary around a tool-using agent: a container, microVM or user-space kernel that runs tool-authored code with its own filesystem, network and least privilege.

AI Security Frameworks: What Actually Changes Your Build

NIST AI RMF, ISO/IEC 42001 and the EU AI Act, mapped to the AI security controls you can test: log retention, evaluation cadence, deployment gates and supplier evidence.

Agent Identity and Authorization: NHI Design for AI Agents

Agent identity is the non-human identity an autonomous workload authenticates with. This page covers credential issuance, rotation, least-privilege scopes, audit attribution and agent-to-agent trust.

Model Supply Chain Security: Provenance, Poisoning, Integrity

A model supply chain is the weights, data, adapters and tool definitions an app trusts. How to pin versions, verify integrity and defend against poisoning.

AI Agent Observability: Steps, Cost, and Failure Traces

AI agent observability answers three questions about one run: what happened, where the money went, and which step failed. Steps, retries, cost attribution, latency and privacy.

LLM Pricing: A Taxonomy of How Providers Charge

LLM pricing is a family of billing mechanisms, not one rate: per-token input and output, cache and batch tiers, subscription seats, and gateway markups layered on top.

Claude Pricing: Seats, API Tokens, and the Cache Asymmetry

Claude pricing has two entry points for the same model: a monthly subscription seat and a per-token API key. Here is how they differ, and how a cache write reprices a workload.

OpenAI API Pricing: Generations, Entries and Units

OpenAI API pricing is one card read across generations and entry points: chat, responses, batch and cached input each carry a rate, and per-million versus per-token is one number in two units.

Azure OpenAI Pricing: Managed vs Direct, Two Shapes of Cost

Azure OpenAI pricing is a managed service, not one rate: pay-as-you-go tokens, hourly provisioned capacity, regional deployment types and subscription quota all move the number.

DeepSeek Pricing: The Shape of a Low-Price Tier

DeepSeek pricing is a two-rate cache, an off-peak window and per-family tiers. This page maps the shape of a low-price tier and the workload shapes where cheap stops holding.

OpenRouter Pricing: Three Components Behind One Model

OpenRouter pricing is three stacked components: a pass-through model rate, a platform fee on credits or usage, and whether you bring your own key. What makes one model show two prices.

Groq Pricing: Paying for Tokens at the Speed They Arrive

Groq pricing bills per token but sells tokens per second: how speed tiers, batch rates and rate limits fold into the unit cost of a completion, and when faster is really cheaper.

Free LLM API: How to Read a Free Tier Before You Commit

A free LLM API is not a price of zero but a contract of limits: a rate limit, a quota or trial credit, a model subset, data-use terms, and the cost of leaving.