Concurrency Control in AI Backend Systems: Limits, Queues, Pools
concurrency control in an AI backend is five layers, and they answer different questions. A plan catalog decides the numbers; check_rate_limit and check_rate_limit_mcp enforce them per team and per API key; rateLimitApi stops anonymous bursts at the edge with a sliding window; queues with a readable length turn overload into backpressure instead of timeouts; and a bounded connection pool plus…
Short answer: concurrency control in an AI backend is five layers, and they answer
different questions. A plan catalog decides the numbers; check_rate_limit and
check_rate_limit_mcp enforce them per team and per API key; rateLimitApi stops anonymous
bursts at the edge with a sliding window; queues with a readable length turn overload into
backpressure instead of timeouts; and a bounded connection pool plus batched work keeps the
process itself from thrashing. rate limiting is searched about 4,400 times a month and
api rate limiting about 720; the layers below are what those readers are actually trying to
assemble.
Key takeaways
- Counters in Redis, decisions in code. Every limit here is a window-bucketed Redis counter with a doubled TTL - cheap, atomic, and identical for REST and MCP traffic.
- Two levels for agent traffic.
check_rate_limit_mcpenforces a per-key limit and a team ceiling, and reports which one tripped (limit_scope), so an incident is diagnosable from the response alone. - Limits come from the plan, not from a constant.
merge_rate_limitsreads the catalog entry for the plan and lets a team's feature overrides win;resolve_rate_limitscaches the result for 60 seconds and fails safe to the free tier. - Fail open, always. Every limiter returns
allowed: Truewhen Redis is unavailable. A rate limiter that becomes an availability risk is worse than the traffic it was protecting against. - Bound the work, not just the rate.
compute_candidate_limitandextract_entities_batchshow the second half of concurrency control: cap the candidate set per request and process text in batches instead of one call at a time. - Queues need a gauge.
enqueue/dequeueare a Redis list;queueLengthis what turns the queue into a signal you can alert on before users feel it. - Do this next: write down, for one endpoint, what happens when it is called twice as fast as today - which of the five layers catches it, and what the caller sees.
The short version for whoever owns reliability
Concurrency problems in an AI backend rarely look like concurrency problems. They look like p95 latency, a database connection error, a provider 429 passed through to a user, or a sudden monthly bill. The five layers in this article exist because each of those symptoms has a different fix, and applying the wrong one wastes the incident: rate limits protect the provider and the budget, queues protect your latency, pools protect the database, batching protects the CPU, and the plan catalog decides who gets how much.
Those five layers also answer a cost question, because each of them bounds work that would otherwise be billed; pricing, caching and routing the calls themselves is the other half, worked through in Build a Low-Cost AI Backend Architecture.
The design principle behind all five is the same: make the limit visible to the caller, and never let the enforcement mechanism become the outage. Every limiter below fails open, every refusal carries a retry hint, and every number comes from configuration rather than from a constant someone typed during an incident.
Where AI traffic differs from ordinary API traffic
Three properties make agent traffic harder to shape than a REST API. It is bursty by nature,
because an agent plans several calls and fires them together. It is multi-tenant in a way
that spans two identities - a team and an API key - so a single limit cannot express "this
team may do 600 rpm, but no single key may do more than 60". And it is expensive per call in a
way that ordinary requests are not, so the limits double as budget protection. None of that is an argument about the
interface itself: if you are still deciding whether agent traffic belongs behind a
protocol endpoint at all, the trade-off is set out in
MCP vs REST API. The Model
Context Protocol puts the tool surface behind one endpoint
(MCP tools), which is
convenient for enforcement: one place to count. The counter primitives are Redis's atomic
INCR and EXPIRE
(INCR,
EXPIRE), and queueing is just a list you push
and pop (Redis lists).
check_rate_limit: one window, one counter, one answer
# backend/smartgate/core/rate_limiter.py — source lines 18–48 (check_rate_limit)
async def check_rate_limit(
team_id: str,
max_requests: int = 60,
window_seconds: int = 60,
) -> dict:
try:
redis = await get_redis()
except Exception as exc:
logger.warning("rate_limit fail-open: redis unavailable: %s", exc)
return {"allowed": True, "fail_open": True}
try:
bucket = int(time.time() / window_seconds)
key = f"rate_limit:{team_id}:{bucket}"
count = await redis.incr(key)
if count == 1:
await redis.expire(key, window_seconds * 2)
if count > max_requests:
ttl = await redis.ttl(key)
return {
"allowed": False,
"retry_after": max(1, ttl),
"limit_scope": "rest_team",
}
return {"allowed": True}
except Exception as exc:
logger.warning("rate_limit fail-open: redis command failed: %s", exc)
return {"allowed": True, "fail_open": True}
The simplest layer, and the shape all the others copy. A team id and a window define a key
(rate_limit:<teamId>:<bucket>, where the bucket is derived from the current time divided by
the window), INCR returns the count, and the first increment sets an expiry of twice the
window so the key cannot vanish while it is still authoritative. Over the limit, the response
carries retry_after from the key's TTL and names the scope (rest_team). Two failure paths -
Redis unreachable, and a Redis command error - both return allowed: True, fail_open: True
with a warning, which is the deliberate trade discussed above: the limiter never becomes the
outage.
check_rate_limit_mcp: a per-key limit inside a team ceiling
# backend/smartgate/core/rate_limiter.py — source lines 51–91 (check_rate_limit_mcp)
async def check_rate_limit_mcp(
key_id: str,
team_id: str,
*,
per_key_limit: int,
team_ceiling: int,
window_seconds: int = 60,
) -> dict:
try:
redis = await get_redis()
except Exception as exc:
logger.warning("mcp rate_limit fail-open: %s", exc)
return {"allowed": True, "fail_open": True}
try:
bucket = int(time.time() / window_seconds)
key_bucket = f"rate_limit_mcp:{key_id}:{bucket}"
team_bucket = f"rate_limit_mcp_team:{team_id}:{bucket}"
key_count = await _incr_bucket(redis, key_bucket, window_seconds)
if key_count > per_key_limit:
ttl = await redis.ttl(key_bucket)
return {
"allowed": False,
"retry_after": max(1, ttl),
"limit_scope": "mcp_key",
}
team_count = await _incr_bucket(redis, team_bucket, window_seconds)
if team_count > team_ceiling:
ttl = await redis.ttl(team_bucket)
return {
"allowed": False,
"retry_after": max(1, ttl),
"limit_scope": "mcp_team",
}
return {"allowed": True}
except Exception as exc:
logger.warning("mcp rate_limit fail-open: %s", exc)
return {"allowed": True, "fail_open": True}
Agent traffic needs two limits because it has two identities. This function increments a
per-key bucket first and refuses when the key is over its own limit (limit_scope: mcp_key),
then increments the team bucket and refuses when the team is over its ceiling
(limit_scope: mcp_team). Both are window counters with a doubled TTL, and both refusal
responses carry the retry hint. Reporting the scope is the operational detail that pays for
itself: "one key is hot" and "the whole team is at the ceiling" are different incidents with
different owners, and the response says which one you are in. That labelling is also what lets
an ai agent observability view attribute a refusal to the
right identity instead of lumping every 429 together.
A gateway is where those two limits can sit in front of every client: an MCP Gateway exposes one upstream URL and enforces per-key and per-team limits there.
getPlanRateLimits: the numbers come from a catalog
# lib/plan-rate-limits.ts — source lines 5–12 (getPlanRateLimits)
function getPlanRateLimits(plan: Plan) {
const f = getPlanCatalogEntry(plan).features;
return {
restWriteRpm: f.max_team_rpm,
mcpRpmPerKey: f.mcp_rpm_per_key,
mcpRpmTeamCeiling: f.mcp_rpm_team_ceiling,
};
}
Three values, one lookup: the team's REST write rate, the MCP rate per key, and the MCP team ceiling - all read from the plan's catalog entry rather than assembled from constants. Keeping them together is what stops the three limits from drifting apart when a plan changes, and it means a pricing decision is a configuration change rather than a code change.
The same catalog entry also carries the monthly token ceiling a team is measured against, and how that ceiling is read, checked and refused before a call is spent is set out in enforcing a token quota per team.
merge_rate_limits: catalog defaults, team overrides
# backend/smartgate/core/plan_entitlements.py — source lines 34–41 (merge_rate_limits)
def merge_rate_limits(plan: str, features: dict | None) -> RateLimits:
base = CATALOG_RATE_LIMITS.get(plan.upper(), CATALOG_RATE_LIMITS["FREE"])
f = features or {}
return RateLimits(
rest_write_rpm=int(f.get("max_team_rpm", base["rest_write_rpm"])),
mcp_rpm_per_key=int(f.get("mcp_rpm_per_key", base["mcp_rpm_per_key"])),
mcp_rpm_team_ceiling=int(f.get("mcp_rpm_team_ceiling", base["mcp_rpm_team_ceiling"])),
)
The merge is the interesting part: the catalog entry for the plan provides a base, and the
team's own feature record may override any of the three values - but only those three, and each
falls back individually. That gives you three levels of control without three code paths: the
published plan numbers, a per-team exception, and a safe default when neither exists (the
catalog's FREE entry). If you have ever had to ship a code change to raise one customer's
limit for a week, this function is the answer.
resolve_rate_limits: cache for a minute, fail safe to free
# backend/smartgate/core/plan_entitlements.py — source lines 78–110 (resolve_rate_limits)
async def resolve_rate_limits(team_id: str) -> RateLimits:
"""Resolve merged rate limits for *team_id*; cache 60s; fail-safe to FREE."""
free_defaults = merge_rate_limits("FREE", {})
if not team_id:
return free_defaults
cached = _get_cached(team_id)
if cached is not None:
return cached
try:
pool = await get_pool()
async with pool.acquire() as conn:
row = await conn.fetchrow(
"SELECT plan, features FROM teams WHERE id = $1",
team_id,
)
if row is None:
limits = free_defaults
else:
plan = str(row["plan"] or "FREE")
features = _parse_features(row["features"])
limits = merge_rate_limits(plan, features)
except Exception as exc:
logger.warning(
"Rate limits PG lookup failed for team %s: %s",
team_id,
exc,
)
return free_defaults
_set_cached(team_id, limits)
return limits
Resolution reads the team's plan and features from Postgres, merges them, caches the result for 60 seconds, and returns the free-tier limits if anything goes wrong - no team id, no row, or a database error. Two choices deserve copying. The cache is short because limits are the kind of data that must change within a minute of a plan upgrade. And the failure mode is "free tier" rather than "unlimited": when the entitlement service is unavailable, the safe direction is fewer requests, not a blank cheque.
rateLimitApi: a sliding window at the edge
# middleware.ts — source lines 15–70 (rateLimitApi)
async function rateLimitApi(pathname: string, request: NextRequest) {
// Only rate-limit /api/ paths, exclude webhooks (verified separately)
if (!pathname.startsWith("/api/") || pathname.startsWith("/api/webhooks/")) {
return null;
}
const url = process.env.UPSTASH_REDIS_REST_URL || process.env.REDIS_URL;
const token = process.env.UPSTASH_REDIS_REST_TOKEN || process.env.REDIS_TOKEN;
if (!url || !token) return null;
const ip = request.headers.get("x-forwarded-for") ?? "anonymous";
const key = `ratelimit:middleware:${ip}`;
const now = Date.now();
const window = 60_000; // 1 minute
const limit = 60;
const windowStart = now - window;
try {
const redisUrl = url.startsWith("https") ? url : `https://${url}`;
const body = JSON.stringify([
["zremrangebyscore", key, 0, windowStart],
["zcard", key],
["zadd", key, now, `${now}-${Math.random().toString(36).slice(2)}`],
["expire", key, 60],
]);
const res = await fetch(`${redisUrl}/pipeline`, {
method: "POST",
headers: { Authorization: `Bearer ${token}`, "Content-Type": "application/json" },
body,
});
if (!res.ok) return null;
const results: Array<number | null> = await res.json();
const count = results[1] ?? 0;
if (Number(count) >= limit) {
return NextResponse.json(
{ error: "Too many requests" },
{
status: 429,
headers: {
"Retry-After": "60",
"X-RateLimit-Limit": String(limit),
"X-RateLimit-Remaining": "0",
},
},
);
}
} catch {
// Rate limiting failure should not block requests
}
return null;
}
This layer runs in middleware, before the application, and it exists for a different class of
traffic: anonymous bursts against /api/ paths that have no team id to key on yet. The
implementation is a sorted-set sliding window - remove entries older than the window, count
what remains, add the current request, refresh the expiry - executed as a single Redis pipeline
over HTTP. Over the limit, the caller gets a 429 with Retry-After, X-RateLimit-Limit and
X-RateLimit-Remaining. Three practical notes: webhook paths are excluded (they authenticate
separately and cannot be rate-limited by IP), the whole block is wrapped so that any failure
returns null and lets the request through, and the redis request line carries a token that
this pipeline's redaction replaces with a placeholder, which is why it is quoted as ***.
compute_candidate_limit: bound the work per request
# backend/smartgate/modules/dedup/utils.py — source lines 88–113 (compute_candidate_limit)
def compute_candidate_limit(
total: int,
selection_size: int,
fraction: float = 0.1,
min_candidates: int = 100,
max_candidates: int = 1000,
) -> int:
"""
Compute the 'auto' candidate limit based on the total number of records.
:param total: Total number of records.
:param selection_size: Number of representatives to select.
:param fraction: Fraction of total records to consider as candidates.
:param min_candidates: Minimum number of candidates.
:param max_candidates: Maximum number of candidates.
:return: Computed candidate limit.
"""
# 1) fraction of total
limit = int(total * fraction)
# 2) ensure enough to pick selection_size
limit = max(limit, selection_size)
# 3) enforce lower bound
limit = max(limit, min_candidates)
# 4) enforce upper bound (and never exceed the dataset)
limit = min(limit, max_candidates, total)
return limit
Rate limits bound how often requests arrive; this bounds how much each one costs. The function computes an "auto" candidate limit in four steps: a fraction of the total (10% by default), raised to at least the selection size, raised again to a floor of 100, and finally clamped to the ceiling and to the dataset size. The result is a request whose cost grows with the data but never exceeds a known maximum - which is the property that lets you reason about tail latency at all. Any per-request step that scales with corpus size deserves this treatment.
Bounding the work per request is one half of the bill; deciding what the request is allowed to send is the other, and that half is collected in token optimization techniques.
get_pool: pool size is a concurrency decision
# backend/smartgate/core/db.py — source lines 9–18 (get_pool)
async def get_pool() -> asyncpg.Pool:
"""Return the shared asyncpg connection pool, creating it on first call."""
global _pool
if _pool is None:
_pool = await asyncpg.create_pool(
settings.database_url,
min_size=2,
max_size=10,
)
return _pool
The pool has a floor of two connections and a ceiling of ten, created lazily on first use and shared for the process's lifetime. Small on purpose: database concurrency that exceeds what the database can serve turns into queueing inside the driver, which is invisible and worse than an explicit limit. Ten connections per process with a documented ceiling is a number you can multiply by your replica count; an unbounded pool is a number you discover during an incident.
enqueue: concurrency starts by handing work over
# lib/queue.ts — source lines 12–15 (enqueue)
async function enqueue(queueName: string, data: JobData): Promise<void> {
const redis = getRedis();
await redis.lpush(`queue:${queueName}`, JSON.stringify(data));
}
The queue is a Redis list with a JSON payload, pushed on the left. The interesting part is what the absence of code means: there is no retry loop, no timeout and no in-request wait - the caller hands work over and returns. When an operation can take longer than a user's patience (a bulk import, a report, a model call you want to batch), this is the boundary that turns a concurrency problem into a throughput question.
The same boundary appears in any SaaS product that moves a send or a webhook-driven action off the request path, which is the shape AI automation for SaaS operations is built on.
dequeue: FIFO concurrency from one Redis list
# lib/queue.ts — source lines 21–26 (dequeue)
async function dequeue(queueName: string): Promise<JobData | null> {
const redis = getRedis();
const raw = await redis.rpop(`queue:${queueName}`);
if (!raw) return null;
return JSON.parse(raw as string) as JobData;
}
Pop from the right and you have first-in-first-out: the list pushed on the left is drained on
the right. Returning null when the queue is empty is the contract a worker loop needs - poll,
get nothing, sleep, poll again - and it keeps the worker free of error handling for the
ordinary case of "no work right now".
queueLength: the gauge that makes backpressure possible
# lib/queue.ts — source lines 31–34 (queueLength)
async function queueLength(queueName: string): Promise<number> {
const redis = getRedis();
return redis.llen(`queue:${queueName}`);
}
One line against the same key, and the most operationally useful function in the layer: the length of the queue is the signal you alert on. Without it a queue is a black box that either looks fine or is already hours deep; with it you can define "backlog above N for M minutes" and act before users feel the delay. If you ship a queue without a way to read its depth, you have shipped half a queue.
Depth is not the only signal this layer produces; the refusal counts and the limit_scope
values on every response are the raw material of an
LLM observability view of which limit is actually tripping.
Which product renders that view best is its own question, taken up in
llm observability tools.
extract_entities_batch: batching beats parallelism
# backend/smartgate/modules/memory/entity_extraction.py — source lines 147–174 (extract_entities_batch)
def extract_entities_batch(texts: List[str], batch_size: int = 32) -> List[List[Tuple[str, str]]]:
"""Extract entities from multiple texts using spaCy's nlp.pipe() for batched NER.
Uses spaCy's efficient batch processing pipeline instead of calling
nlp() individually per text. Significantly faster for multiple texts.
Args:
texts: List of input texts to extract entities from.
batch_size: Number of texts to process in each spaCy batch.
Returns:
List of entity lists, one per input text. Each entity list contains
(entity_type, entity_text) tuples. Returns list of empty lists if
spaCy is unavailable.
"""
if not texts:
return []
from mem0.utils.spacy_models import get_nlp_full
nlp = get_nlp_full()
if nlp is None:
return [[] for _ in texts]
results = []
for doc in nlp.pipe(texts, batch_size=batch_size):
results.append(_extract_entities_from_doc(doc))
return results
The last layer is CPU and model work. Rather than calling the NLP pipeline once per text, this function hands the whole list to the library's batched pipe with a batch size (32 by default), which is significantly faster than a loop and uses one pass over the model. The failure behaviour is worth noting too: when the NLP model is unavailable, callers get one empty list per input rather than an exception, so a degraded extractor does not fail the whole ingestion. For agent backends the pattern generalises: batch what the model can batch, and make the degraded path a well-defined empty result instead of an error.
How SmartGate compares
Concurrency control is where "it depends who runs the infrastructure" shows up most clearly.
| What you get | Limits per what | What you pay | |
|---|---|---|---|
| Cloud API gateway (managed rate limiting at the edge) | IP- and route-based limits, DDoS absorption | Route, IP, sometimes an API key | Gateway pricing plus your own per-tenant logic |
| Provider-side limits | The provider's own quotas | Your account, or a provider key | Provider pricing, and a 429 you must translate |
| Self-built middleware | Exactly your rules | Whatever you code - and maintain | Engineering time on the hot path |
| Gateway limits (SmartGate) | Per-plan numbers, per-key and per-team MCP limits, edge sliding window, queues and a usage meter in the same path | Team, API key, endpoint class | Free tier: 2M tokens/mo, all seven tools, no card; Pro from $18/mo (300 req/min/key), Teams from $55/mo (600 req/min/key), Enterprise 1200 req/min/key |
The advantage of limits living in the gateway is that they see the same traffic the meter sees, so a rate limit and a spend limit cannot disagree about who was calling.
The same shared view is what lets a spend owner reconcile the month, and the invoice-side reading of these meters is written up for the FinOps lead. Your plan's numbers, including the per-key and per-team MCP rates quoted above, are on the pricing page.
How to get started
- Put the numbers in a catalog. Plan → three or four rate values, merged with per-team overrides, cached for a minute, failing safe to the free tier.
- Enforce in Redis, decide in code. Window counters with doubled TTLs; refuse with a retry hint and a scope name.
- Add a second level for keys. A team ceiling alone cannot stop one runaway key from consuming the team's whole allowance.
- Stop anonymous bursts at the edge. A sliding window on IP for the paths that have no identity yet, with webhooks exempt.
- Make queues observable. Push, pop, and read the depth; alert on depth, not on latency.
- Bound per-request work and batch the rest. A clamped candidate limit and batched inference keep the tail latency a number you can predict.
Start on the free tier and watch the rate-limit headers come back on your own key: start free; the per-plan RPM values and the tool surface are in the docs and on the pricing page, and contract traffic shapes (team ceilings, dedicated pools) start with the contact form.
Frequently Asked Questions
Should a rate limiter fail open or fail closed?
Fail open, unless the thing it protects is more important than availability. Every limiter in this article logs a warning and allows the request when Redis is unreachable, because the alternative is turning a cache outage into a full outage. If your threat model disagrees, make that explicit and accept the availability cost.
Why a window counter instead of a token bucket?
Predictability. A window is explainable to the customer who hit it ("your retry window resets at :00") and trivial to debug in production; the cost is a burst at a window boundary, which the edge sliding window in front of it largely absorbs.
What is the difference between the team limit and the key limit?
Scope and blast radius. The key limit constrains one credential (a runaway agent, a leaked key); the team ceiling constrains the account. Both are needed: with only a team ceiling, one bad client can starve the rest of the team.
How do I choose max_size for the database pool?
Start from the database's own connection budget divided by your expected replica count, not from what feels fast. A small pool with an explicit ceiling fails visibly; an unbounded pool fails as latency inside the driver.
When should work go into a queue instead of being done inline?
When the caller cannot wait for it - anything measured in tens of seconds or more, anything that fans out per item, and anything you want to retry. The queue is also the cheapest way to introduce batching later without changing the API.
Does batching change the result quality?
Batched inference processes the same text with the same model; the usual caveat is ordering and memory, not accuracy. Keep the batch size configurable, and test with a real corpus before raising it - the win is throughput, not magic.
Limitations and what this does not do
- Window counters allow boundary bursts. Two windows can pass roughly twice the limit in the second around a boundary. If that is unacceptable, add the sliding window at the edge, or move to a token bucket and accept the debugging cost.
- Fail-open limits have a blind spot. During a Redis outage the per-key and per-team limits do not apply; the spend limit downstream is your remaining protection.
- One quoted line carries a redacted placeholder. The edge limiter builds a bearer auth
header; the slice API redacts credential-shaped strings, so the line shows a
***placeholder. It is quoted as-is rather than reconstructed. - Ten connections per process is a starting point, not a rule. The right pool size depends on your database, your replica count and your query latency; the value here is the shape - a documented floor, a documented ceiling, created lazily.
- The queue has no dead-letter path in these functions. Push, pop and length are shown; a production queue also needs failure handling, which belongs in the worker.
- Redis is a hard dependency for every layer except batching. If you self-host, size and monitor it as a critical service, not as a cache.
Sources
- Redis — INCR, the atomic counter behind every limiter: https://redis.io/docs/latest/commands/incr/
- Redis — EXPIRE, and why window keys get a doubled TTL: https://redis.io/docs/latest/commands/expire/
- Redis — lists, the data structure behind the queue: https://redis.io/docs/latest/develop/data-types/lists/
- Redis — key patterns and expiry practices: https://redis.io/docs/latest/develop/use/patterns/
- Model Context Protocol — the tool surface this traffic arrives through: https://modelcontextprotocol.io/specification/2026-07-28/server/tools
- asyncpg — the PostgreSQL pool used by the backend: https://magicstack.github.io/asyncpg/current/
- SmartGate — documentation, pricing and sales contact: https://smartgate.network/docs · https://smartgate.network/pricing · https://smartgate.network/contact
Method note
The code in this article is not transcribed. Each block was cut directly out of the slice body
returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body
before publication; the first line inside every fence records the file and the exact source
lines. Symbols were pinned by whole-name containment (rule A level 2) and confirmed by the
service's slot-proof endpoint before being written into the prose - 12 of 12 planned sections
pinned, no abstentions. Section 7 quotes compute_candidate_limit rather than the export rate
limiter that a first pass selected, because that symbol is quoted in the
private-cloud-deployment rebuild published the same day and repeating a block across two
articles helps neither reader.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | check_rate_limit check rate limit for concurrent ai requests | check_rate_limit |
backend/smartgate/core/rate_limiter.py |
18–48 | rule A L2 → slot-proof | 30f6ca7805cc |
| 2 | check_rate_limit_mcp check rate limit mcp per api key | check_rate_limit_mcp |
backend/smartgate/core/rate_limiter.py |
51–91 | rule A L2 → slot-proof | b540f8154beb |
| 3 | getPlanRateLimits get plan rate limits for concurrency control | getPlanRateLimits |
lib/plan-rate-limits.ts |
5–12 | rule A L2 → slot-proof | 6d29a819fd91 |
| 4 | merge_rate_limits merge rate limits from plan features | merge_rate_limits |
backend/smartgate/core/plan_entitlements.py |
34–41 | rule A L2 → slot-proof | 6d65b319cf07 |
| 5 | resolve_rate_limits resolve rate limits per team plan | resolve_rate_limits |
backend/smartgate/core/plan_entitlements.py |
78–110 | rule A L2 → slot-proof | 7b076d21efc6 |
| 6 | rateLimitApi rate limit api middleware for burst control | rateLimitApi |
middleware.ts |
15–70 | rule A L2 → slot-proof | 9262de788a00 |
| 7 | compute_candidate_limit compute candidate limit for bounded work per request | compute_candidate_limit |
backend/smartgate/modules/dedup/utils.py |
88–113 | rule A L2 → slot-proof | 9800d86e921c |
| 8 | get_pool get the database pool for concurrency control | get_pool |
backend/smartgate/core/db.py |
9–18 | rule A L2 → slot-proof | e1682a8177aa |
| 9 | enqueue enqueue jobs for concurrency control in ai backend systems | enqueue |
lib/queue.ts |
12–15 | rule A L2 → slot-proof | adf10f9a3fa3 |
| 10 | dequeue dequeue jobs for concurrency control in ai backend systems | dequeue |
lib/queue.ts |
21–26 | rule A L2 → slot-proof | 8b1b3d488baa |
| 11 | queueLength queue length for backpressure in an ai backend | queueLength |
lib/queue.ts |
31–34 | rule A L2 → slot-proof | 7c0be9247d0c |
| 12 | extract_entities_batch extract entities batch for batched ai work | extract_entities_batch |
backend/smartgate/modules/memory/entity_extraction.py |
147–174 | rule A L2 → slot-proof | 21e636336488 |
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.