SmartGateSmartGate

Concurrency Control in AI Backend Systems: Limits, Queues, Pools

concurrency control in an AI backend is five layers, and they answer different questions. A plan catalog decides the numbers; check_rate_limit and check_rate_limit_mcp enforce them per team and per API key; rateLimitApi stops anonymous bursts at the edge with a sliding window; queues with a readable length turn overload into backpressure instead of timeouts; and a bounded connection pool plus…

Short answer: concurrency control in an AI backend is five layers, and they answer different questions. A plan catalog decides the numbers; check_rate_limit and check_rate_limit_mcp enforce them per team and per API key; rateLimitApi stops anonymous bursts at the edge with a sliding window; queues with a readable length turn overload into backpressure instead of timeouts; and a bounded connection pool plus batched work keeps the process itself from thrashing. rate limiting is searched about 4,400 times a month and api rate limiting about 720; the layers below are what those readers are actually trying to assemble.

Key takeaways

  • Counters in Redis, decisions in code. Every limit here is a window-bucketed Redis counter with a doubled TTL - cheap, atomic, and identical for REST and MCP traffic.
  • Two levels for agent traffic. check_rate_limit_mcp enforces a per-key limit and a team ceiling, and reports which one tripped (limit_scope), so an incident is diagnosable from the response alone.
  • Limits come from the plan, not from a constant. merge_rate_limits reads the catalog entry for the plan and lets a team's feature overrides win; resolve_rate_limits caches the result for 60 seconds and fails safe to the free tier.
  • Fail open, always. Every limiter returns allowed: True when Redis is unavailable. A rate limiter that becomes an availability risk is worse than the traffic it was protecting against.
  • Bound the work, not just the rate. compute_candidate_limit and extract_entities_batch show the second half of concurrency control: cap the candidate set per request and process text in batches instead of one call at a time.
  • Queues need a gauge. enqueue / dequeue are a Redis list; queueLength is what turns the queue into a signal you can alert on before users feel it.
  • Do this next: write down, for one endpoint, what happens when it is called twice as fast as today - which of the five layers catches it, and what the caller sees.

The short version for whoever owns reliability

Concurrency problems in an AI backend rarely look like concurrency problems. They look like p95 latency, a database connection error, a provider 429 passed through to a user, or a sudden monthly bill. The five layers in this article exist because each of those symptoms has a different fix, and applying the wrong one wastes the incident: rate limits protect the provider and the budget, queues protect your latency, pools protect the database, batching protects the CPU, and the plan catalog decides who gets how much.

Those five layers also answer a cost question, because each of them bounds work that would otherwise be billed; pricing, caching and routing the calls themselves is the other half, worked through in Build a Low-Cost AI Backend Architecture.

The design principle behind all five is the same: make the limit visible to the caller, and never let the enforcement mechanism become the outage. Every limiter below fails open, every refusal carries a retry hint, and every number comes from configuration rather than from a constant someone typed during an incident.

Where AI traffic differs from ordinary API traffic

Three properties make agent traffic harder to shape than a REST API. It is bursty by nature, because an agent plans several calls and fires them together. It is multi-tenant in a way that spans two identities - a team and an API key - so a single limit cannot express "this team may do 600 rpm, but no single key may do more than 60". And it is expensive per call in a way that ordinary requests are not, so the limits double as budget protection. None of that is an argument about the interface itself: if you are still deciding whether agent traffic belongs behind a protocol endpoint at all, the trade-off is set out in MCP vs REST API. The Model Context Protocol puts the tool surface behind one endpoint (MCP tools), which is convenient for enforcement: one place to count. The counter primitives are Redis's atomic INCR and EXPIRE (INCR, EXPIRE), and queueing is just a list you push and pop (Redis lists).

check_rate_limit: one window, one counter, one answer

# backend/smartgate/core/rate_limiter.py — source lines 18–48 (check_rate_limit)
async def check_rate_limit(
    team_id: str,
    max_requests: int = 60,
    window_seconds: int = 60,
) -> dict:
    try:
        redis = await get_redis()
    except Exception as exc:
        logger.warning("rate_limit fail-open: redis unavailable: %s", exc)
        return {"allowed": True, "fail_open": True}

    try:
        bucket = int(time.time() / window_seconds)
        key = f"rate_limit:{team_id}:{bucket}"

        count = await redis.incr(key)
        if count == 1:
            await redis.expire(key, window_seconds * 2)

        if count > max_requests:
            ttl = await redis.ttl(key)
            return {
                "allowed": False,
                "retry_after": max(1, ttl),
                "limit_scope": "rest_team",
            }

        return {"allowed": True}
    except Exception as exc:
        logger.warning("rate_limit fail-open: redis command failed: %s", exc)
        return {"allowed": True, "fail_open": True}

The simplest layer, and the shape all the others copy. A team id and a window define a key (rate_limit:<teamId>:<bucket>, where the bucket is derived from the current time divided by the window), INCR returns the count, and the first increment sets an expiry of twice the window so the key cannot vanish while it is still authoritative. Over the limit, the response carries retry_after from the key's TTL and names the scope (rest_team). Two failure paths - Redis unreachable, and a Redis command error - both return allowed: True, fail_open: True with a warning, which is the deliberate trade discussed above: the limiter never becomes the outage.

check_rate_limit_mcp: a per-key limit inside a team ceiling

# backend/smartgate/core/rate_limiter.py — source lines 51–91 (check_rate_limit_mcp)
async def check_rate_limit_mcp(
    key_id: str,
    team_id: str,
    *,
    per_key_limit: int,
    team_ceiling: int,
    window_seconds: int = 60,
) -> dict:
    try:
        redis = await get_redis()
    except Exception as exc:
        logger.warning("mcp rate_limit fail-open: %s", exc)
        return {"allowed": True, "fail_open": True}

    try:
        bucket = int(time.time() / window_seconds)
        key_bucket = f"rate_limit_mcp:{key_id}:{bucket}"
        team_bucket = f"rate_limit_mcp_team:{team_id}:{bucket}"

        key_count = await _incr_bucket(redis, key_bucket, window_seconds)
        if key_count > per_key_limit:
            ttl = await redis.ttl(key_bucket)
            return {
                "allowed": False,
                "retry_after": max(1, ttl),
                "limit_scope": "mcp_key",
            }

        team_count = await _incr_bucket(redis, team_bucket, window_seconds)
        if team_count > team_ceiling:
            ttl = await redis.ttl(team_bucket)
            return {
                "allowed": False,
                "retry_after": max(1, ttl),
                "limit_scope": "mcp_team",
            }

        return {"allowed": True}
    except Exception as exc:
        logger.warning("mcp rate_limit fail-open: %s", exc)
        return {"allowed": True, "fail_open": True}

Agent traffic needs two limits because it has two identities. This function increments a per-key bucket first and refuses when the key is over its own limit (limit_scope: mcp_key), then increments the team bucket and refuses when the team is over its ceiling (limit_scope: mcp_team). Both are window counters with a doubled TTL, and both refusal responses carry the retry hint. Reporting the scope is the operational detail that pays for itself: "one key is hot" and "the whole team is at the ceiling" are different incidents with different owners, and the response says which one you are in. That labelling is also what lets an ai agent observability view attribute a refusal to the right identity instead of lumping every 429 together.

A gateway is where those two limits can sit in front of every client: an MCP Gateway exposes one upstream URL and enforces per-key and per-team limits there.

getPlanRateLimits: the numbers come from a catalog

# lib/plan-rate-limits.ts — source lines 5–12 (getPlanRateLimits)
function getPlanRateLimits(plan: Plan) {
  const f = getPlanCatalogEntry(plan).features;
  return {
    restWriteRpm: f.max_team_rpm,
    mcpRpmPerKey: f.mcp_rpm_per_key,
    mcpRpmTeamCeiling: f.mcp_rpm_team_ceiling,
  };
}

Three values, one lookup: the team's REST write rate, the MCP rate per key, and the MCP team ceiling - all read from the plan's catalog entry rather than assembled from constants. Keeping them together is what stops the three limits from drifting apart when a plan changes, and it means a pricing decision is a configuration change rather than a code change.

The same catalog entry also carries the monthly token ceiling a team is measured against, and how that ceiling is read, checked and refused before a call is spent is set out in enforcing a token quota per team.

merge_rate_limits: catalog defaults, team overrides

# backend/smartgate/core/plan_entitlements.py — source lines 34–41 (merge_rate_limits)
def merge_rate_limits(plan: str, features: dict | None) -> RateLimits:
    base = CATALOG_RATE_LIMITS.get(plan.upper(), CATALOG_RATE_LIMITS["FREE"])
    f = features or {}
    return RateLimits(
        rest_write_rpm=int(f.get("max_team_rpm", base["rest_write_rpm"])),
        mcp_rpm_per_key=int(f.get("mcp_rpm_per_key", base["mcp_rpm_per_key"])),
        mcp_rpm_team_ceiling=int(f.get("mcp_rpm_team_ceiling", base["mcp_rpm_team_ceiling"])),
    )

The merge is the interesting part: the catalog entry for the plan provides a base, and the team's own feature record may override any of the three values - but only those three, and each falls back individually. That gives you three levels of control without three code paths: the published plan numbers, a per-team exception, and a safe default when neither exists (the catalog's FREE entry). If you have ever had to ship a code change to raise one customer's limit for a week, this function is the answer.

resolve_rate_limits: cache for a minute, fail safe to free

# backend/smartgate/core/plan_entitlements.py — source lines 78–110 (resolve_rate_limits)
async def resolve_rate_limits(team_id: str) -> RateLimits:
    """Resolve merged rate limits for *team_id*; cache 60s; fail-safe to FREE."""
    free_defaults = merge_rate_limits("FREE", {})
    if not team_id:
        return free_defaults

    cached = _get_cached(team_id)
    if cached is not None:
        return cached

    try:
        pool = await get_pool()
        async with pool.acquire() as conn:
            row = await conn.fetchrow(
                "SELECT plan, features FROM teams WHERE id = $1",
                team_id,
            )
        if row is None:
            limits = free_defaults
        else:
            plan = str(row["plan"] or "FREE")
            features = _parse_features(row["features"])
            limits = merge_rate_limits(plan, features)
    except Exception as exc:
        logger.warning(
            "Rate limits PG lookup failed for team %s: %s",
            team_id,
            exc,
        )
        return free_defaults

    _set_cached(team_id, limits)
    return limits

Resolution reads the team's plan and features from Postgres, merges them, caches the result for 60 seconds, and returns the free-tier limits if anything goes wrong - no team id, no row, or a database error. Two choices deserve copying. The cache is short because limits are the kind of data that must change within a minute of a plan upgrade. And the failure mode is "free tier" rather than "unlimited": when the entitlement service is unavailable, the safe direction is fewer requests, not a blank cheque.

rateLimitApi: a sliding window at the edge

# middleware.ts — source lines 15–70 (rateLimitApi)
async function rateLimitApi(pathname: string, request: NextRequest) {
  // Only rate-limit /api/ paths, exclude webhooks (verified separately)
  if (!pathname.startsWith("/api/") || pathname.startsWith("/api/webhooks/")) {
    return null;
  }

  const url = process.env.UPSTASH_REDIS_REST_URL || process.env.REDIS_URL;
  const token = process.env.UPSTASH_REDIS_REST_TOKEN || process.env.REDIS_TOKEN;
  if (!url || !token) return null;

  const ip = request.headers.get("x-forwarded-for") ?? "anonymous";
  const key = `ratelimit:middleware:${ip}`;
  const now = Date.now();
  const window = 60_000; // 1 minute
  const limit = 60;
  const windowStart = now - window;

  try {
    const redisUrl = url.startsWith("https") ? url : `https://${url}`;
    const body = JSON.stringify([
      ["zremrangebyscore", key, 0, windowStart],
      ["zcard", key],
      ["zadd", key, now, `${now}-${Math.random().toString(36).slice(2)}`],
      ["expire", key, 60],
    ]);

    const res = await fetch(`${redisUrl}/pipeline`, {
      method: "POST",
      headers: { Authorization: `Bearer ${token}`, "Content-Type": "application/json" },
      body,
    });

    if (!res.ok) return null;

    const results: Array<number | null> = await res.json();
    const count = results[1] ?? 0;

    if (Number(count) >= limit) {
      return NextResponse.json(
        { error: "Too many requests" },
        {
          status: 429,
          headers: {
            "Retry-After": "60",
            "X-RateLimit-Limit": String(limit),
            "X-RateLimit-Remaining": "0",
          },
        },
      );
    }
  } catch {
    // Rate limiting failure should not block requests
  }

  return null;
}

This layer runs in middleware, before the application, and it exists for a different class of traffic: anonymous bursts against /api/ paths that have no team id to key on yet. The implementation is a sorted-set sliding window - remove entries older than the window, count what remains, add the current request, refresh the expiry - executed as a single Redis pipeline over HTTP. Over the limit, the caller gets a 429 with Retry-After, X-RateLimit-Limit and X-RateLimit-Remaining. Three practical notes: webhook paths are excluded (they authenticate separately and cannot be rate-limited by IP), the whole block is wrapped so that any failure returns null and lets the request through, and the redis request line carries a token that this pipeline's redaction replaces with a placeholder, which is why it is quoted as ***.

compute_candidate_limit: bound the work per request

# backend/smartgate/modules/dedup/utils.py — source lines 88–113 (compute_candidate_limit)
def compute_candidate_limit(
    total: int,
    selection_size: int,
    fraction: float = 0.1,
    min_candidates: int = 100,
    max_candidates: int = 1000,
) -> int:
    """
    Compute the 'auto' candidate limit based on the total number of records.

    :param total: Total number of records.
    :param selection_size: Number of representatives to select.
    :param fraction: Fraction of total records to consider as candidates.
    :param min_candidates: Minimum number of candidates.
    :param max_candidates: Maximum number of candidates.
    :return: Computed candidate limit.
    """
    # 1) fraction of total
    limit = int(total * fraction)
    # 2) ensure enough to pick selection_size
    limit = max(limit, selection_size)
    # 3) enforce lower bound
    limit = max(limit, min_candidates)
    # 4) enforce upper bound (and never exceed the dataset)
    limit = min(limit, max_candidates, total)
    return limit

Rate limits bound how often requests arrive; this bounds how much each one costs. The function computes an "auto" candidate limit in four steps: a fraction of the total (10% by default), raised to at least the selection size, raised again to a floor of 100, and finally clamped to the ceiling and to the dataset size. The result is a request whose cost grows with the data but never exceeds a known maximum - which is the property that lets you reason about tail latency at all. Any per-request step that scales with corpus size deserves this treatment.

Bounding the work per request is one half of the bill; deciding what the request is allowed to send is the other, and that half is collected in token optimization techniques.

get_pool: pool size is a concurrency decision

# backend/smartgate/core/db.py — source lines 9–18 (get_pool)
async def get_pool() -> asyncpg.Pool:
    """Return the shared asyncpg connection pool, creating it on first call."""
    global _pool
    if _pool is None:
        _pool = await asyncpg.create_pool(
            settings.database_url,
            min_size=2,
            max_size=10,
        )
    return _pool

The pool has a floor of two connections and a ceiling of ten, created lazily on first use and shared for the process's lifetime. Small on purpose: database concurrency that exceeds what the database can serve turns into queueing inside the driver, which is invisible and worse than an explicit limit. Ten connections per process with a documented ceiling is a number you can multiply by your replica count; an unbounded pool is a number you discover during an incident.

enqueue: concurrency starts by handing work over

# lib/queue.ts — source lines 12–15 (enqueue)
async function enqueue(queueName: string, data: JobData): Promise<void> {
  const redis = getRedis();
  await redis.lpush(`queue:${queueName}`, JSON.stringify(data));
}

The queue is a Redis list with a JSON payload, pushed on the left. The interesting part is what the absence of code means: there is no retry loop, no timeout and no in-request wait - the caller hands work over and returns. When an operation can take longer than a user's patience (a bulk import, a report, a model call you want to batch), this is the boundary that turns a concurrency problem into a throughput question.

The same boundary appears in any SaaS product that moves a send or a webhook-driven action off the request path, which is the shape AI automation for SaaS operations is built on.

dequeue: FIFO concurrency from one Redis list

# lib/queue.ts — source lines 21–26 (dequeue)
async function dequeue(queueName: string): Promise<JobData | null> {
  const redis = getRedis();
  const raw = await redis.rpop(`queue:${queueName}`);
  if (!raw) return null;
  return JSON.parse(raw as string) as JobData;
}

Pop from the right and you have first-in-first-out: the list pushed on the left is drained on the right. Returning null when the queue is empty is the contract a worker loop needs - poll, get nothing, sleep, poll again - and it keeps the worker free of error handling for the ordinary case of "no work right now".

queueLength: the gauge that makes backpressure possible

# lib/queue.ts — source lines 31–34 (queueLength)
async function queueLength(queueName: string): Promise<number> {
  const redis = getRedis();
  return redis.llen(`queue:${queueName}`);
}

One line against the same key, and the most operationally useful function in the layer: the length of the queue is the signal you alert on. Without it a queue is a black box that either looks fine or is already hours deep; with it you can define "backlog above N for M minutes" and act before users feel the delay. If you ship a queue without a way to read its depth, you have shipped half a queue.

Depth is not the only signal this layer produces; the refusal counts and the limit_scope values on every response are the raw material of an LLM observability view of which limit is actually tripping. Which product renders that view best is its own question, taken up in llm observability tools.

extract_entities_batch: batching beats parallelism

# backend/smartgate/modules/memory/entity_extraction.py — source lines 147–174 (extract_entities_batch)
def extract_entities_batch(texts: List[str], batch_size: int = 32) -> List[List[Tuple[str, str]]]:
    """Extract entities from multiple texts using spaCy's nlp.pipe() for batched NER.

    Uses spaCy's efficient batch processing pipeline instead of calling
    nlp() individually per text. Significantly faster for multiple texts.

    Args:
        texts: List of input texts to extract entities from.
        batch_size: Number of texts to process in each spaCy batch.

    Returns:
        List of entity lists, one per input text. Each entity list contains
        (entity_type, entity_text) tuples. Returns list of empty lists if
        spaCy is unavailable.
    """
    if not texts:
        return []

    from mem0.utils.spacy_models import get_nlp_full

    nlp = get_nlp_full()
    if nlp is None:
        return [[] for _ in texts]

    results = []
    for doc in nlp.pipe(texts, batch_size=batch_size):
        results.append(_extract_entities_from_doc(doc))
    return results

The last layer is CPU and model work. Rather than calling the NLP pipeline once per text, this function hands the whole list to the library's batched pipe with a batch size (32 by default), which is significantly faster than a loop and uses one pass over the model. The failure behaviour is worth noting too: when the NLP model is unavailable, callers get one empty list per input rather than an exception, so a degraded extractor does not fail the whole ingestion. For agent backends the pattern generalises: batch what the model can batch, and make the degraded path a well-defined empty result instead of an error.

How SmartGate compares

Concurrency control is where "it depends who runs the infrastructure" shows up most clearly.

What you get Limits per what What you pay
Cloud API gateway (managed rate limiting at the edge) IP- and route-based limits, DDoS absorption Route, IP, sometimes an API key Gateway pricing plus your own per-tenant logic
Provider-side limits The provider's own quotas Your account, or a provider key Provider pricing, and a 429 you must translate
Self-built middleware Exactly your rules Whatever you code - and maintain Engineering time on the hot path
Gateway limits (SmartGate) Per-plan numbers, per-key and per-team MCP limits, edge sliding window, queues and a usage meter in the same path Team, API key, endpoint class Free tier: 2M tokens/mo, all seven tools, no card; Pro from $18/mo (300 req/min/key), Teams from $55/mo (600 req/min/key), Enterprise 1200 req/min/key

The advantage of limits living in the gateway is that they see the same traffic the meter sees, so a rate limit and a spend limit cannot disagree about who was calling.

The same shared view is what lets a spend owner reconcile the month, and the invoice-side reading of these meters is written up for the FinOps lead. Your plan's numbers, including the per-key and per-team MCP rates quoted above, are on the pricing page.

How to get started

  1. Put the numbers in a catalog. Plan → three or four rate values, merged with per-team overrides, cached for a minute, failing safe to the free tier.
  2. Enforce in Redis, decide in code. Window counters with doubled TTLs; refuse with a retry hint and a scope name.
  3. Add a second level for keys. A team ceiling alone cannot stop one runaway key from consuming the team's whole allowance.
  4. Stop anonymous bursts at the edge. A sliding window on IP for the paths that have no identity yet, with webhooks exempt.
  5. Make queues observable. Push, pop, and read the depth; alert on depth, not on latency.
  6. Bound per-request work and batch the rest. A clamped candidate limit and batched inference keep the tail latency a number you can predict.

Start on the free tier and watch the rate-limit headers come back on your own key: start free; the per-plan RPM values and the tool surface are in the docs and on the pricing page, and contract traffic shapes (team ceilings, dedicated pools) start with the contact form.

Frequently Asked Questions

Should a rate limiter fail open or fail closed?

Fail open, unless the thing it protects is more important than availability. Every limiter in this article logs a warning and allows the request when Redis is unreachable, because the alternative is turning a cache outage into a full outage. If your threat model disagrees, make that explicit and accept the availability cost.

Why a window counter instead of a token bucket?

Predictability. A window is explainable to the customer who hit it ("your retry window resets at :00") and trivial to debug in production; the cost is a burst at a window boundary, which the edge sliding window in front of it largely absorbs.

What is the difference between the team limit and the key limit?

Scope and blast radius. The key limit constrains one credential (a runaway agent, a leaked key); the team ceiling constrains the account. Both are needed: with only a team ceiling, one bad client can starve the rest of the team.

How do I choose max_size for the database pool?

Start from the database's own connection budget divided by your expected replica count, not from what feels fast. A small pool with an explicit ceiling fails visibly; an unbounded pool fails as latency inside the driver.

When should work go into a queue instead of being done inline?

When the caller cannot wait for it - anything measured in tens of seconds or more, anything that fans out per item, and anything you want to retry. The queue is also the cheapest way to introduce batching later without changing the API.

Does batching change the result quality?

Batched inference processes the same text with the same model; the usual caveat is ordering and memory, not accuracy. Keep the batch size configurable, and test with a real corpus before raising it - the win is throughput, not magic.

Limitations and what this does not do

  • Window counters allow boundary bursts. Two windows can pass roughly twice the limit in the second around a boundary. If that is unacceptable, add the sliding window at the edge, or move to a token bucket and accept the debugging cost.
  • Fail-open limits have a blind spot. During a Redis outage the per-key and per-team limits do not apply; the spend limit downstream is your remaining protection.
  • One quoted line carries a redacted placeholder. The edge limiter builds a bearer auth header; the slice API redacts credential-shaped strings, so the line shows a *** placeholder. It is quoted as-is rather than reconstructed.
  • Ten connections per process is a starting point, not a rule. The right pool size depends on your database, your replica count and your query latency; the value here is the shape - a documented floor, a documented ceiling, created lazily.
  • The queue has no dead-letter path in these functions. Push, pop and length are shown; a production queue also needs failure handling, which belongs in the worker.
  • Redis is a hard dependency for every layer except batching. If you self-host, size and monitor it as a critical service, not as a cache.

Sources

Method note

The code in this article is not transcribed. Each block was cut directly out of the slice body returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body before publication; the first line inside every fence records the file and the exact source lines. Symbols were pinned by whole-name containment (rule A level 2) and confirmed by the service's slot-proof endpoint before being written into the prose - 12 of 12 planned sections pinned, no abstentions. Section 7 quotes compute_candidate_limit rather than the export rate limiter that a first pass selected, because that symbol is quoted in the private-cloud-deployment rebuild published the same day and repeating a block across two articles helps neither reader.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 check_rate_limit check rate limit for concurrent ai requests check_rate_limit backend/smartgate/core/rate_limiter.py 18–48 rule A L2 → slot-proof 30f6ca7805cc
2 check_rate_limit_mcp check rate limit mcp per api key check_rate_limit_mcp backend/smartgate/core/rate_limiter.py 51–91 rule A L2 → slot-proof b540f8154beb
3 getPlanRateLimits get plan rate limits for concurrency control getPlanRateLimits lib/plan-rate-limits.ts 5–12 rule A L2 → slot-proof 6d29a819fd91
4 merge_rate_limits merge rate limits from plan features merge_rate_limits backend/smartgate/core/plan_entitlements.py 34–41 rule A L2 → slot-proof 6d65b319cf07
5 resolve_rate_limits resolve rate limits per team plan resolve_rate_limits backend/smartgate/core/plan_entitlements.py 78–110 rule A L2 → slot-proof 7b076d21efc6
6 rateLimitApi rate limit api middleware for burst control rateLimitApi middleware.ts 15–70 rule A L2 → slot-proof 9262de788a00
7 compute_candidate_limit compute candidate limit for bounded work per request compute_candidate_limit backend/smartgate/modules/dedup/utils.py 88–113 rule A L2 → slot-proof 9800d86e921c
8 get_pool get the database pool for concurrency control get_pool backend/smartgate/core/db.py 9–18 rule A L2 → slot-proof e1682a8177aa
9 enqueue enqueue jobs for concurrency control in ai backend systems enqueue lib/queue.ts 12–15 rule A L2 → slot-proof adf10f9a3fa3
10 dequeue dequeue jobs for concurrency control in ai backend systems dequeue lib/queue.ts 21–26 rule A L2 → slot-proof 8b1b3d488baa
11 queueLength queue length for backpressure in an ai backend queueLength lib/queue.ts 31–34 rule A L2 → slot-proof 7c0be9247d0c
12 extract_entities_batch extract entities batch for batched ai work extract_entities_batch backend/smartgate/modules/memory/entity_extraction.py 147–174 rule A L2 → slot-proof 21e636336488

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.