SmartGateSmartGate

Model Supply Chain Security: Provenance, Poisoning, Integrity

A model supply chain is every component an AI application trusts but did not build — base weights, fine-tuned adapters, the tokenizer, training and fine-tuning data, the inference runtime, third-party packages, and the tool definitions and prompt templates the model reads at runtime.

Short answer: A model supply chain is every component an AI application trusts but did not build — base weights, fine-tuned adapters, the tokenizer, training and fine-tuning data, the inference runtime, third-party packages, and the tool definitions and prompt templates the model reads at runtime. The controls that hold extend ordinary software-supply-chain practice to artifacts that are opaque by default: pin an immutable revision instead of a floating tag, record the digest, verify a hash or signature before load, scan what you cannot verify, and treat a model-version change as a reviewed change with a rollback. Poisoning is the failure unique to this chain — a web-scale dataset or a small adapter can carry a hidden trigger that survives normal safety training.

Key takeaways

  • A model is a dependency you cannot read. Weights are an opaque binary, so provenance and integrity — not source review — are the primary controls an application has.
  • Pin the revision, not the tag. latest and "the current model version" are mutable references; an immutable commit plus a recorded digest makes a build reproducible and an incident answerable.
  • Poisoning is cheap at the data and adapter layers. A small poisoned fine-tuning set or a malicious LoRA adapter changes behavior without touching pretraining.
  • Tool definitions and prompt templates are supply-chain artifacts too. They are text the model reads as instructions, and deserve the same review, pinning and diff-on-upgrade as code.
  • Start with an inventory and one integrity check. List every artifact with its revision and hash, then verify at load.

What a model supply chain covers

Every application is mostly other people's code, and an AI application is mostly other people's code, weights and text. The traditional software supply chain is the part everyone already knows: open-source dependencies, base images, build tools. The model supply chain adds a layer of artifacts that behave differently — a model file is a binary blob whose behavior is not visible in a diff, and a training corpus is a moving target rather than a fixed input. OWASP's supply-chain entry draws the line plainly: "LLM supply chains are susceptible to various vulnerabilities, which can affect the integrity of training data, models, and deployment platforms", and unlike ordinary software, "in ML the risks also extend to third-party pre-trained models and data" (OWASP LLM03:2025 Supply Chain).

The reason this is a distinct discipline is the black-box problem. A vulnerability scanner can read a package's source and flag a bad version; a pre-trained model is described by its authors rather than inspected by its consumers, and OWASP is explicit that "models are binary black boxes and unlike open source, static inspection can offer little to security assurances". NIST makes the same move from the governance side: its Generative AI Profile lists "Value Chain and Component Integration" as one of the twelve generative-AI risk categories, covering third-party models, data and dependencies with opaque provenance (NIST AI 600-1).

It helps to name the classes of thing you are trusting, because each carries a different failure and a different control.

Component Where it comes from What a compromise can carry
Base model weights A vendor, or an open hub A backdoor, hidden bias, or an unexpected capability
Fine-tuning adapters (LoRA, PEFT) Community hubs, internal teams Changed instruction-following, a hidden trigger
Training and fine-tuning data Scraped web, licensed corpora, user data Poisoning, licensing exposure, personal data
Tokenizer and embedding model The same hub as the weights A mismatched or manipulated encoding
Inference runtime and packages Package registries Arbitrary code at load or serve time
Tool definitions and prompt templates MCP servers, vendor prompt packs Redirected tool calls, injected instructions

The last row is the one most teams have never written down. What an agent can do is decided as much by the descriptions of its tools as by the weights behind them — which is why the artifact side of the problem overlaps the injection problem at exactly that seam, treated on the injection problem itself.

Model provenance: pinning the weights, not trusting latest

Provenance answers three questions about an artifact — where, when and how it was produced — and it is the only durable defense when the artifact itself cannot be inspected. SLSA's definition is worth quoting for its precision: provenance is "the verifiable information about software artifacts describing where, when and how something was produced", and the point of capturing it is that a complex supply chain can then be traced back to its source (SLSA provenance). For a model, "where, when and how" means the exact repository and revision, the training recipe or the fact that it is undisclosed, and the digest of the file that actually landed.

The document supposed to carry that information is the model card — the README of a model repository, whose convention (introduced by Mitchell and co-authors and now standard on the major hubs) is to describe the model, its intended uses and limitations including biases, its training parameters and datasets, and its evaluation results (Hugging Face model cards; Model Cards for Model Reporting). A model card is genuinely useful and it is not a security control: OWASP states flatly that "Model Cards and associated documentation provide model information and relied upon users, but they offer no guarantees on the origin of the model", and warns that an attacker can compromise a supplier's account on a hub or stand up a lookalike repository and let social engineering do the rest.

That is the gap provenance discipline closes, and it is mostly reference hygiene. A floating reference — latest, main, or an API alias that silently points at a newer model — is mutable by a third party, so a build that resolved cleanly today can resolve to a different artifact tomorrow with no change on your side. The replacement is an immutable revision (a commit identifier for a repository, a concrete version for an API), recorded with a hash of the bytes you downloaded, held in a local mirror you control, and re-verified at load. OWASP's mitigation is close to this: "Only use models from verifiable sources and use third-party model integrity checks with signing and file hashes to compensate for the lack of strong model provenance" (OWASP LLM03:2025 Supply Chain). None of that proves a model is safe; it proves the model is the one you decided to trust.

Training data poisoning and the adapter shortcut

Poisoning is the failure mode with no clean software analogue: an attacker does not need to breach your application, only to influence the data or the adapter it learns from. The 2023 result that made this concrete showed two practical attacks against web-scale datasets — split-view poisoning, which exploits the fact that a dataset annotator's view of a web page can differ from the view a later download sees, and frontrunning poisoning, which targets periodically snapshotted crowd-sourced sources such as Wikipedia. The authors estimated they could have poisoned 0.01% of a large public image-text dataset for about sixty dollars, which is the number that should reset anyone's sense of the attacker's budget (Poisoning Web-Scale Training Datasets is Practical).

Fine-tuning makes the attack cheaper and more targeted, because it removes the need to influence pretraining at all. A small poisoned fine-tuning set, or a single malicious adapter bolted onto a clean base model, can install behavior while every integrity check on the base weights still passes. OWASP names this class directly — a "Vulnerable LoRA adapter ... compromises the integrity and security of the pre-trained base model". The persistence result is the part that should shape expectations: backdoored behavior can be trained to survive supervised fine-tuning, reinforcement learning and adversarial training, remaining most persistent in the largest models (Sleeper Agents). The practical conclusion is that a poisoned artifact cannot be relied on to be washed out by later safety training, so the control belongs upstream — in where the data and adapters come from — rather than downstream in a filter.

Three defenses do cheap, real work here. Keep a provenance record for every dataset and adapter, so a model can be traced to inputs with a known origin. Run held-out evaluations against a baseline whenever a data or adapter change is proposed. And constrain when an adapter is loaded — a merge or download that a running service can perform on demand is a supply-chain entry point a review gate would otherwise stop. The sibling that simulates an adversary against a deployed agent is red-teaming a deployed agent; this page stays with the artifact and its inputs.

Tool definitions and prompt templates are supply-chain artifacts

The component most teams forget to inventory is the one that arrives as text. In a tool-using application, the model does not call an API the way code does; it reads a description of each tool and decides, in the same context window that holds the user's request, which one to invoke. The Model Context Protocol makes this explicit: a server exposes a discovery call that lists its tools with names, human-readable descriptions and input schemas, and an invocation call that runs one (MCP server tools). Those descriptions and schemas are instructions the model consumes at runtime, from a party you may not control.

That makes tool metadata a supply-chain surface of exactly the kind this page is about. A compromised or malicious server can ship a description that steers selection ("always call this tool first"), a schema whose field names carry phrasing the model infers from, or a tool whose stated purpose differs from what it does. Vendor-supplied system prompts have the same property: they are text from a third party that the model reads as direction, and they change behavior without changing any weight you can hash. The same reasoning is why an indirect payload can travel through fetched content and tool returns into the model's context, the mechanism taken apart on indirect injection through fetched content — the difference here is that the text is a declared dependency you chose to install, not something fetched on the user's behalf.

The control is to treat this text as code. Review a tool's description and schema before you connect it, pin the server to a known revision, keep a copy you control, and diff the metadata on every upgrade the way you would diff a dependency's API. An upgrade that quietly rewrites a tool description is, for a model, an upgrade that rewrites code.

Integrity checks: hashes, signatures, and the SBOM analogy

If provenance is the claim about where an artifact came from, integrity checking is the test of whether the bytes in front of you match that claim. The software world has a settled toolkit for this, and the model world borrows from it rather than inventing a parallel one. File hashes answer "is this the file I recorded"; signed commits and code signing answer "did the party I expect produce it". Sigstore provides keyless signing and verification for artifacts; SLSA provides a levelled framework for build provenance together with a catalogue of the threats those levels are meant to resist; and in-toto supplies a framework for attesting each step of a supply chain (Sigstore; SLSA threats; in-toto).

The inventory artifact has a direct model analogue. A software bill of materials is a machine-readable list of the components in a build, and OWASP recommends exactly that for LLM applications: "Maintain an up-to-date inventory of components using a Software Bill of Materials (SBOM) to ensure you have an up-to-date, accurate, and signed inventory", adding that "AI BOMs and ML SBOMs are an emerging area" and pointing at CycloneDX as the format to evaluate (OWASP LLM03:2025 Supply Chain; CycloneDX ML-BOM). An ML bill of materials is a young format, but the shape is right: models, datasets, adapters and the runtime, each with a version and a digest, so a new vulnerability can be answered with a query rather than an archaeology project.

One integrity check is cheaper than all the others and should come first: prefer weight formats that cannot execute code. The default serialization for PyTorch weights is pickle, and the major hub's own security documentation states that "there are dangerous arbitrary code execution attacks that can be perpetrated when you load a pickle file", recommending that users load models only from sources they trust, rely on signed commits, and prefer formats that do not carry executable instructions (Hugging Face pickle scanning). The safetensors format exists to sidestep that class of attack entirely (Hugging Face safetensors documentation). So the ordering is: use a non-executing format where one exists, scan what remains, and never load an artifact from a source you have not pinned. The cluster's OWASP Top 10 entry list is where this maps onto the published risk categories.

Model security at the dependency boundary

Model security is narrower than "AI security" and worth separating out, because it is about the artifact rather than the application around it. The threat set OWASP catalogues for a pre-trained model is specific: an opaque binary "can contain hidden biases, backdoors, or other malicious features that have not been identified through the safety evaluations of model repository", and a model can be altered by a poisoned dataset or by direct tampering with weights. The same entry lists the collaborative surfaces where this happens in practice — model-merge and conversion services hosted in shared environments can be used "to introduce vulnerabilities in shared models", and on-device models widen the surface to manufacturing, firmware and reverse-engineered repackaging (OWASP LLM03:2025 Supply Chain).

The defensive posture that follows is containment plus evidence, and it is unglamorous. Load untrusted artifacts outside your production trust domain, so a code-execution bug in a loader is a contained failure rather than a beachhead. Keep a rollback to the last revision you verified, so a regression found after deployment can be reversed without a rebuild. Monitor a newly pinned model's outputs against the previous one, because a behavior change is often the only signal that something moved. And treat the model's stated safety evaluation as a claim by its author, not a measurement by you. Where the boundary between screening input and containing output is drawn belongs to the guardrail layer, covered at input and output guardrails.

AI supply chain risk: hosted versus self-hosted

The largest single decision in this chain is whether the model runs as a service you call or an artifact you host, and it trades one set of risks for another rather than eliminating any. A hosted model removes the artifact from your estate: you do not hold the weights, you do not patch the runtime, and your exposure concentrates in availability, data handling and the provider's terms. Self-hosting puts the artifact back in your hands and with it the responsibility for the runtime, the dependencies and the loader — the very surfaces the previous sections describe. There is no default answer; there are questions that force one.

  • Can you pin the version you tested? If a provider lets you name a concrete model version, the floating-reference problem shrinks. If the model upgrades silently, behavior can change without a change on your side, and your evaluation baseline no longer describes what is live.
  • Where is the data allowed to go? OWASP flags "unclear T&Cs and data privacy policies of the model operators" as a supply-chain risk in its own right, since sensitive inputs may be used for training and later surface elsewhere. If that is not acceptable, the model has to run somewhere you control.
  • Do you need an attestation? If a regulator or a customer asks for provenance of the model behind a decision, a provider that will not supply a digest or a version makes that question unanswerable.
  • Who absorbs the operational risk? Self-hosting means you own GPU capacity, runtime CVEs and the upgrade cadence; hosting means you own the evaluation that detects a silent change.

A common arrangement is a split: call a hosted model for general reasoning, self-host a small model for a sensitive or high-volume task, and put both behind one control plane so the calls are counted and recorded in the same place. That keeps the expensive artifact off your floor without leaving the question of which model answered unanswered.

Model governance: what the record has to prove

Governance is the part of the supply chain that produces evidence, and here the evidence is concrete. A model governance program that actually holds has a model inventory — every model, adapter and dataset with its revision and digest; an approval trail for each new dependency, so "who added this" has an answer; a policy that a model- or adapter-version change is a reviewed change with a rollback, not a configuration edit; and records retained long enough to answer a question asked months later. NIST's AI Risk Management Framework organises the work into Govern, Map, Measure and Manage, which is useful because it separates the decision to accept a risk from the act of measuring it (NIST AI RMF).

The record is also what makes an incident survivable. When a model starts behaving differently, the questions are which revision was live, what data and adapters it was built from, who approved the change and what the previous baseline measured — and every one of those is a fact you either recorded or did not. The control plane is where the runtime half of that record naturally lands, and the framework that maps these controls onto an auditable program, rather than the artifact controls themselves, is the subject of mapping controls to a governance framework. This page's part of the answer is narrower and prior to it: you cannot govern an artifact you cannot identify.

Where SmartGate fits

SmartGate does not verify model weights, scan a hub or sign an artifact, and this page will not pretend otherwise. What it provides is the control plane that makes an agent's dependency on external tools visible and bounded — the runtime half of the record this page keeps asking for. SmartGate is an MCP-native algorithm gateway for token control, traffic shaping and agent audit: a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit row written as calls happen. Seven tools are exposed through it — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and each call is counted against the caller's key.

That matters to a supply chain for two specific reasons. First, a tool definition is a component you install, and a gateway is where the set of installed components is one list rather than a scatter of client configuration, which is what an inventory needs. Second, a compromised or merely misbehaving component tends to show up as a call pattern — a loop, an over-fetch, a burst — and a per-key limit bounds what it can spend before anyone reads an invoice. Plan limits move with the tier — monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, and audit-log retention of 7, 30, 90 or 180 days — and the four plans run from a free tier through Pro and Teams to Enterprise, so the pricing page is the authoritative table. The surfaces are documented rather than described here: call control on token control, the record on audit and compliance, and the connection steps in connect an MCP client.

The model supply chain inputs we pin, read from the source

The sections above describe the artifact side of an AI supply chain as a discipline. This one is the part of it we can show: the input surface the gateway resolves against is written into the repository rather than fetched from a vendor at run time. The reading is of origin/main, on 2026-10-08.

The model manifest is a file, not a catalogue call. backend/smartgate/modules/budget_guard/model_prices.json is a fixed set of model entries, and each carries its input and output cost per token and its context bounds: deepseek-chat at 65,536 input and 8,192 output tokens, gpt-4o at 128,000 and 4,096, gpt-4o-mini at 128,000 and 16,384, claude-3-5-sonnet-20241022 at 200,000 and 8,192, and claude-3-haiku-20240307 at 200,000 and 4,096. It is the shape this page keeps asking for — a named list of models with a bound on each — except that here the bound is the context window and the price, and the record is a file in the repo that shows up in a diff.

Where it is read, and how the version is pinned. backend/smartgate/modules/budget_guard/algorithm.py loads that file by relative path in _load_model_prices() — Path(__file__).parent / "model_prices.json" — and memoises it in a module-level _model_prices cache through _get_model_prices(), so the numbers a cost call sees come from the checked-in file rather than a live rate card. That is a provenance property, the narrow kind this page is about: the version you tested is the version in the tree, and changing what a model costs or how many tokens it accepts is a code change that goes through review rather than a vendor-side move.

The tokenizer is pinned to the model family too. _get_count_function(model) selects the tiktoken encoding by family rather than trusting the caller: gpt-4o and deepseek names take tiktoken.get_encoding("o200k_base"), other names go through tiktoken.encoding_for_model(model), and an unknown name falls back to cl100k_base through the module's _default_encoding. The same text therefore counts the same way across the boundary of the gateway, because the encoding is chosen by the pinned catalog rather than by whoever sends the request.

What this does not claim. Pinning the model list, its prices, its context limits and its tokenizer fixes the input surface — which models exist and what they accept — and it says nothing about the weights. Nothing here verifies a digest, signs an artifact or scans a hub, which is the line this page draws elsewhere: SmartGate does not verify model weights. The value is narrower and prior: the thing the application reasons about is a reviewed list, so the supply-chain question of which model, at what limit, has an answer in the source.

How to get started

The useful first move is an inventory, because every later control reads from it, and the sequence below is ordered so each step makes the next one cheap.

  1. Write the inventory. List every base model, adapter, dataset, runtime and installed tool server, each with its source and the revision in use. This is the ML bill of materials in its simplest form, and a spreadsheet is enough to start.
  2. Replace floating references with pinned ones. Swap latest and default aliases for an immutable commit or a concrete version, and record the digest of the bytes you actually resolved. Add a local mirror for anything you cannot re-fetch reliably.
  3. Verify at load. Check the recorded hash, and a signature where one exists, before an artifact is used. Prefer a weight format that cannot execute code, and scan the ones you cannot avoid.
  4. Review the text artifacts. Treat every tool description, input schema and vendor prompt template as code: read it before you connect it, pin the server, and diff the metadata on upgrade.
  5. Bind the runtime. Put one metering and audit layer in front of tool calls so the installed set is one list and a runaway component is capped.
  6. Make a version change a reviewed change. Require a held-out evaluation against a baseline before a model or adapter version moves, and keep a rollback to the last verified revision. If you route tool calls through a gateway, start free and watch one call end to end before trusting the limits.

Frequently Asked Questions

Is a model supply chain the same as a software supply chain?

No, it is a superset. It includes the ordinary software chain — packages, base images, build tools — and adds artifacts that behave differently: model weights that are opaque binaries rather than readable source, training and fine-tuning data that is a moving input rather than a fixed dependency, and tool definitions and prompt templates that are text a model reads as instructions. The familiar controls still apply; they are just not sufficient on their own.

Can we trust a model because it comes from a large public hub?

A hub is a distribution channel, not a guarantee. OWASP notes that model cards offer no guarantees about a model's origin and that an attacker can compromise a supplier account or create a lookalike repository. The workable rule is to trust a specific artifact you have verified, not a hosting location: pin the revision, record the digest, and check it, because the same hub serves verified and unverified uploads through the same interface.

Does fine-tuning carry the same poisoning risk as pretraining?

The risk is at least as real and cheaper to trigger. Poisoning a web-scale pretraining corpus takes access to the corpus, but a small poisoned fine-tuning set or a single malicious adapter applied to a clean base model can install behavior while every check on the base weights still passes. The persistence result is the sobering part: backdoored behavior can be trained to survive normal safety fine-tuning, reinforcement learning and adversarial training.

What is the smallest integrity check worth doing?

Record a hash of every artifact you install and verify it before use. It costs almost nothing, it detects the most common real failure — the artifact that changed between the one you tested and the one you deployed — and it is the prerequisite for every larger control, because a signature you cannot check against a known value proves nothing.

Limitations

This page is a control framework for the artifact side of an AI supply chain, not a scanner and not a benchmark: it names no product's detection rate and ranks no format or hub. The categories above overlap in real systems — a single runtime may load weights, apply an adapter and read tool metadata in one process — and the right split between hosted and self-hosted depends on data sensitivity and how a specific provider versions its models, neither of which this page can decide for you.

Integrity checking proves that an artifact is the one you recorded; it does not prove the artifact is safe, and provenance documentation is only as trustworthy as the party that produced it. A held-out evaluation reduces the chance of shipping a poisoned change without making it impossible, and no single control here removes the risk that a trusted source is compromised. This page also makes no claim about any specific product's behavior, including ours — SmartGate's role is described narrowly and only where this project records the facts.

The external descriptions above are each source's own published wording, read at the linked pages, and they describe scope rather than quality. Where a source names a technique without a published measurement, this page supplies no number.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections (0 abstentions, 7 no-slice verdicts): rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section is written from external, linkable sources, which is the house rule for an unpinned section.

Product facts were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-checked against the live pricing page on 2026-10-04; the demand figures are this project's own measurement. Every external statement quoted above is taken from the URL cited beside it. No code, batch fingerprints, auction data or internal hosts are transcribed.

The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.