Supply Chain Security for AI Model Artifacts

Model artifacts evade traditional supply chain defenses.

Staff Writer · · 10 min read
Cover illustration for “Supply Chain Security for AI Model Artifacts”
Security Posture · October 3, 2026 · 10 min read · 2,270 words

AI model artifacts are not passive code sitting on a shelf waiting to be read. A library dependency is a known quantity: you can read its source, diff its versions, and trace what it touches. A model checkpoint is a few gigabytes of numbers wrapped in a serialization format, and there's no way to eyeball a weight matrix and know what it does.

Traditional supply chain security rests on four assumptions: dependencies form a graph, versions are stable, builds are reproducible, and the underlying files are readable as text. AI artifacts break all four.

Start with reproducibility. Research from Cesarano and Monperrus at KTH shows that trained artifacts cannot be reproduced bit-for-bit. A hash tells you a file hasn't changed since it was hashed. It tells you nothing about whether the file was ever legitimate in the first place. Call this the verifiability gap.

Next comes versioning. The system just behaves differently. Call this the versioning gap.

Then there's observability. Call this the observability gap.

Last is traceability. Software lineage tools like SLSA and in-toto were built around a simple mental model: one parent, one build step, one deterministic output. AI lineage looks nothing like that. A single deployed model might descend from multiple base models, multiple fine-tuning runs, and multiple organizations, none of it reducible to a clean chain of custody. These lineage graphs are multi-parent, non-deterministic, and cross-organizational, and neither SLSA nor in-toto has a way to represent that structure.

The attack surface compounds the problem further. A 2025 guide from Glacis maps the AI supply chain across five distinct layers, data, models, libraries, infrastructure, and tooling, each with its own independent way to be compromised. Each of those is its own software project with its own dependency tree, and a compromise in any one of them reaches the final output.

The scale of what's running unexamined is hard to overstate. That's functionally the same as running software from an anonymous developer with no code review, except the file in question is a few gigabytes of weights nobody can read.

The industry has started to name this formally. OWASP ranks supply chain vulnerabilities as LLM03 in its Top 10 for LLM Applications, and MITRE ATLAS catalogs known attack techniques specific to AI systems. These frameworks exist because the risk has already produced incidents, not because someone anticipated a theoretical problem.

Those incidents are documented. In February 2024, roughly one hundred models on Hugging Face were found to be carrying malicious payloads, some of which could execute arbitrary code on any machine that loaded them. Back in December 2022, the PyTorch ecosystem itself was hit: a malicious package called torchtriton exfiltrated environment variables and credentials from machines that installed it. The frameworks sitting underneath nearly every production AI system carry the same vulnerability.

Diagram: Four Gaps Where AI Artifacts Break Traditional Supply Chain Security. Visualizes: Visualize four named gaps that explain why AI model artifacts defeat traditional software supply chain security: (1) the Verifiability Gap — a hash confirms…

Pickle-format deserialization and conjunctive metadata attacks as the two primary artifact-level exploit mechanisms

Two distinct ways to compromise a model artifact exist, and they call for two distinct kinds of defense. Treating them as one problem means one of them goes unaddressed.

The first is deserialization through Python's pickle format. A large share of models distributed today are stored in pickle, and pickle was never built with security in mind: loading a pickle file can execute arbitrary code embedded inside it. Opening the file is the attack.

A backdoor can sit inside a model without touching its normal performance. A team can run their usual accuracy benchmarks, see numbers that look fine, and ship a model that behaves exactly as poisoned as it was before testing. Accuracy testing was never built to catch this, and it doesn't.

Most model scanning tools available today weren't built to catch this consistently either. A scanner tuned to catch one known signature in one known format misses formats it wasn't built for, and new formats show up constantly.

The second exploit path is entirely different, and it's the one that's easy to miss if you're only checking weight hashes. In a conjunctive attack, the model weights themselves stay completely untouched. What gets tampered with is the wrapper code and metadata sitting around the weights, before the whole bundle gets reuploaded to a public platform under the appearance of an official release. But the wrapper logic running around those weights at inference time has been altered, and it deterministically changes how the model behaves in production.

This is the exploit class that defeats weight-hash verification by design. If a control only inspects the weight file, a conjunctive attack sails through every check that control can run. Catching it requires a separate verification step aimed specifically at the wrapper and metadata layer in addition to the weights inside it.

Detecting both of these exploit types reliably depends on watching what a model actually does when it runs, beyond what it looks like on disk. Within each phase, a model's interactions with the host system follow highly structured, predictable patterns. A loading phase that suddenly opens network connections, or an inference phase that starts writing files it has no reason to write, is a strong signal something is wrong, regardless of what format the file was stored in.

This lifecycle-aware approach was tested at real scale: it caught every evaluated attack class across tens of thousands of real-world model artifacts pulled from the Hugging Face Hub, a set of proof-of-concept exploits drawn from CVEs, and 334 models from a state-of-the-art benchmark dataset, all while keeping its false-positive rate close to zero. That result matters because it shows the lifecycle-phase approach generalizes across formats and frameworks in a way signature-based scanning never could.

Diagram: Two Exploit Paths, Two Distinct Defenses. Visualizes: Show two parallel attack paths against a model artifact and the specific control that addresses each.

The data layer as the most consequential and least defended stage of the AI supply chain

Training data poisoning sits upstream of every other control in this piece, and that position is what makes it the most consequential attack vector in the entire AI supply chain. A backdoor planted in training data gets baked into the model itself before any artifact exists to sign, scan, or hash. Every control downstream of that point, model signing, format verification, lifecycle monitoring, becomes a way to manage the damage rather than a way to stop it from happening.

What makes this threat counterintuitive is how little poisoned data it actually takes. The 2025 International AI Safety Report found that poisoning efficacy depends on the absolute number of poison samples rather than the percentage of the dataset those samples represent. A trillion-token training run isn't safer from poisoning just because it's trillion-token; the attacker only needs to get a fixed, small number of documents into the mix, regardless of how big everything else is.

Web-scraped data makes this easy to pull off at scale. Glacis documents a pattern called indirect data poisoning, where adversaries build web pages full of carefully crafted content, designed specifically to influence any model that later trains on a crawl of the open web. The attacker doesn't need access to anyone's training pipeline. They just need their page indexed somewhere a crawler will find it.

Most AI teams, when asked basic questions about their own training data, don't have answers. What's the licensing status of the source material? These are the provenance record a team needs in order to even begin assessing whether its training data was compromised, and most teams simply don't keep it.

Regulation has started to catch up to this gap faster than most engineering practice has. Glacis notes that the EU AI Act requires data-governance and management practices addressing data origin and collection for AI systems that fall under its high-risk category. The legal bar is now higher than what most teams have actually built.

Synthetic data adds a feedback loop on top of all this. Each generation of synthetic data potentially compounds whatever distortions existed in the generation before it, and that chain is harder to audit than a dataset built from original source material, because there's no clean point where a human curated the content directly.

Some detection tools do exist for this layer, but they require knowing what pattern to look for. For the messier problem of open web-scraped corpora, the practical mitigations are blunter: statistical outlier detection, filtering by source reputation, and deduplication to catch content that's been seeded in multiple places to amplify its influence.

None of these are substitutes for provenance tracking. They're mitigations layered on top of a process that, for most organizations, doesn't yet document where its training data came from. The signing and SBOM controls covered next are built to address that absence once an artifact exists.

The controls that address artifact integrity at the registry and download stage

Once a model artifact exists and reaches a registry, three controls form the minimum baseline for checking it before it enters a pipeline: cryptographic signing, safe serialization formats, and hash verification. Each one covers a different piece of the attack surface, and none of them, alone, is enough.

Cryptographic signing answers a specific question: who produced this artifact? The OpenSSF Model Signing specification, known as OMS, defines a signature format built on Sigstore bundles, applied over file-level hashes of an entire model directory. It supports both keyless and key-based signing flows, so teams can choose the approach that fits their existing identity infrastructure.

The keyless flow works like this: a publisher authenticates through an OIDC identity, such as a GitHub Actions identity tied to a specific workflow. On the consuming side, a team verifies the artifact against that public log rather than having to manage its own set of trusted keys.

This solves identity, but it has a specific limit. Signing tells a team who made something. It doesn't tell them whether the copy sitting in front of them took a clean path to get there, and even where signing is fully implemented, enforcement on the actual download path tends to stay a manual step unless a team builds automation around it specifically.

SafeTensors addresses a different problem entirely: the pickle deserialization risk covered in the previous section. SafeTensors is a serialization format built specifically to block code execution during deserialization. It closes off the arbitrary-code-execution path at the file format level rather than trying to detect it after the fact. Making SafeTensors the required format for anything pulled from a public registry removes the deserialization attack class for that artifact. This is a preventive control: the attack simply can't happen in that format, instead of being caught after it starts.

Hash verification covers a third, narrower question: has this specific file been modified since it was published? But hash verification has a blind spot that matters: it says nothing about the conjunctive attack described earlier. A file hash confirms the weight file itself wasn't touched. It says nothing about whether the wrapper code and metadata sitting around those weights were tampered with before the whole bundle got reuploaded. Hash checks and format controls need to run alongside wrapper verification, not instead of it.

An AI SBOM rounds out this stage by giving a team a way to assess impact across every artifact affected. If a base model turns out to contain a backdoor, an AI SBOM shows which downstream models and applications inherited that risk by building on top of it. Without that inventory, finding one compromised base model doesn't tell a team anything about how far the damage spread through everything built on top of it.

The gap here is in how SBOM data actually gets used. An SBOM that nobody's pipeline reads provides no protection at the moment it would matter.

Model cards fill a different role again: they're a signal that still requires verification. A model card documents lineage, training data, intended use, known limitations, and evaluation results, and gives a team a baseline for what to expect from a model's behavior. A model that claims to be a fine-tune of a known architecture but ships with no model card from a recognized organization should be treated as untrusted on that basis alone. But model cards can be falsified just as easily as any other metadata, so they need corroboration from the signature and hash checks covered above. They're a starting point for trust that still requires verifying it.

Hardening the CI/CD pipeline as the primary enforcement point for model artifact policy

Everything covered so far, signing, format checks, hash verification, SBOM tracking, model card review, only functions as real protection if something actually enforces it before an artifact reaches production. That enforcement point is the CI/CD pipeline. It's the one stage in the entire lifecycle where an automated gate can block a compromised artifact outright, no matter where upstream the compromise was introduced, whether that's poisoned training data, a tampered wrapper, or a registry serving a backdoored release.

That position in the lifecycle is also what makes the pipeline itself a target. The LiteLLM incident from early 2026 is the clearest evidence of this: attackers went after the pipeline instead of the model artifact. They compromised a Trivy dependency running inside LiteLLM's own CI/CD pipeline, used that foothold to steal PyPI publishing credentials, and uploaded backdoored releases directly to a public package index, bypassing the normal release process. The pipeline wasn't a bystander in that attack. It was the point of entry.

That's why pipeline hardening has to extend past checking artifacts for signs of tampering, into checking the pipeline's own dependencies, credentials, and build tools with the same scrutiny applied to the model artifacts it processes. A registry with 500,000 unaudited models and a pipeline pulling from it daily cannot rely on a person checking each download by hand. The gate has to run automatically, every time, or it fails as a gate.

Sources

  1. AI Supply Chain Security: Securing Models, Data & Infrastructure (2026)
  2. AI Supply Chain Security Guide 2026 — GLACIS
  3. Lifecycle-Aware Dynamic Analysis for Secure ML Model Execution
  4. The Grand Software Supply Chain of AI Systems
  5. Conjunctive Poisoning in AI Supply-Chain Applications
Filed underSecurity Posture

More in Security Posture