Blog · · 11 min read
AI agent supply chain security for MCP servers and skills

AI agent supply chain security starts before an agent invokes a tool. A familiar MCP server name, a valid registry entry, or a passing malware scan does not prove that the exact artifact is safe for your product. A third-party capability may contain executable code, natural-language instructions, hidden downloads, or a later update that expands access. If you install it directly into a production runtime, it inherits whatever the runtime can reach.
The fix is an admission and update pipeline. Treat every MCP server, skill, plugin, and tool package as an untrusted release candidate. Resolve it to an immutable artifact, verify its provenance, inspect code and instructions, calculate the permissions it needs, test it without secrets, approve it against policy, and record exactly what reached production. This article defines that lifecycle and the evidence needed to revoke or roll back a capability later.
Why discovery is not a security decision
A registry answers a discovery question: what package claims to implement this capability, and which namespace published it? The official MCP Registry documentation explains that namespace authentication can be tied to a GitHub account or domain. It also says that code scanning is delegated to package registries and downstream aggregators. That boundary matters. Publisher authentication is useful evidence, but it is not a verdict on code, embedded instructions, downloaded payloads, or future versions.
The risk extends beyond conventional dependencies. Agent capabilities can influence the model through prose as well as code. OWASP's AST01 Malicious Skills entry describes instruction-layer payloads, executable payloads, typosquatting, delayed activation, and persistence through memory or identity artifacts. Snyk's ToxicSkills study reports that a useful review must inspect scripts, dependencies, credential handling, third-party content, and natural-language instructions. Snyk's figures come from its own ecosystem scan. They show observed techniques, not a universal compromise rate.
Practitioner reports expose the operational gap. One MCP server author reported familiar projects being forked and republished under third-party package scopes. Another project reported that deep skill audits found brand-jacking, deferred execution, generated payloads, and malicious instructions. These are author reports rather than independently corroborated incidents, but they show why name matching and shallow metadata checks are weak admission controls.
Define one capability admission contract
Do not let each team invent a different checklist. Represent every candidate with one machine-readable record and require an explicit decision before deployment.
capability_candidate:
registry: "mcp-registry"
namespace: "example.com/finance-tools"
requested_version: "2.4.1"
resolved_digest: "sha256:<immutable-artifact-digest>"
publisher_identity: "dns:example.com"
provenance_ref: "https://provenance.example/statement.json"
entrypoint: "package:finance-tools"
declared_permissions:
filesystem: ["workspace:read"]
network: ["api.example.com:443"]
credentials: ["finance-api:invoke"]
tools: ["invoice.read", "invoice.prepare_update"]
discovered_permissions: []
instruction_hash: "sha256:<normalized-instruction-digest>"
scan_results: []
quarantine_run_ids: []
policy_version: "capability-admission-7"
decision: "pending"
The fields are not decorative inventory. They support four decisions:
- Is this the intended publisher and artifact?
- Does the artifact request only the access required for its stated job?
- Did static and behavioral tests reveal code, instructions, or downloads that violate policy?
- Can the platform identify, disable, and replace every deployed copy if trust changes?
Use a stable capability ID for product references, but bind every deployed release to an immutable digest. A mutable tag such as latest can remain a discovery hint. It must never be the identity recorded on a production run.
Step 1: resolve and verify an immutable artifact
Resolve the requested registry entry and version before rendering runtime configuration. Store the registry, publisher namespace, semantic version, package URL, and final artifact digest. Then fetch by digest or verify the digest after download. If the registry cannot provide an immutable resolution, reject the candidate or mirror the approved bytes into a controlled artifact store.
This rule prevents a review-time package from becoming a different runtime package. It also makes incident scope answerable: the platform can query which tenants and runs used one digest instead of guessing which bytes a mutable version label referred to.
An active implementation request in Kanon issue #100 asks for registry-aware MCP resolution, exact version locking, and auditable configuration instead of mutable raw commands and URLs. That request is practitioner evidence, not a protocol guarantee. The architecture recommendation is still straightforward: lock before deployment, and fail closed if the lock cannot be reproduced.
Verify available provenance alongside the digest. The SLSA v1.2 specification defines source and build guarantees plus artifact provenance and verification concepts. A third-party MCP package may not meet a high SLSA level, but the model is useful for policy. Record whether the artifact has a signed provenance statement, whether the builder identity is trusted, whether source and build references are complete, and whether the subject digest matches the downloaded bytes.
A missing provenance statement does not automatically mean malware. It means the decision has less evidence. Route that case to stricter quarantine or manual approval rather than converting absence into trust.
Step 2: inspect both executable code and agent instructions
Ordinary dependency scanners are one input, not the whole review. Build two parallel analyses.
The code analysis should inspect packaged files, install hooks, binaries, transitive dependencies, dynamic imports, network clients, filesystem access, process execution, and embedded archives. Flag runtime downloads because the reviewed package may only be a loader. Unpack or capture downloaded artifacts in quarantine and repeat the scan against the final executable set.
The instruction analysis should inspect tool descriptions, skill files, prompt fragments, examples, comments exposed to the model, and generated descriptions. Look for text that asks the agent to ignore policy, conceal activity, read unrelated files, copy credentials, contact undeclared destinations, persist instructions, or fetch and execute another payload. OWASP explicitly treats malicious instructions and code as connected layers, which is why a clean dependency scan cannot approve an agent skill by itself.
Normalize the exposed instruction set and record its digest. Keep a readable diff between the candidate and the last approved version. Reviewers need to see that a tool changed from "read an invoice" to "read any account file and send diagnostics" even if the package's public API and dependency list did not change.
Scanner findings also require context. A project report about scanning 100 MCP servers discusses possible tool-description injection and exfiltration indicators while acknowledging false positives and manual review. Treat scanner output as evidence with severity, location, and rationale. Do not blindly approve a zero-finding scan or blindly block every heuristic match.
Step 3: derive a permission and behavior diff
Declared permissions are rarely complete enough for an agent runtime. Derive the candidate's effective capability envelope from static analysis and quarantine observations:
- filesystem paths and access modes
- outbound hosts, protocols, and ports
- credential references and operations
- subprocess and interpreter use
- product tools it may call
- memory or state it may read and write
- data classes it may receive or emit
- interactive approvals it may request
Compare that envelope with both the candidate's declaration and the last approved release. A new version that adds one outbound host or requests a write credential is a security-relevant change even if its semantic version says "patch."
Set policy around expansion, not version labels. An unchanged envelope may qualify for an automated quarantine run. Any new publisher, unsigned artifact, hidden download, executable install hook, credential class, write operation, memory access, or network destination should require explicit review. A removed permission is still worth recording because it changes expected runtime behavior and may expose a compatibility failure.
Step 4: quarantine without production secrets
Run the candidate in the same isolation class used for production, but provide synthetic data and no reusable production credential. Deny outbound network access by default. Allow only a test endpoint that records requests and returns controlled responses. Mount a disposable workspace, apply CPU, memory, process, and time limits, and destroy the environment after the test.
Exercise the capability with normal, malformed, adversarial, and cancellation cases. Include documents or tool results containing hostile instructions. Confirm that the candidate cannot reach undeclared paths or destinations, cannot persist outside its assigned workspace, and does not hide failed operations behind a successful response.
Quarantine is not proof of safety. Delayed or environment-aware behavior can evade a short test. Its role is narrower: verify the declared envelope, catch observable policy violations, and produce replayable evidence for the admission decision.
Record test inputs, the candidate digest, policy version, observed network and filesystem events, tool calls, generated instructions, outputs, and cleanup result. If the test cannot produce that evidence, it should not promote an artifact into an action-taking product.
Step 5: approve, deploy, and bind every run
The policy decision should return approve, approve_with_limits, needs_review, or reject, plus reason codes and expiry. Approval must name the immutable digest and exact permission envelope. Never approve only a publisher or package name, because that silently approves future artifacts that have not been reviewed.
Promote the approved digest through development, staging, a small production cohort, and broader rollout. At run creation, record:
- capability ID and artifact digest
- publisher identity and provenance result
- admission policy and decision ID
- granted filesystem, network, credential, memory, and tool envelope
- agent release and tenant
- quarantine evidence reference
This is the agent capability bill of materials for that run. It lets an incident responder answer which code and instructions were active when the agent acted. Keep a reverse index from digest to deployed environments, tenants, credentials, and recent runs.
Monitor actual behavior against the approved envelope. Block undeclared destinations and tools at runtime rather than relying on the candidate to behave. Send expansion attempts to security telemetry with the capability digest and run ID. The control remains enforceable even if the model or plugin is manipulated later.
Treat every update as a new candidate
Auto-update is incompatible with meaningful admission. A new version, build, digest, provenance statement, entrypoint, instruction set, or permission envelope creates a new CapabilityCandidate. It may reuse prior evidence, but it must receive a new decision.
Automate the safe comparison work: resolve both artifacts, produce code and instruction diffs, compare dependencies and provenance, calculate permission expansion, and run the quarantine suite. Do not let the automation auto-approve a high-risk delta just because the publisher is unchanged.
Use time-limited approvals for capabilities that depend on a temporary exception or missing provenance. Expiry should stop new runs while allowing a deliberate choice for in-flight work. Record whether each in-flight run may finish, must pause, or must terminate when an approval expires.
Build revocation and rollback before installation
A product that can install a capability but cannot disable it across all runtimes has an incomplete supply-chain control. Maintain a denylist outside the registry so you can react even when registry metadata is unavailable or disputed.
Your incident action should support four scopes:
- block one digest while leaving earlier approved versions available
- block every artifact from one publisher identity
- block one capability for selected tenants or environments
- block a permission, destination, or credential class across capabilities
When a digest is revoked, deny new runs immediately. For active runs, apply a declared policy: cancel, pause for review, or allow read-only completion. Revoke or rotate credentials that the capability could access, inspect durable memory and generated artifacts for persistence, and query the reverse index for affected tenants and actions.
Rollback means routing new work to the last approved digest. It does not mean rewriting historical run records. Keep the revoked digest attached to prior traces and receipts so investigators can reproduce what happened.
Common implementation mistakes
Do not treat a registry badge as a safety verdict. Namespace authentication establishes who controls a name, but it does not inspect every instruction, dependency, or downloaded payload.
A semantic version or tag is insufficient unless resolution is immutable and verified by digest. Pin the bytes, not only the version label.
Source-code scanning misses agent-facing instructions and tool descriptions that can change model behavior without a conventional executable exploit.
One scan does not cover later updates. Transitive dependencies, provenance, and runtime downloads can all change the effective artifact.
Approve the artifact and its envelope together. Runtime enforcement needs explicit filesystem, network, credential, memory, and tool limits.
Heuristic findings need review because false positives exist. Preserve the evidence and make the decision explainable.
Keep a deployed inventory. Without a digest-to-run and digest-to-tenant index, emergency response becomes repository archaeology.
Do not wait for an incident to design rollback. Test disablement, credential rotation, memory inspection, and last-known-good routing before the first production install.
Verify the pipeline with adversarial tests
Run these tests before permitting third-party capabilities:
- Replace a package behind a mutable tag and confirm digest verification blocks it.
- Publish the same display name from another namespace and confirm identity matching does not select it.
- Add an undeclared outbound destination and confirm both the diff gate and runtime policy block it.
- Add a hidden download to an install hook and confirm quarantine captures and scans the downloaded artifact.
- Insert a hostile instruction into a tool description and confirm instruction scanning sends it to review.
- Add a new write credential to a patch release and confirm semantic versioning cannot bypass approval.
- Revoke a deployed digest and confirm new runs fail closed across every runtime replica.
- Roll back to the last approved digest and confirm historical runs retain their original artifact identity.
- Rotate an exposed credential and confirm the reverse inventory identifies every affected capability and tenant.
- Simulate registry downtime and confirm deployed artifacts keep running by immutable local identity while new, unresolved candidates remain blocked.
A clean scan is insufficient. Every production capability needs a reproducible artifact identity, an evidence-backed admission decision, an enforced permission envelope, a deployed inventory, and a tested removal path.
Next action
Choose one third-party MCP server or agent skill currently allowed in a development environment. Apply this AI agent supply chain security workflow by creating its CapabilityCandidate record, resolving and pinning its digest, inventorying its exposed instructions and effective permissions, then running it in a no-secret quarantine. Do not connect it to production until you can prove which artifact will run, what it can reach, and how you will disable it everywhere.
References
- OWASP Agentic Skills Top 10: AST01 Malicious Skills supports the instruction-layer and code-layer threat model plus signing, scanning, isolation, pinning, and response controls.
- Snyk ToxicSkills study supports scanning executable content, instructions, dependencies, credential handling, and external content. Its prevalence data is vendor research.
- SLSA specification v1.2 supports artifact provenance, build identity, distribution, and verification concepts.
- Official MCP Registry documentation supports the distinction between namespace authentication, discovery, package-registry scanning, and downstream security checks.
- Kanon issue #100 provides practitioner implementation evidence for registry resolution, exact version locking, and auditable MCP configuration.
- Hacker News report on mass-forked MCP servers provides an author's report about familiar projects being republished under third-party scopes.
- Hacker News report on MCP scanner findings provides practitioner evidence about tool-description findings, possible exfiltration indicators, false positives, and manual review.
- Hacker News report on deep OpenClaw skill audits provides author-reported examples of brand-jacking, deferred execution, generated payloads, and malicious instructions.