AI Coding Agents Now Ship Software. Plugin4Shell Shows Nobody Checks the Package
A zero-click plugin swap bypasses SHA pinning in four major AI coding agents. The fix has to ship in the agent itself, and updating is only the complete remedy where one exists.
TL;DR
- On September 17, Air Security disclosed Plugin4Shell, a zero-click, high-severity remote code execution flaw affecting all four major AI coding agents: Anthropic's Claude Code, OpenAI's Codex, GitHub Copilot, and Google's Gemini CLI.
- The mechanism is a plugin SHA-pinning bypass: agents check out the pinned commit of a reviewed plugin but never verify the code that lands actually matches it. An attacker who controls the plugin's repository can swap it, and the pin still looks honored.
- Because background auto-update is the default in Claude Code and Codex, no user click is required. The malicious plugin runs with the same permissions as the developer who installed the trusted one.
- As of disclosure, Claude Code (2.1.179) and Codex (0.146.0) are patched. Copilot has no fix, and Google is deprecating Gemini CLI rather than patching it. No CVE had been assigned, and no vendor advisory was published.
- The practical read for a team: update the two agents that have fixes, treat unpatchable ones as an inventory problem, and stop assuming a version pin is proof of integrity.
Air Security's Plugin4Shell disclosure, published September 17, 2026, is the first supply-chain vulnerability this blog has seen that attacks the AI agent distribution layer rather than the model or the agent runtime itself. It is worth reading carefully because it is almost the inverse of the usual agent-security story: the plugin was reviewed, the version was pinned, the marketplace was trusted — and none of that mattered.
The four affected agents install plugins from online marketplaces. The marketplaces lock each plugin to a specific reviewed commit using SHA pinning. That is the control that is supposed to ensure a reviewed plugin cannot change underneath you. Air Security found that each of the four agents fetches the pinned version but never verifies that the code it ends up running actually matches it. Someone who controls a plugin's code repository can therefore swap what the agent installs, even when the agent is locked to a reviewed version. Background auto-update, which is on by default in Claude Code and Codex, turns the swap into a zero-click event: no user action at all is required, because the update path delivers the substituted code.

The scope claim deserves attention. Air Security's researchers built working proof-of-concept exploits against all four agents in May 2026, disclosed the flaw to each vendor the following month, and published on September 17 with two of four fixes shipped. Anthropic fixed Claude Code in 2.1.179. OpenAI fixed Codex in 0.146.0, with its own public fix describing the same underlying behavior. Microsoft, notified in June, had not shipped a fix at disclosure, leaving GitHub Copilot users with no patch. Google has not patched Gemini CLI and is instead deprecating the tool and pointing users to Antigravity, which Air says is not reachable by this attack. As of September 18, no CVE identifier had been assigned and none of the four vendors had published a security advisory, according to The Hacker News.
The pin you rely on is a statement of intent, not proof of integrity
The load-bearing point is not the exploit mechanics; it is where the trust model broke. The whole point of SHA pinning is that once a plugin passes review, it cannot change without the developer knowing. Air Security's finding is that every one of the four agents checks out the pinned commit without verifying the checkout landed there. The reviewed snapshot and the executed snapshot were assumed to be the same thing, and nothing in the agent's code ever checked that assumption. The primary source is more precise: for Claude Code, Codex and Copilot, Git resolves a branch named exactly like the pinned SHA as a ref rather than as the commit object; Gemini CLI’s variant redirects git checkout FETCH_HEAD through a default branch named FETCH_HEAD, even though the pinned commit was fetched.

This is a category of failure that the SBOM discussion on this blog flagged in advance: inventory is not provenance, and provenance is not verification. An SBOM tells you what a package claims to contain. A pin tells you which version was reviewed. Neither tells you what actually executed. Plugin4Shell is what happens when the industry's standard answer to plugin integrity — pin the reviewed commit — is itself never verified at execution time.
There is a second structural fact: because each agent checks the pin on the user's own machine and not at the marketplace, no marketplace can close this for users. The fix has to ship in the agent itself. That means the two unpatched agents are not a remediation backlog a team can route around with better registry hygiene; until the vendor ships a fix, the agent's pin is simply not doing the job its design intended.
Two fixes, two open doors, and a deprecation that is not a patch
The patch matrix is the uneven part of this disclosure, and it is worth stating plainly rather than framing it as a vendor contest.
Claude Code fixed in 2.1.179. That is a complete fix where it exists: updating is the only remedy that fully closes the gap. Codex fixed in 0.146.0, and the pull request describes the same underlying behavior the update closes. For Copilot, Microsoft was told in June and had not shipped a fix at disclosure. For Gemini CLI, Google is deprecating the tool rather than repairing it — which is a support decision, not a security one. Users who rely on CLI parity across a fleet do not get a patch from that; they get a migration requirement, and the interim exposure remains.

Screenshot of the OpenAI Codex fix pull request #34644, "Verify Git plugin SHA checkouts", which describes Git resolving a requested commit SHA as a branch name; retrieved September 26, 2026.- The other load-bearing fact from the CSA research note is that no CVE had been assigned as of disclosure, three months after vendors were first notified. That is a disclosure-process gap working against defenders: no CVE means no central tracking, no vendor advisory to anchor a fleet-inventory ticket, and no distributed signal. The baseline discipline here is mechanical: know which agents your developers run, which marketplace plugins are installed, and whether those agents update the pinned code without verification.
- The full status picture, as of September 18, 2026: Claude Code fixed in 2.1.179; Codex fixed in 0.146.0; Copilot unpatched; Gemini CLI deprecated rather than fixed. That is the matrix in one line, and it is the first thing a defender should record.
What a team running these agents can verify today
None of the actionable checks depend on knowing exactly how the swap works. That is the useful property of this disclosure.
- Update the two agents with fixes: Claude Code to 2.1.179 or later, Codex to 0.146.0 or later. This is the only complete remediation that exists.
- Treat Copilot and Gemini CLI as exposed until a fix ships: inventory which plugins are installed and from which marketplaces, and weigh disabling nonessential plugins or turning off auto-update.
- Audit plugin-marketplace usage across the team. Map which agents are in use, which marketplaces they pull from, and whether installs come from sources you control. The pin was the assurance you were relying on — treat marketplace provenance as an open question.

Screenshot of the Air Security Plugin4Shell disclosure, "Zero Click RCE Vulnerability found in top 4 most popular coding agents" (air.security, published September 17, 2026); retrieved September 26, 2026.- Contain the blast radius. Because reach equals the running user's access, keep long-lived credentials and secrets out of the environments where these agents run, and prefer short-lived, scoped tokens.
There are open questions the sources do not close, and they should be named as open: active exploitation is unconfirmed; Microsoft and Google have not committed to patch timelines; whether Cursor, Windsurf, or Cline share the weakness is untested; and the sources do not say whether updating a patched agent removes an already-swapped plugin or only stops future swaps. None of those unknowns change the checklist above, which is why the checklist is the deliverable of this article.
The deeper lesson for the industry is the one Air Security chose to name: the development workflow, not just the registry, is now part of the supply-chain attack surface. Agents select dependencies and execute build commands with a degree of implicit trust the industry spent a decade learning not to extend to human contributors. Front-line teams should treat every AI coding agent as a distribution channel that needs the same update, inventory, and provenance discipline as any other package manager in the stack.
Sources
The claims in this article come from the Air Security Plugin4Shell disclosure (September 17, 2026), the CSA research note on Plugin4Shell (September 19, 2026), The Hacker News coverage (September 18, 2026), and OpenAI's Codex fix pull request. No CVE, vendor advisory, or official Microsoft or Google statement existed as of September 18, 2026, so none is cited here beyond what the named sources report.
Source: Air Security — air.security, 2026-09-17. Corroborated by Help Net Security and The Hacker News, 2026-09-18, and the Cloud Security Alliance research note, 2026-09-19.
Explore the Netics approach to agent and infrastructure security