An Unauthenticated KV Cache Server Puts Code Execution on Your Cluster Network

CVE-2026-105192 lets one unauthenticated ZeroMQ message run code as the LMCache cache process, and no fixed release exists. The containment is the bind address.

Netics editorial card showing a key-value cache server listening on a cluster network port in front of a model server.
Netics editorial card on the LMCache cache server: one ZeroMQ port, no authentication, and no fixed release to install.

TL;DR

  • LMCache is a caching layer for large language model servers that keeps the attention key-value cache outside the model process, so a long prompt or a restarted engine does not pay for the same prefill twice.
  • JFrog disclosed CVE-2026-105192 on 7 October 2026: a flaw rated 9.8 that lets a single unauthenticated message to the cache server run code as the user the LMCache process runs as.
  • The transport opens a ZeroMQ socket with no authentication, and one message type is decoded with Python's pickle while the server is still reading the arguments of the request.
  • Versions 0.3.9 through 0.5.5 carry the flaw, along with the 0.5.6 release candidates and the development branch. No fixed version exists.
  • Exposure follows one setting. The server binds to localhost unless an operator passes a routable address, and the project's own Kubernetes deployment runs the daemon with hostNetwork enabled, which listens on every interface of the node.
  • Official container images run the cache process as root, so code arriving on that socket starts with the privileges of the process.
  • The control available to an operator today is placement: keep the transport port on the loopback interface or on a trusted cluster network, and treat any host that can open a connection as a host that can run code.
  • For a team serving models with vLLM, the flaw turns the KV cache into an inventory question inside the inference stack: which cache servers exist, where they listen, and who can reach them.

A cache server that answers any caller on the network

The CVE record is unusually specific about the mechanism. LMCache's multiprocess mode, which the documentation also calls distributed mode, "opens an unauthenticated ZeroMQ ROUTER so worker processes can register and share KV cache blocks". Messages on that socket are msgpack, and one extension code is handed to a deserialiser that calls Python's pickle library while the server is still decoding the arguments of the incoming request, before the handler for that message type runs at all.

Official LMCache architecture diagram from the project's blog, comparing in-process offload with four separate KV buffers against multiprocess mode with one unified KV pool in host memory.
Official LMCache project diagram (blog.lmcache.ai, published 3 April 2026, retrieved 2026-10-07): the unified pool that produces the performance gain is the same object the cache server exposes over the network.

That ordering is the whole vulnerability. A pickle payload executes during decoding, so the code runs before anything checks what kind of message arrived. One ZeroMQ DEALER message to the transport port, which defaults to 5555 in the JFrog record, is enough. The record states the outcome plainly: code runs as the user the LMCache process runs as, and on the project's official container images that process runs as root.

The affected range starts at 0.3.9, the release that shipped the multiprocess ZMQ transport and the pickle-backed extension in October 2025, and runs through 0.5.5, the current stable release. The 0.5.6 release candidates and the development branch carry it as well. There is no fixed version to install, which is the part that changes the operational conversation.

The documented deployment is the exposed configuration

The exposure test is short. A cache server that binds to the local machine cannot be reached from another host, and that is the default. It becomes reachable when an operator passes a routable address, which is exactly what multi-node deployments do so that peers can share cached blocks.

The project's Kubernetes guidance walks into that configuration on purpose. LMCache's deployment page describes a DaemonSet with one cache server per node, shared by several vLLM pods, and records that the DaemonSet uses hostNetwork so the model pods can find the server through the node's own address. Combined with a server that listens on all interfaces, that is a cache port reachable by anything that can route to the node. The same page confirms the transfer mode that matters here: when the engine-driven path is loaded, KV transfers between server and workers use a pickle-based path by default.

Nobody in that design made an obvious mistake. The mistake is the sequence. A performance feature that started as an in-process buffer became a shared service, the service acquired a port, and the port arrived without authentication. Our own reading of KV-cache pressure as a routing signal covered the day that shared cache started deciding where requests land; this disclosure is the other half of the same architecture, and it is the half that needs a network control.

Official LMCache benchmark chart from the project's blog, showing average time to first token reduced 13.6 times and average decoding speed increased 3.8 times with multiprocess mode.
Official LMCache project chart (blog.lmcache.ai, published 3 April 2026, retrieved 2026-10-07): the measured case for a shared cache layer, reported by the project behind it.

Why the KV cache moved out of the model process

The reason teams run this at all is measurable, and LMCache published the numbers itself. In multiprocess mode the cache runs as an independent service that registers serving engines at runtime and puts their KV blocks into one host-memory pool instead of a buffer per process.

The published benchmark used Qwen3-235B-A22B-Instruct-2507-FP8 across eight H100 GPUs with a multi-turn chat workload. Mean time to first token fell from 3.98 seconds with in-process offloading to 0.29 seconds with multiprocess mode, and mean decoding speed rose from 9.81 to 37.47 tokens per second. Those are the numbers that justify the extra process, the extra port and the extra operational surface.

The architecture diagram the project published alongside the benchmark shows the same shift. In-process offloading gives every data-parallel rank its own cache slice, and multiprocess mode replaces those slices with one pool the server owns. A single pool is what makes the speedup possible, and it is also what makes one reachable port interesting to somebody else.

Sizing follows from the same numbers. The documentation recommends allocating as much host memory to the L1 cache as remains after the operating system and the model server, with a documented example of 60GB and a benchmark run at 400GB. A pool that size sitting on a node is worth an inventory entry on its own, before anyone discusses authentication.

What an operator can do while no patch exists

Containment here is network work, and it has to be honest about its limits. JFrog's guidance is to keep the multiprocess server off a routable address and to hold its port on the local machine or on a trusted cluster network. A firewall that restricts who can reach the port lowers the risk, and the CVE record is explicit that it does not remove it, because any host that can still open a connection can run code.

Official LMCache serving architecture diagram from the project's blog, showing attention layers per data-parallel rank and sharded expert layers behind a load balancer.
Official LMCache project diagram (blog.lmcache.ai, published 3 April 2026, retrieved 2026-10-07): several serving processes share one cache server, which is why one unauthenticated port is enough to matter.

Two details sharpen the priority. First, a process that runs as root in the container means arriving code inherits the privileges of a cluster-relevant process. Second, the CVE record offers no way to tell whether a server has already been attacked, and JFrog's advisory does not add one. Log review therefore starts from the socket rather than from an indicator: who connected to the cache port, from which address, and when.

There is also a version trap worth naming. A related flaw in vLLM, tracked as CVE-2026-105756, was fixed in version 0.30.0 and concerned a malformed cache salt value crashing the engine on deployments using the LMCache multiprocess connector. That one is patched. The cache server flaw is not, so an upgrade of the model server closes one issue and leaves the other exactly where it was.

Where the cache server sits in your inference stack

The practical output of this disclosure is an inventory with four columns. Does a cache server run at all, and in which mode. Which address does it bind to, since the default is safe and the routable one is not. Which network can reach its port, and whether that network contains anything you would not trust to run commands as root. And which container image the process comes from, which decides whose privileges those commands inherit.

Teams that cannot answer those questions have a bigger problem than the CVE: the inference stack has grown a component class that nobody owns at the network layer. The response that fits is the same one we apply to any security audit and hardening pass, where every listening service is named, its bind address recorded, and its reachability decided rather than inherited from a container-runtime default.

Netics editorial checklist card of five questions that decide whether an LMCache cache server is reachable from another host.
Netics editorial checklist built from the CVE record and the project's deployment documentation: the exposure test for a cache server in multiprocess mode.

That is also the durable lesson. Cache layers, vector stores and model routers now sit inside production networks because they make inference fast, and each of them is a service with a port. The ones that arrive with authentication in place stay boring. The ones that arrive as a local buffer and later learn to listen need somebody to notice the transition, and the CVE tracker is a late place to notice it.

Sources

Source: CVE-2026-105192, LMCache Unauthenticated RCE in multiprocess mode via pickle deserialization — cveawg.mitre.org/api/cve/CVE-2026-105192, CNA JFrog, published 7 October 2026, retrieved 2026-10-07 (the unauthenticated ZeroMQ ROUTER, msgpack messages, extension code 1 passed to DeviceIPCWrapper.Deserialize calling pickle.loads before the handler runs, the default transport port 5555, the single unauthenticated ZMQ DEALER message, root in official container images, the localhost default and the routable --host option, affected from 0.3.9, CVSS 3.1 9.8, CWE-502 and CWE-306, the finder credit to Yuval Moravchick of JFrog Security Research, and CISA's ADP assessment of Exploitation none, Automatable yes and Technical Impact total). Source: LMCache multiprocess deployment documentation — docs.lmcache.ai/mp/deployment.html, retrieved 2026-10-07 (the standalone server command with --l1-size-gb, --eviction-policy, --max-workers and --port, the Docker --network host requirement, the DaemonSet and Deployment pattern with one cache server per node, the DaemonSet using hostNetwork true and discovery through status.hostIP, GPUs not requested in the DaemonSet, the pickle-based engine-driven transfer path, and the L1 sizing guidance from 60GB to 400GB). Source: LMCache's New Architecture Boosts MoE Inference Performance by 10x — blog.lmcache.ai, published 3 April 2026, retrieved 2026-10-07 (multiprocess mode as an independent service that registers serving engines at runtime, the unified host-memory KV pool shared across data-parallel ranks and independent instances, the Qwen3-235B-A22B-Instruct-2507-FP8 benchmark on 8 H100 GPUs, mean time to first token 3.98s to 0.29s, mean decoding speed 9.81 to 37.47 tokens per second, and the official architecture and benchmark diagrams). Source: Unpatched Critical LMCache Flaw Lets Unauthenticated Attackers Run Code Remotely — thehackernews.com, published 7 October 2026, retrieved 2026-10-07 (the disclosure date of 7 October 2026, the 9.8 severity score, the affected range from 0.3.9 through 0.5.5 plus the 0.5.6 release candidates and development branch, the absence of a fixed version and of a project security advisory, the guidance to avoid a routable address and to keep the port locally reachable, the firewall limitation, the absence of any way to determine prior attack, and the related vLLM fix CVE-2026-105756 rated 6.5 and fixed in 0.30.0).

Source: JFrog's CVE record for CVE-2026-105192 — cveawg.mitre.org, 7 October 2026; LMCache's multiprocess deployment documentation and architecture post — docs.lmcache.ai and blog.lmcache.ai, retrieved 2026-10-07; The Hacker News report on the disclosure — thehackernews.com, 7 October 2026. All retrieved 2026-10-07. Figures: the unified-pool architecture diagram, the serving architecture diagram and the benchmark chart are LMCache project images from blog.lmcache.ai, plus two Netics editorial diagrams rendered from the same sources.