Critical Flaw in LMCache AI Software Lets Attackers Run Code Without a Password, and No Fix Exists Yet
A critical vulnerability in LMCache, open-source software that speeds up large language model servers such as vLLM, allows an attacker to run commands on the cache server without logging in. JFrog disclosed the flaw, tracked as CVE-2026-105192, on October 7 and rated it 9.8 out of 10. It affects versions 0.3.9 through 0.5.5, the latest stable release, as well as the 0.5.6 release candidates and the development branch. No fixed version is available.
The problem sits in LMCache's multiprocess mode, where the cache runs as a standalone server and AI workers connect over the ZeroMQ messaging library. The server's socket has no authentication. One type of message is decoded using pickle, a Python format that can carry code and run it during decoding, and this happens before the message type is checked. A single crafted network message can therefore run commands with the privileges of the LMCache process. JFrog says that process runs as root on the project's official container images.
Exposure depends on one setting. By default the server listens only on the local machine, so other hosts cannot reach it. It becomes reachable when an operator sets a routable address, as multi-node deployments do. LMCache's own example Kubernetes deployment listens on every network interface. Running LMCache inside a single vLLM process does not open the port at all.