On March 24, 2026, production systems running LiteLLM started falling over. Processes multiplied until CPUs sat at 100%, and containers were killed for running out of memory. The cause turned out to be a bug in someone else’s malware.
Two versions of LiteLLM, 1.82.7 and 1.82.8, had been published to PyPI with a credential stealer inside. A mistake in the payload made it spawn processes without limit, and that noise is what got it noticed. A researcher at FutureSearch spotted it while testing a Cursor MCP plugin that pulled in LiteLLM as a transitive dependency. Without the bug, the stealer could have run quietly for much longer.
The incident is worth studying closely, because the gateway is the part of the AI stack most teams trust without a second thought. This piece covers how the attack worked, why a gateway is so expensive to lose, and the controls worth putting in place now, no matter which gateway you run.
The group behind the attack, tracked as TeamPCP, never went after LiteLLM directly. It went after Trivy, a popular open-source vulnerability scanner that thousands of projects run in CI through a GitHub Action.
In late February, someone exploited a misconfigured workflow in Trivy’s repository and stole a bot account’s access token. Aqua Security rotated credentials, but not all at once, and a working token survived. On March 19, TeamPCP used it to force-push 76 of the 77 release tags of trivy-action to malicious commits. Every pipeline that referenced the action by tag now ran the attackers’ code. The real scan still ran afterward, so the output looked normal.
LiteLLM’s CI was one of those pipelines. The poisoned scanner ran in a job that also had access to the project’s PyPI publishing token, and it took the token. Five days later, the attackers used it to publish the two malicious releases. CloudSEK summed up the chain as “one unrevoked token, three tools deep”.
The payload was careful work. Version 1.82.7 hid the malicious code in the proxy server module. Thirteen minutes later, 1.82.8 added a .pth file, which Python executes every time the interpreter starts, whether or not anything imports LiteLLM. It swept each machine for more than 50 kinds of secrets, including cloud credentials, SSH keys, Kubernetes tokens, database passwords, and the LLM API keys for every configured provider. It encrypted the haul and sent it to a lookalike domain, models.litellm.cloud. On non-CI machines, it installed a systemd user service called sysmon.service that checked in for new payloads every 50 minutes. The code could also create privileged pods on every Kubernetes node it could reach.
How long the packages stayed live depends on whose timeline you read. CloudSEK puts it at about 40 minutes, while Trend Micro’s analysis has PyPI quarantining them roughly three hours after publication. Either way, it was long enough. CloudSEK reconstructed potential exposure across more than 2,500 organizations and about 434,000 CI/CD pipelines, and carefully notes this is exposure rather than confirmed compromise. At the time of its analysis, Trend Micro had found no public confirmation that stolen credentials were used, and the Kubernetes spreading was present in the code but not confirmed in the wild. That uncertainty is part of the problem. Most affected teams had no reliable way to know what had been taken.
Most compromised libraries leak whatever happens to be lying around on the machine. A gateway is different, because concentrating secrets is its job. A LiteLLM proxy typically holds the API keys for every model provider the company uses and the virtual keys issued to internal teams and agents. It also sees every prompt and response in plain text.
Trend Micro draws a useful distinction here. If you use LiteLLM as a library and pass credentials per request, the blast radius is smaller. If you run it as a central proxy, as most production deployments do, one compromised host exposes every model relationship at once.
Transitive dependencies widen the net. DSPy, MLflow, CrewAI, and OpenHands all pull in LiteLLM, and several of them opened pull requests the same day to pin away from the bad versions. Wiz reported finding LiteLLM in 36% of the cloud environments it analyzed. Plenty of teams that would have said “we don’t use LiteLLM” were running it anyway.
Security teams usually learn about stolen cloud keys from an alert. With LLM keys, finance often finds out first.
Sysdig’s threat research team named the pattern LLMjacking in 2024. Attackers use stolen credentials to run models on someone else’s account, often reselling the access through a reverse proxy. In its first analysis, Sysdig estimated a single case could cost the victim more than $46,000 a day. Later research put the figure above $100,000 a day for the most expensive models, and found attackers switching on models the victim had never enabled. Some checked the account’s logging configuration before using the keys, which tells you they know where defenders look.
The bill is only part of it. To a model provider, traffic sent with your keys is your traffic. If the resold access is used for phishing kits or malware, the abuse report lands on your account, and the provider may suspend it while you work out what happened. Sysdig found that at least some LLMjacking users were based in Russia, where sanctions restrict access to Western AI services, which can turn a billing problem into a compliance one. And if you run systems classed as high-risk under the EU AI Act, Article 12 expects automatic event logging over the system’s lifetime. Logs kept on an attacker-controlled host won't carry much weight when you need to show which activity was yours.
None of these is tied to one product. They apply to LiteLLM, to other open-source gateways, and to anything you buy.
1. Pin CI actions to commit SHAs, not tags. The Trivy compromise worked because tags are mutable. The attackers rewrote them, and every pipeline followed. A full commit SHA cannot be moved.
2. Keep scanners away from publishing secrets. The scanner never needed the PyPI token to do its job. Run scanning and publishing in separate jobs, and give the publish job nothing else. PyPI’s trusted publishing replaces long-lived tokens with short-lived ones, which shrinks the window, although it will not help if malicious code runs inside the publishing job itself.
3. Pin dependencies by hash, and wait before adopting new releases. pip install — require-hashes refuses anything that does not match your lockfile. A short delay before new versions reach production, which uv supports with — exclude-newer, lets the rest of the ecosystem trip over a bad release before you do.
4. Restrict egress from gateway hosts. A gateway needs to reach a known list of provider endpoints. The stolen data went to a domain set up for the attack, and an outbound allowlist would have blocked it.
5. Put a hard ceiling on every key. Set per-key and per-team spend caps in the gateway, and use provider-side limits and alerts where they exist. Prefer short-lived, narrowly scoped keys for apps and agents. A capped key turns an open-ended loss into a known one. Test that the caps hold under concurrent requests, not just that the dashboard displays them.
6. Keep the audit log outside the gateway’s blast radius. Ship request logs to a separate account or an append-only store that the gateway host cannot modify. If the host is compromised, those logs are how you prove what was yours.
7. Treat resource anomalies as security signals. This attack was caught because of a fork bomb. An unexplained CPU spike on a gateway host, or a sudden jump in token usage on one key, deserves the same attention as a burst of failed logins.
If you think one of the affected versions reached your systems, these checks are a starting point. Run them in every environment that might have it, including each virtualenv, container image, and CI runner, and run the systemd checks for each user account. # Is LiteLLM installed, which version, and what pulled it in?
pip show litellm # check “Version” and “Required-by”
# The .pth file used by 1.82.8
find / -iname “litellm_init.pth” 2>/dev/null
# Persistence left on non-CI machines
systemctl —— user status sysmon.service
ls -la ~/.config/sysmon/
If anything shows up, removing the package isn't enough. Rotate every credential the process could have read, including cloud IAM, repository tokens, Kubernetes service accounts, SSH keys, registry logins, and LLM API keys. Then review provider usage for activity you do not recognize. In May, the source code of Shai-Hulud, an offensive framework attributed to TeamPCP, appeared on GitHub under a permissive MIT license, with a README signed by the group. The tooling behind this kind of campaign is now available to anyone who wants it.
LiteLLM’s maintainers disclosed publicly and moved quickly, and open-source gateways remain a reasonable choice for many teams. The incident changes the question worth asking about any gateway. Code quality matters, but the more useful question is what a compromise would cost you and who would notice first. If the honest answer is “our finance team, next month,” the seven controls above are a good place to start.
Disclosure: I run a company that builds AI gateway and governance software, so this is a problem I think about professionally, and yes, with some bias. I’ve kept this piece about the attack and the controls any team can apply, regardless of which gateway they run.
What the LiteLLM Breach Teaches Us About Securing LLM Gateways was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.