cd /news/ai-safety/the-ai-stack-security-ledger-eight-o… · home › topics › ai-safety › article
[ARTICLE · art-147266] src=provenbrief.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug

Eight of the ten worst CVEs in the first build of a new AI-stack security ledger are the same class of bug: remote code execution where software executes or obeys model output, according to the ledger compiled from NIST's National Vulnerability Database, GitHub Security Advisories, and CISA's Known Exploited Vulnerabilities Catalog. The ledger's archetype entries run from LangChain's 2023 LLMMathChain prompt-injection code execution at CVSS 9.8 through Langflow's 2024 and 2025 custom_component and /api/v1/validate/code flaws at CVSS 9.8 to LiteLLM's 2026 MCP command injection, and across the three entries CISA lists as exploited in the wild the median gap from CVE publication to confirmed exploitation is 28 days. Wiz Threat Research, which runs honeypots across LiteLLM, Flowise, LangChain, Langflow, ChromaDB, and Ollama, reports most AI tools ship unauthenticated and advises defenders to treat an unauthenticated internet exposure as compromised, while Wiz's State of AI in the Cloud report puts 90 percent of cloud environments on self-hosted AI software.

read8 min views5 publishedOct 8, 2026
The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug
Image: Provenbrief (auto-discovered)

The software layer that runs AI - model servers like Ollama and vLLM, agent frameworks like LangChain, vector databases, and MCP servers - has accumulated its own CVE record, and it keeps producing the same dangerous class of bug: remote code execution where the stack executes or obeys model output. We compile every published CVE across the AI/ML software stack into one dated ledger, each row tagged by vulnerability class, CVSS severity, affected product, time-to-patch, and whether it is known to be exploited in the wild.

The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug #

Eight of the ten worst CVEs in the first build of our AI-stack ledger are the same bug: software that executes whatever it is handed. The record runs from LangChain's 2023 prompt-injection code execution 1 to LiteLLM's 2026 MCP command injection 2, and the same shape appears in the archetype entries below, from LangChain through Langflow to LiteLLM.

That is the first finding of our dated ledger of AI-stack CVEs, built from NIST's National Vulnerability Database, GitHub Security Advisories, and CISA's Known Exploited Vulnerabilities Catalog 3.

Vendors were same-day. Fleets were not. Across the three ledger entries CISA lists as exploited in the wild, the median gap from CVE publication to confirmed exploitation is 28 days, and every one of the three shipped its fix in the same advisory that disclosed the bug.

Why the same bug keeps coming back #

The archetype entries read as a pattern rather than a coincidence:

  • LangChain, 2023: the LLMMathChain utility let prompt injection reach Python's exec method, arbitrary code execution at CVSS 9.8 1 . Two more 2023 criticals repeated the shape through a Jira wrapper and load_prompt45 .
  • Langflow, 2024 and 2025: a custom_component endpoint executed a Python script supplied in the request, CVSS 9.8 6 . Ten months later the /api/v1/validate/code endpoint did it again without authentication, CVSS 9.8, and put Langflow on CISA's exploited catalog7 .
  • Marimo, 2026: the reactive Python notebook's terminal WebSocket endpoint skipped the authentication check its sibling endpoints ran and handed out a full shell, CVSS 9.8 8 .
  • LiteLLM, 2026: the LLM gateway's MCP test endpoints spawned the command field of a submitted server configuration as a subprocess, so any key holder, including low-privilege internal-user keys, could run arbitrary commands on the host 2 .
  • Anyscale Ray, 2023: the job submission API executes submitted code, which is its purpose. CVSS 9.8, tagged disputed after Anyscale maintained Ray is "not intended for use outside of a strictly controlled network environment" 9 .

An agent framework is built to follow instructions embedded in text. A gateway is built to act on request bodies. A test endpoint is built to run a configuration to see whether it works. In this layer execution is the product, so an input-handling slip becomes code execution instead of a garbled output. The amplifier is shipping posture: Wiz Threat Research, which runs honeypots across LiteLLM, Flowise, LangChain, Langflow, ChromaDB, and Ollama, reports that most AI tools ship unauthenticated and tells defenders to treat an unauthenticated internet exposure as compromised 10. Wiz's State of AI in the Cloud report puts 90 percent of cloud environments on self-hosted AI software 10.

The AI stack ledger's first build: highest-severity CVEs per product — CVSS v3.1 base score recorded in NVD (primary score where assigned, secondary otherwise) (CVSS v3.1 base score)
Product and CVE CVSS 3.1 base score Source
|---|---|---|
| LangChain LLMMathChain (CVE-2023-29374) | 9.8 (prompt injection executes arbitrary code via Python exec; disclosed 2023-04-05) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-29374) | 
| LangChain JiraAPIWrapper (CVE-2023-34540) | 9.8 (remote code execution via crafted input; disclosed 2023-06-14) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-34540) | 
| LangChain load_prompt (CVE-2023-34541) | 9.8 (arbitrary code execution; disclosed 2023-06-20) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-34541) | 

| Langflow custom_component endpoint (CVE-2024-37014) | 9.8 (endpoint executes a Python script supplied in the request; disclosed 2024-06-10) | nvd.nist.gov | | Langflow /api/v1/validate/code (CVE-2025-3248) | 9.8 (unauthenticated code injection; CISA KEV added 2025-05-05) | nvd.nist.gov | | Anyscale Ray job submission API (CVE-2023-48022) | 9.8 (arbitrary code execution; record carries the disputed tag after vendor objection) | nvd.nist.gov | | MLflow file overwrite (CVE-2023-6018) | 9.8 (unauthenticated overwrite of any file on the host; finder rates OS command injection (CWE-78)) | nvd.nist.gov | | Marimo /terminal/ws (CVE-2026-39987) | 9.8 (pre-auth remote code execution via terminal WebSocket; CVSS 4.0 score 9.3 Critical; CISA KEV added 2026-04-23) | nvd.nist.gov | | LiteLLM MCP test endpoints (CVE-2026-42271) | 8.8 (command injection; CVSS 4.0 secondary score 8.7 High; CISA KEV added 2026-06-08; patched in 1.83.7) | nvd.nist.gov | | Ollama model path digest (CVE-2024-37032) | 8.8 (path traversal via unvalidated digest; score from NVD's secondary scoring source) | nvd.nist.gov |

Download this table: CSV · JSON (all tables) · free to cite with attribution.

The clock that matters runs from disclosure to your patch #

Days from CVE publication to CISA listing the bug as exploited in the wild — calendar days between the CVE's NVD publication date and CISA's KEV addition date, as computed for this ledger from the two dates each NVD record carries (days)
AI-stack CVE days from publication to KEV Source
--- --- ---
Marimo CVE-2026-39987 14 (published 2026-04-09, KEV added 2026-04-23; Sysdig observed first exploitation attempt 9h41m after the GitHub advisory) nvd.nist.gov
| Langflow CVE-2025-3248 | 28 (published 2025-04-07, KEV added 2025-05-05) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2025-3248) | 
| LiteLLM CVE-2026-42271 | 31 (published 2026-05-08, KEV added 2026-06-08) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-42271) | 

Download this table: CSV · JSON (all tables) · free to cite with attribution.

CISA's catalog dates put numbers on the deployment lag. Marimo's CVE published April 9, 2026 and reached the catalog April 23: 14 days 8. Langflow's code-injection CVE published April 7, 2025 and was catalogued May 5: 28 days 7. LiteLLM's command injection published May 8, 2026 and made the catalog June 8: 31 days 2. All three advisories carried their own fix, Langflow 1.3.0, Marimo 0.23.0, LiteLLM 1.83.7 11. The window those clocks measured was not vendor response time. It was the time operators took to apply a patch that already existed.

Exploitation starts earlier than any official clock. Sysdig's Threat Research Team observed the first attempt against Marimo's terminal bug 9 hours and 41 minutes after the GitHub advisory published, with no public proof-of-concept in existence: the attacker built a working exploit from the advisory's description and stole credentials in under three minutes 12. Wiz's 90-day honeypot telemetry caught the LiteLLM chain in active use, including a fake MCP server whose command field downloaded a Monero miner and returned a valid handshake so the connection test looked successful; external researchers have linked the Qilin ransomware group to that chain 10. Wiz's warning is blunter than the CVE dates: attackers weaponize new vulnerabilities as soon as fixes appear in code, often ahead of CVE assignment 10.

The exploited set is also a map. CISA's catalog filter lists Langflow, BerriAI, Marimo, MLflow, Ray-Project, and n8n among vendors with confirmed exploited CVEs, out of 1,734 catalog entries at the October 2026 build 3. Ollama, vLLM, and LangChain are absent from that vendor list. Absence is not safety: Ollama's record includes a digest-validation path traversal scored 8.8 13, and vLLM's published entries so far are availability bugs, such as a completions request with an empty prompt crashing the server 14. Those three names simply have no entry yet caught being exploited.

How the ledger decides what counts as the AI stack #

Scope is the software layer that runs AI: model servers and inference engines, agent and workflow frameworks, LLM gateways and proxies, MCP servers, vector stores, and the ML tooling that ships beside them. Rows come from NVD and GitHub Security Advisories, cross-checked against CISA's catalog. Severity is the CVSS 3.1 base score NVD records, and each product keeps its representative worst entries rather than every near-duplicate. The two non-execution exceptions are MLflow's unauthenticated file overwrite (9.8) 15 and Ollama's digest path traversal (8.8) 13. NVD indexes by product and CVE number and carries no AI-infrastructure grouping, which is the gap this ledger maintains; the standing count updates as the stack grows.

The row to watch for in every new AI tool is any feature that tests, runs, or executes a user-supplied string. That is where the stack keeps bleeding.

References

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-29374)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2026-42271)nvd.nist.gov ↗

[CISA](https://www.cisa.gov/known-exploited-vulnerabilities-catalog)cisa.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-34540)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-34541)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-37014)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2025-3248)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2026-39987)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-48022)nvd.nist.gov ↗

[Wiz](https://www.wiz.io/blog/ai-infrastructure-honeypot)wiz.io ↗

[GitHub Advisory GHSA-v4p8-mg3p-g94g](https://github.com/BerriAI/litellm/security/advisories/GHSA-v4p8-mg3p-g94g)github.com ↗

[Sysdig](https://www.sysdig.com/blog/marimo-oss-python-notebook-rce-from-disclosure-to-exploitation-in-under-10-hours)sysdig.com ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-37032)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-8768)nvd.nist.gov ↗

[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-6018)nvd.nist.gov ↗

Cite this story

ProvenBrief (2026). "The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug." ProvenBrief. https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Compare the other ProvenBrief trackers

Browse the maintained datasets, changelogs and downloadable source material behind the stories.

Explore all datasets → 39 factual claims in this story were independently checked against primary sources before publication;

3 unverifiable claims were removed during fact-checking. Read our

editorial standards.

This tracker changes

The figures above move. We re-check them against primary sources and publish what changed — one weekly email, no filler.

This story

[WordsSam Rivera· Staff Writer](https://provenbrief.com/team/sam)

[Fact-checkElena Volkov· Standards & Verification Editor](https://provenbrief.com/team/elena)

[EditingDiana Okafor· Editor-in-Chief](https://provenbrief.com/team/diana)

Standards reviewJames Whitfield· Standards & Compliance Officer

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.

── more in #ai-safety 4 stories · sorted by recency
── more on @langchain 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-ai-stack-securit…] indexed:0 read:8min 2026-10-08 · —