{"slug": "the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug", "title": "The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug", "summary": "Eight of the ten worst CVEs in the first build of a new AI-stack security ledger are the same class of bug: remote code execution where software executes or obeys model output, according to the ledger compiled from NIST's National Vulnerability Database, GitHub Security Advisories, and CISA's Known Exploited Vulnerabilities Catalog. The ledger's archetype entries run from LangChain's 2023 LLMMathChain prompt-injection code execution at CVSS 9.8 through Langflow's 2024 and 2025 custom_component and /api/v1/validate/code flaws at CVSS 9.8 to LiteLLM's 2026 MCP command injection, and across the three entries CISA lists as exploited in the wild the median gap from CVE publication to confirmed exploitation is 28 days. Wiz Threat Research, which runs honeypots across LiteLLM, Flowise, LangChain, Langflow, ChromaDB, and Ollama, reports most AI tools ship unauthenticated and advises defenders to treat an unauthenticated internet exposure as compromised, while Wiz's State of AI in the Cloud report puts 90 percent of cloud environments on self-hosted AI software.", "body_md": "# The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug\n\nThe software layer that runs AI - model servers like Ollama and vLLM, agent frameworks like LangChain, vector databases, and MCP servers - has accumulated its own CVE record, and it keeps producing the same dangerous class of bug: remote code execution where the stack executes or obeys model output. We compile every published CVE across the AI/ML software stack into one dated ledger, each row tagged by vulnerability class, CVSS severity, affected product, time-to-patch, and whether it is known to be exploited in the wild.\n\n## The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug\n\nEight of the ten worst CVEs in the first build of our AI-stack ledger are the same bug: software that executes whatever it is handed. The record runs from LangChain's 2023 prompt-injection code execution [1](#ref-1) to LiteLLM's 2026 MCP command injection [2](#ref-2), and the same shape appears in the archetype entries below, from LangChain through Langflow to LiteLLM.\n\nThat is the first finding of our dated ledger of AI-stack CVEs, built from NIST's National Vulnerability Database, GitHub Security Advisories, and CISA's Known Exploited Vulnerabilities Catalog [3](#ref-3).\n\nVendors were same-day. Fleets were not. Across the three ledger entries CISA lists as exploited in the wild, the median gap from CVE publication to confirmed exploitation is 28 days, and every one of the three shipped its fix in the same advisory that disclosed the bug.\n\n## Why the same bug keeps coming back\n\nThe archetype entries read as a pattern rather than a coincidence:\n\n- LangChain, 2023: the LLMMathChain utility let prompt injection reach Python's exec method, arbitrary code execution at CVSS 9.8 [1](#ref-1) . Two more 2023 criticals repeated the shape through a Jira wrapper and load_prompt[4](#ref-4)[5](#ref-5) .\n- Langflow, 2024 and 2025: a custom_component endpoint executed a Python script supplied in the request, CVSS 9.8 [6](#ref-6) . Ten months later the /api/v1/validate/code endpoint did it again without authentication, CVSS 9.8, and put Langflow on CISA's exploited catalog[7](#ref-7) .\n- Marimo, 2026: the reactive Python notebook's terminal WebSocket endpoint skipped the authentication check its sibling endpoints ran and handed out a full shell, CVSS 9.8 [8](#ref-8) .\n- LiteLLM, 2026: the LLM gateway's MCP test endpoints spawned the command field of a submitted server configuration as a subprocess, so any key holder, including low-privilege internal-user keys, could run arbitrary commands on the host [2](#ref-2) .\n- Anyscale Ray, 2023: the job submission API executes submitted code, which is its purpose. CVSS 9.8, tagged disputed after Anyscale maintained Ray is \"not intended for use outside of a strictly controlled network environment\" [9](#ref-9) .\n\nAn agent framework is built to follow instructions embedded in text. A gateway is built to act on request bodies. A test endpoint is built to run a configuration to see whether it works. In this layer execution is the product, so an input-handling slip becomes code execution instead of a garbled output. The amplifier is shipping posture: Wiz Threat Research, which runs honeypots across LiteLLM, Flowise, LangChain, Langflow, ChromaDB, and Ollama, reports that most AI tools ship unauthenticated and tells defenders to treat an unauthenticated internet exposure as compromised [10](#ref-10). Wiz's State of AI in the Cloud report puts 90 percent of cloud environments on self-hosted AI software [10](#ref-10).\n\n| The AI stack ledger's first build: highest-severity CVEs per product — CVSS v3.1 base score recorded in NVD (primary score where assigned, secondary otherwise) (CVSS v3.1 base score) |  |  | \n|---|---|---|\n| Product and CVE | CVSS 3.1 base score | Source | \n|---|---|---|\n| LangChain LLMMathChain (CVE-2023-29374) | 9.8 (prompt injection executes arbitrary code via Python exec; disclosed 2023-04-05) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-29374) | \n| LangChain JiraAPIWrapper (CVE-2023-34540) | 9.8 (remote code execution via crafted input; disclosed 2023-06-14) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-34540) | \n| LangChain load_prompt (CVE-2023-34541) | 9.8 (arbitrary code execution; disclosed 2023-06-20) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-34541) | \n| Langflow custom_component endpoint (CVE-2024-37014) | 9.8 (endpoint executes a Python script supplied in the request; disclosed 2024-06-10) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2024-37014) | \n| Langflow /api/v1/validate/code (CVE-2025-3248) | 9.8 (unauthenticated code injection; CISA KEV added 2025-05-05) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2025-3248) | \n| Anyscale Ray job submission API (CVE-2023-48022) | 9.8 (arbitrary code execution; record carries the disputed tag after vendor objection) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-48022) | \n| MLflow file overwrite (CVE-2023-6018) | 9.8 (unauthenticated overwrite of any file on the host; finder rates OS command injection (CWE-78)) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2023-6018) | \n| Marimo /terminal/ws (CVE-2026-39987) | 9.8 (pre-auth remote code execution via terminal WebSocket; CVSS 4.0 score 9.3 Critical; CISA KEV added 2026-04-23) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-39987) | \n| LiteLLM MCP test endpoints (CVE-2026-42271) | 8.8 (command injection; CVSS 4.0 secondary score 8.7 High; CISA KEV added 2026-06-08; patched in 1.83.7) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-42271) | \n| Ollama model path digest (CVE-2024-37032) | 8.8 (path traversal via unvalidated digest; score from NVD's secondary scoring source) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2024-37032) | \n\nDownload this table: [CSV](https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug/data.csv?table=ai-stack-criticals) · [JSON (all tables)](https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug/data.json) · free to cite with attribution.\n\n## The clock that matters runs from disclosure to your patch\n\n| Days from CVE publication to CISA listing the bug as exploited in the wild — calendar days between the CVE's NVD publication date and CISA's KEV addition date, as computed for this ledger from the two dates each NVD record carries (days) |  |  | \n|---|---|---|\n| AI-stack CVE | days from publication to KEV | Source | \n|---|---|---|\n| Marimo CVE-2026-39987 | 14 (published 2026-04-09, KEV added 2026-04-23; Sysdig observed first exploitation attempt 9h41m after the GitHub advisory) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-39987) | \n| Langflow CVE-2025-3248 | 28 (published 2025-04-07, KEV added 2025-05-05) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2025-3248) | \n| LiteLLM CVE-2026-42271 | 31 (published 2026-05-08, KEV added 2026-06-08) | [nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-42271) | \n\nDownload this table: [CSV](https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug/data.csv?table=kev-lag) · [JSON (all tables)](https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug/data.json) · free to cite with attribution.\n\nCISA's catalog dates put numbers on the deployment lag. Marimo's CVE published April 9, 2026 and reached the catalog April 23: 14 days [8](#ref-8). Langflow's code-injection CVE published April 7, 2025 and was catalogued May 5: 28 days [7](#ref-7). LiteLLM's command injection published May 8, 2026 and made the catalog June 8: 31 days [2](#ref-2). All three advisories carried their own fix, Langflow 1.3.0, Marimo 0.23.0, LiteLLM 1.83.7 [11](#ref-11). The window those clocks measured was not vendor response time. It was the time operators took to apply a patch that already existed.\n\nExploitation starts earlier than any official clock. Sysdig's Threat Research Team observed the first attempt against Marimo's terminal bug 9 hours and 41 minutes after the GitHub advisory published, with no public proof-of-concept in existence: the attacker built a working exploit from the advisory's description and stole credentials in under three minutes [12](#ref-12). Wiz's 90-day honeypot telemetry caught the LiteLLM chain in active use, including a fake MCP server whose command field downloaded a Monero miner and returned a valid handshake so the connection test looked successful; external researchers have linked the Qilin ransomware group to that chain [10](#ref-10). Wiz's warning is blunter than the CVE dates: attackers weaponize new vulnerabilities as soon as fixes appear in code, often ahead of CVE assignment [10](#ref-10).\n\nThe exploited set is also a map. CISA's catalog filter lists Langflow, BerriAI, Marimo, MLflow, Ray-Project, and n8n among vendors with confirmed exploited CVEs, out of 1,734 catalog entries at the October 2026 build [3](#ref-3). Ollama, vLLM, and LangChain are absent from that vendor list. Absence is not safety: Ollama's record includes a digest-validation path traversal scored 8.8 [13](#ref-13), and vLLM's published entries so far are availability bugs, such as a completions request with an empty prompt crashing the server [14](#ref-14). Those three names simply have no entry yet caught being exploited.\n\n## How the ledger decides what counts as the AI stack\n\nScope is the software layer that runs AI: model servers and inference engines, agent and workflow frameworks, LLM gateways and proxies, MCP servers, vector stores, and the ML tooling that ships beside them. Rows come from NVD and GitHub Security Advisories, cross-checked against CISA's catalog. Severity is the CVSS 3.1 base score NVD records, and each product keeps its representative worst entries rather than every near-duplicate. The two non-execution exceptions are MLflow's unauthenticated file overwrite (9.8) [15](#ref-15) and Ollama's digest path traversal (8.8) [13](#ref-13). NVD indexes by product and CVE number and carries no AI-infrastructure grouping, which is the gap this ledger maintains; the standing count updates as the stack grows.\n\nThe row to watch for in every new AI tool is any feature that tests, runs, or executes a user-supplied string. That is where the stack keeps bleeding.\n\n### References\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-29374)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2026-42271)nvd.nist.gov ↗\n\n[CISA](https://www.cisa.gov/known-exploited-vulnerabilities-catalog)cisa.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-34540)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-34541)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-37014)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2025-3248)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2026-39987)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-48022)nvd.nist.gov ↗\n\n[Wiz](https://www.wiz.io/blog/ai-infrastructure-honeypot)wiz.io ↗\n\n[GitHub Advisory GHSA-v4p8-mg3p-g94g](https://github.com/BerriAI/litellm/security/advisories/GHSA-v4p8-mg3p-g94g)github.com ↗\n\n[Sysdig](https://www.sysdig.com/blog/marimo-oss-python-notebook-rce-from-disclosure-to-exploitation-in-under-10-hours)sysdig.com ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-37032)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2024-8768)nvd.nist.gov ↗\n\n[NVD](https://nvd.nist.gov/vuln/detail/CVE-2023-6018)nvd.nist.gov ↗\n\n### Cite this story\n\nProvenBrief (2026). \"The AI Stack Security Ledger: Eight of Its Ten Worst CVEs Are the Same Bug.\" ProvenBrief. https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug\n\nFree to quote and link with attribution. Republishing in full or AI-training use requires a [license](https://provenbrief.com/contact).\n\n### Compare the other ProvenBrief trackers\n\nBrowse the maintained datasets, changelogs and downloadable source material behind the stories.\n\n[Explore all datasets →](https://provenbrief.com/data)\n\n**39 factual claims** in this story were independently checked against primary sources before publication;\n\n**3** unverifiable claims were removed during fact-checking. Read our\n\n[editorial standards](https://provenbrief.com/standards).\n\n### This tracker changes\n\nThe figures above move. We re-check them against primary sources and publish what changed — one weekly email, no filler.\n\n### This story\n\n[WordsSam Rivera· Staff Writer](https://provenbrief.com/team/sam)\n\n[Fact-checkElena Volkov· Standards & Verification Editor](https://provenbrief.com/team/elena)\n\n[EditingDiana Okafor· Editor-in-Chief](https://provenbrief.com/team/diana)\n\n[Standards reviewJames Whitfield· Standards & Compliance Officer](https://provenbrief.com/team/james)\n\nProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our [editorial standards](https://provenbrief.com/standards).", "url": "https://wpnews.pro/news/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug", "canonical_source": "https://provenbrief.com/story/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug", "published_at": "2026-10-08 01:37:01+00:00", "updated_at": "2026-10-08 01:47:03.057324+00:00", "lang": "en", "topics": ["ai-safety", "ai-infrastructure", "ai-agents", "large-language-models", "ai-tools"], "entities": ["LangChain", "Langflow", "LiteLLM", "Marimo", "Anyscale Ray", "Wiz Threat Research", "NIST National Vulnerability Database", "CISA Known Exploited Vulnerabilities Catalog"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug", "markdown": "https://wpnews.pro/news/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug.md", "text": "https://wpnews.pro/news/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug.txt", "jsonld": "https://wpnews.pro/news/the-ai-stack-security-ledger-eight-of-its-ten-worst-cves-are-the-same-bug.jsonld"}}