{"slug": "every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater", "title": "Every CISO Needs an AIBOM in 2026. Most Vendors Ship Theater.", "summary": "A security practitioner argues that an AI Bill of Materials (AIBOM) must be a continuously updated graph of runtime model deployments — covering model provenance, reachable systems, and change authority — rather than the procurement-based inventories most vendors sell. Drawing on a mid-market fintech CISO's audit that surfaced 73 AI systems including unreviewed Hugging Face downloads and an unmonitored vLLM instance, the account says most \"AI governance\" products track SSO, CASB, and contract data and never inspect the model artifact itself.", "body_md": "A friend of mine runs security at a mid-market fintech. About 900 engineers, Series D, the kind of place where the AI budget tripled in eighteen months and nobody can quite explain where it all went. She called me in August because her board had asked a question she couldn't answer: *What AI are we actually running?*\n\nNot \"what did we buy.\" What are we *running*. In production. Right now.\n\nShe pulled her procurement list. Fourteen vendors. Then she pulled her cloud bills and found another six AI services nobody had logged. Then her platform team mentioned the three internal Llama fine-tunes they'd deployed on an EKS cluster for the fraud team. Then someone remembered the LangChain app the growth team shipped, which calls OpenAI, Anthropic, and a Cohere reranker depending on the query. Then the data science org admitted to a vLLM instance running a quantized Mistral variant that had been up for seven months, serving an internal tool nobody had security-reviewed.\n\nBy the end of the week her \"AI inventory\" was a 73-row spreadsheet, and she still didn't trust it. The part that bothered her most wasn't the sprawl. It was that two of those models had been downloaded directly from Hugging Face by a staff engineer who had since left the company. Nobody knew what training data they'd been exposed to. Nobody knew if they'd been tampered with. They were just... running. Taking customer input. Returning answers.\n\nShe asked me what I'd do. I told her: you need an AIBOM. A real one. And no, your GRC vendor's new \"AI Governance Module\" isn't it.\n\nAn AI Bill of Materials is the single most undervalued security artifact of 2026, and ninety percent of what gets sold under that name is theater — a spreadsheet of vendor names dressed up as governance. A functional AIBOM has to answer four questions continuously, not quarterly: what models are running, where did they come from, what can they reach, and who is allowed to change them. If your AIBOM can't answer those four at any moment, you don't have an AIBOM. You have a compliance screenshot.\n\nThe SBOM analogy is useful but it undersells the problem. An SBOM tracks dependencies that are deterministic — version 2.14.3 of a library behaves the same way across your fleet. A model doesn't. A model has weights, a tokenizer, a config, a prompt scaffold, a serving runtime, system prompts, retrieval sources, tool definitions, and in most production stacks, a half-dozen downstream integrations that materially change its behavior. Two deployments of the \"same\" Llama-3-70B can behave like different systems because one has a RAG pipeline hitting your CRM and the other doesn't.\n\nSo an AIBOM isn't a list. It's a graph. Here's what has to be in it, at minimum:\n\nIf your AIBOM stops at bullet one and two, which is where most vendor products stop, you've described the engine. You haven't described the car, the driver, or the road.\n\nI've evaluated eleven products that pitch themselves as \"AI inventory\" or \"AI governance\" platforms this year. I'll spare the names. Here's what they consistently get wrong.\n\n**They inventory procurement, not runtime.** The majority of these tools integrate with your SSO logs, your CASB, and your contract management system. They produce a list of AI vendors you've paid. That is useful for finance. It is nearly useless for security. The Llama fine-tune running on an unmonitored EKS node doesn't show up in your SSO logs. The quantized GGUF model someone loaded into LM Studio on a developer laptop doesn't show up in your CASB. The models that scare me most are precisely the ones that generate no procurement paper trail.\n\n**They treat the model as a black box.** A real AIBOM has to inspect the model artifact. What's the hash? Does it match a known provenance? Has anyone fine-tuned it, and on what? Is the tokenizer the one that shipped with the base model or has it been swapped? Most governance tools can't answer any of this because they never touch the artifact. They just ingest the name.\n\n**They ignore the application layer.** The riskiest thing in your stack is not the model. It's the code around the model — the LangChain glue, the custom retrievers, the tool definitions, the prompt templates with injected user data, the Flask endpoint that passes a raw user string to a function-calling API with access to your production database. An AIBOM that lists \"GPT-4\" but doesn't catalog the fifty lines of Python that wrap it is lying to you about your attack surface.\n\n**They refresh on a cadence.** Weekly. Daily if you're lucky. In an environment where a developer can deploy a new model in forty minutes, cadence is the wrong primitive. Your AIBOM has to be event-driven. New container with a model weight in it? AIBOM entry. New Ollama endpoint reachable on an internal subnet? AIBOM entry. New Hugging Face download in a build pipeline? AIBOM entry, with the hash, right now, not at the next sync.\n\n**They don't correlate findings.** This one bothers me the most. An AIBOM that's just an inventory is table stakes. An AIBOM that cross-references each entry against known model vulnerabilities, insecure runtime configurations, exposed inference endpoints, and application-layer flaws in the surrounding code — that's security. The inventory is the beginning of the conversation, not the end.\n\nLet me describe how I'd build this, because I think concretely about it and I've watched too many teams get stuck in slide-deck purgatory.\n\nStart at the artifact layer. Every build pipeline that produces a container, lambda, or image gets scanned for model weights, tokenizer files, and the obvious AI framework imports. You're looking for `.gguf`, `.safetensors`, `.bin`, `.onnx`, `pytorch_model.*`, and the Python signatures of transformers, llama-cpp, vllm, langchain, llamaindex, haystack, and the long tail. This is a code-scanning problem. On Cybrium we handle it with cyscan, which runs 1,815 rules across 75+ languages and specifically covers the AI-adjacent code patterns most SAST tools miss. The output is: for every repo and every build, what AI code and what AI artifacts are present.\n\nMove to the runtime layer. You need an active scanner that walks your internal network — your VPCs, your Kubernetes clusters, your developer VLANs — and identifies inference endpoints. Ollama has a signature. vLLM has one. TGI, LocalAI, Triton, LM Studio, llama.cpp — all identifiable. Cybrium's cyradar does this specifically for AI inference surfaces, and the finding we see most often is an Ollama instance bound to 0.0.0.0 on a developer's EC2 that's reachable from a peered VPC. Nobody put that in the procurement spreadsheet.\n\nMove to the application layer. Every web endpoint, every API, every Lambda function that mediates between a user and a model gets fuzzed. Prompt injection, system prompt leakage, tool abuse, output handling flaws, SSRF through retrieval, the full taxonomy. cyweb runs 22 fuzz categories across these, and the finding rate on production LLM apps is substantially higher than most CISOs expect. The AIBOM has to carry those findings as attributes of the component, not as separate tickets.\n\nCorrelate. Every artifact gets a stable ID. Every runtime endpoint gets linked to the artifact it serves. Every application gets linked to the runtime it calls. Every finding gets attached to the right node in the graph. Now when the board asks \"what AI are we running,\" you have an answer. When your IR team asks \"if this CVE drops tomorrow for vLLM, what's exposed,\" you have an answer. When legal asks \"which of our customer data flows touch a model trained on scraped web data,\" you have an answer.\n\nThat graph, continuously updated, event-driven, correlated with findings — that's an AIBOM.\n\nI want to spend a minute on provenance because it's the hardest part and the part vendors handwave hardest.\n\nIf an engineer runs:\n\n```\nollama pull llama3:70b-instruct-q4_K_M\n```\n\nWhat did they just download? They downloaded a quantized GGUF file from Ollama's model library, which was packaged from Meta's release, which was trained on a dataset Meta hasn't fully disclosed. The chain of custody has at least three hops, and at every hop the artifact could have been swapped. Hugging Face has had typosquatted model names. Compromised accounts have pushed malicious weights. Pickle deserialization in older model formats has been a reliable RCE primitive for years.\n\nA real AIBOM pins the hash. Not the name, not the version tag — the hash of the actual bytes on disk. And it tracks that hash against a known-good reference. If the hash doesn't match what Meta published, that's a finding. If the hash matches but the tokenizer has been swapped, that's a finding. If the model was downloaded from a mirror you don't trust, that's a finding.\n\nMost governance tools do none of this. They record \"llama3-70b\" as a string and move on. That string tells you nothing about what's actually running in your cluster.\n\nThe reflex at a mature security org is to buy best-of-breed and glue it together. SBOM tool here, model scanner there, DAST for the LLM apps, posture management for the inference endpoints, a GRC overlay to roll it up. I've watched a lot of teams try this and it collapses for one reason: the correlation layer is the hardest part, and nobody owns it when it lives between five vendors.\n\nThe AIBOM only works as an AIBOM when the code findings, the artifact provenance, the runtime discovery, and the application-layer vulnerabilities land in the same graph. If they land in five different consoles with five different IDs for the same model, you don't have an AIBOM. You have five dashboards and a Jira project to reconcile them.\n\nThis is why Cybrium builds all four layers under one roof and exposes them through a single MCP server with 10 tools, so your AI agents and your SOC can query the same graph. I'm not neutral on this. I built it this way because I spent years watching customers try the five-vendor approach and I never once saw it produce a trustworthy inventory.\n\nHere's what I think is actually going on in the market, and why I'm writing this now.\n\nFor the last two years, \"AI security\" has been a category defined by whoever showed up first with a slide. Red-teaming startups, prompt-injection scanners, data-leakage classifiers, model firewalls, governance overlays. Each one is a feature. None of them is a program.\n\nIn 2026 the center of gravity is shifting toward the AIBOM because every other control depends on it. You can't red-team a model you don't know exists. You can't apply a policy to an endpoint you haven't discovered. You can't do incident response on a system whose provenance you can't reconstruct. The inventory is the substrate. Everything else is a function that takes the inventory as input.\n\nThe CISOs I talk to who are ahead of this are already treating their AIBOM as the primary artifact and everything else — the red-teaming, the policies, the runtime guards — as consumers of that artifact. The CISOs who are behind are still buying point tools and trying to produce a spreadsheet for the board every quarter. The gap between those two postures is going to be the thing regulators, insurers, and acquirers notice next year. I'm already seeing it in diligence questionnaires.\n\nIf your board hasn't asked yet, they will. And \"we have a spreadsheet\" is not going to survive the follow-up question.\n\nIf I were sitting in my friend's chair at that fintech, this is the sequence I'd run.\n\nWeek one, inventory the artifact layer. Scan every repo and build pipeline for model files and AI framework imports. You'll find more than you expect. On Cybrium that's a cyscan deployment and about two days of triage.\n\nWeek two, discover the runtime. Point a scanner at every internal subnet and surface every inference endpoint, hosted or self-hosted. Ollama, vLLM, TGI, LocalAI, Triton, LM Studio, llama.cpp — all of them. That's cyradar, and the finding rate is uncomfortable. Good. That's the point.\n\nWeek three, test the application layer. Every LLM-fronted endpoint gets fuzzed across the 22 categories cyweb covers — prompt injection, tool abuse, output handling, retrieval SSRF, the rest. The findings here are the ones most likely to produce an actual incident.\n\nWeek four, correlate. Build the graph. Link artifacts to runtimes to applications to findings. Give every node a stable ID. Make it queryable by your SOC and by your AI agents through the MCP interface.\n\nThen operate it continuously. Not quarterly. Not weekly. Event-driven, every build, every deploy, every new endpoint. That's the bar for 2026.\n\nThe AIBOM is not a document. It's a system. If you treat it like a document you'll keep producing screenshots for the board and discovering surprises at 2 a.m. If you treat it like a system you'll know what you're running, where it came from, what it can reach, and who can change it — at any moment, for any model, in any environment.\n\nThat's the baseline now. Not aspiration. Baseline.\n\nIf you're building this out and want to compare notes on what's working and what isn't, or if you want to see how we assemble the four layers into a single graph on Cybrium — cyscan for the code, cyradar for the runtime, cyweb for the apps, and the MCP server for querying it all — find me at `anand@cybrium.ai`.", "url": "https://wpnews.pro/news/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater", "canonical_source": "https://dev.to/grumpysage/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater-4mpi", "published_at": "2026-10-08 19:11:11+00:00", "updated_at": "2026-10-08 19:20:07.027391+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "mlops", "ai-infrastructure", "ai-agents"], "entities": ["Hugging Face", "OpenAI", "Anthropic", "Cohere", "LangChain", "vLLM", "Mistral", "Llama"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater", "markdown": "https://wpnews.pro/news/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater.md", "text": "https://wpnews.pro/news/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater.txt", "jsonld": "https://wpnews.pro/news/every-ciso-needs-an-aibom-in-2026-most-vendors-ship-theater.jsonld"}}