# Red Hat Brings Enterprise AI at Scale Into Focus

> Source: <https://techstrong.ai/videos/red-hat-brings-enterprise-ai-at-scale-into-focus/>
> Published: 2026-08-17 15:55:14+00:00

Synopsis: AI pilots have proven the technology works. What comes next is harder — running AI securely, efficiently and alongside the rest of the enterprise IT stack, without treating the model as if it were the whole problem.

AI pilots have proven the technology works. What comes next is harder — running AI securely, efficiently and alongside the rest of the enterprise IT stack, without treating the model as if it were the whole problem. Enterprises are quickly finding that the model is only one piece of the picture. Governance, monitoring, policy controls and the platform underneath everything are what separate a working proof of concept from something a business can actually depend on.

Mike Vizard sits down with Tushar Katarki, head of product for Red Hat AI Platforms, to work through what enterprise AI at scale really looks like once the pilot stage is over. Katarki’s central point is that open source models and open source platforms are improving faster than most teams planned for, which gives organizations more control over cost, data and infrastructure than they had a year ago — and reshapes how IT leaders should be thinking about the whole stack, not just the model on top of it.

They get into the mechanics of why inference has become the center of gravity. Katarki compares vLLM to the Linux kernel, arguing that just as Linux abstracted applications from hardware, an open source inference engine abstracts models from the growing variety of AI accelerators underneath them. Distributed inference through projects like llm-d matters because enterprises are not running one model on one kind of hardware — they are routing many models across many accelerators, balancing low-latency workloads against high-throughput ones, and stitching small specialized models together with larger distributed ones.

The last stretch turns to AI gateways and the shape of the control plane that agentic workloads actually need. Token tracking, rate limits, quota management, chargeback, guardrails, tool-calling controls — none of that is optional once agents move past chatbots into longer-running workflows. Katarki’s framing is that IT teams need to start thinking like model service providers and agent service providers, delivering AI as a new class of workload with the same cost, security and governance discipline that any other production service earns.
