# Kong AI Gateway 2.2: Built for what comes next in AI

> Source: <https://konghq.com/blog/product-releases/kong-ai-gateway-2-2>
> Published: 2026-09-30 16:00:00+00:00

# Kong AI Gateway 2.2: Built for what comes next in AI

Alex Drag

Head of Product Marketing

AI infrastructure is changing fast.

Enterprises are no longer building around a small set of models exposed through relatively consistent interfaces. New model architectures are emerging (Here’s looking at you Jev!). Providers are introducing new APIs and capabilities. Enterprises are running self-hosted models and specialized AI services. And the interfaces connecting all of this are becoming more diverse, not less.

For platform teams, that creates a challenge: *how do you maintain a consistent layer of control when the AI underneath it keeps changing?*

Kong AI Gateway 2.2 is built for this increasingly diverse AI ecosystem. With native support for emerging models and provider capabilities, passthrough support for custom and evolving AI interfaces, and greater extensibility for enterprise-specific requirements, this release expands the range of AI traffic organizations can bring under a common control point.

It is also another step toward **strategic portability**. Developers should be able to adopt the right AI technology for the job without creating a new infrastructure and governance stack every time the underlying model, provider, or interface changes.

## Support what’s new: TypeSafe Jev and Skills

AI is not just producing different answers. It is starting to do different kinds of work.

### Native support for TypeSafe Jev

JEV introduces a different model for using AI: instead of generating text, it returns typed, probabilistic decisions that software can act on directly. By returning only the structured decision an application needs, JEV reduces ambiguity and token overhead while making AI decisions easier to use directly in application logic.

That makes it useful anywhere an application needs AI to make a fast, structured choice. For example, JEV can work with Kong AI Gateway to inform intelligent LLM routing — deciding which model should handle a request, while Kong AI Gateway executes the routing to the selected model.

AI Gateway 2.2 adds native support for TypeSafe JEV as a provider and format, including its decisions capability, allowing teams to adopt this model while keeping the same gateway infrastructure, governance, and operational controls they use for the rest of their AI traffic.

### Support for Skills APIs

Models themselves are not the only thing evolving. Providers are adding new APIs that expand what models can do.

AI Gateway 2.2 adds support for Skills APIs, giving developers a consistent way to work with reusable, versioned Skills across supported AI providers. Kong handles differences between provider APIs, including translating between Kong’s common representation and provider-specific formats.

That means teams can adopt emerging capabilities like Skills without hardwiring applications to a single provider’s implementation, helping preserve strategic portability as provider capabilities evolve.

Together, native support for emerging models and capabilities helps organizations preserve choice as the AI ecosystem evolves. That is strategic portability in practice: adopt what’s new without rebuilding the control layer around it.

## Support what’s next with passthrough mode

Native integrations matter. But the AI ecosystem is evolving too quickly for every new model, API, and endpoint to fit an established provider format on day one.

Developers should not have to wait for their AI gateway to catch up before they can use them.

AI Gateway 2.2 introduces passthrough mode for AI workloads that use custom or non-standard interfaces. That includes self-hosted model servers such as vLLM, Ollama, and NVIDIA NIM, provider preview APIs with evolving schemas, and specialized non-LLM AI endpoints.

In passthrough mode, Kong forwards requests and responses without transforming the body, while preserving capabilities such as upstream provider authentication, Kong authentication, rate limiting, and logging.

The result is a more open model for AI governance: your AI workload does not have to conform to a predefined provider format to sit behind Kong.

This gives organizations the freedom to introduce new models and AI services while maintaining a consistent control layer, reinforcing strategic portability as their AI architecture evolves.

## Make AI Gateway yours with custom plugins

Supporting a broader AI ecosystem is only half the challenge. Enterprises also have their own requirements for how that traffic should be handled.

AI Gateway 2.2 adds support for custom plugins in the AI Gateway control plane, including support for streaming custom plugins from the control plane to data planes as well as plugins installed directly on the data plane.

This brings the extensibility enterprises expect from Kong’s API infrastructure into AI Gateway. Platform teams can add organization-specific logic and policies, giving them the flexibility to adapt AI traffic management to their own architecture, security, and governance requirements.

## More control over AI consumption

As AI adoption expands, platform teams are not just being asked to connect more AI. They are being asked to control who can consume it and what that consumption costs.

AI Gateway 2.2 extends **AI Rate Limiting Advanced** with credential-based matching. Policies can now use credential as a matching dimension, allowing organizations to apply AI consumption controls at a more granular credential level.

This gives teams another way to translate shared AI infrastructure into meaningful consumption boundaries, particularly where multiple credentials are accessing the same models or AI services.

### Tech Preview: Reduce AI token consumption with Headroom

AI Gateway 2.2 also introduces a Tech Preview integration with **Headroom** for AI compression.

Prompt compression reduces the number of tokens sent to or returned from an LLM while preserving the information needed to complete the task. Because token consumption directly affects inference cost, and can also affect latency, compression can make AI workloads more efficient without requiring teams to change models or applications.

Headroom uses output shaping and adaptive verbosity to reduce output token consumption. Kong integrates Headroom with its AI prompt compression capabilities while remaining the primary LLM egress point.

The integration captures metrics including tokens saved, compression ratio, and transformations applied, making the impact of compression measurable. If Headroom is unavailable or times out, Kong can forward the original content rather than interrupting the AI request.

For teams facing growing inference bills, this provides another lever for improving AI economics: reduce the tokens you are paying for while preserving the information the application needs.

## Built for an AI ecosystem that won’t stop changing

No one knows exactly what the dominant AI architecture will look like two years from now.

What we do know is that new models, provider APIs, protocols, interaction patterns, and specialized AI services will continue to emerge. Organizations need the freedom to take advantage of those innovations without tying their infrastructure to any single model, provider, or interface.

That is the core of strategic portability.

AI Gateway 2.2 expands both the AI workloads Kong can support natively and the workloads organizations can bring under governance without waiting for native support. Developers get greater freedom to choose the AI technologies that fit the job. Platform teams retain a consistent place to apply control as those choices change.

And once that traffic is behind the gateway, teams have more ways to extend it, control consumption, secure it, and improve its economics.

Alongside the 2.2 release, we’ve also updated our support policy with the latest guidance on supported versions, maintenance, and upgrade expectations.

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

# Event-Driven AI: Agents Shouldn’t Have to Keep Asking if Something Happened

Consider a bank using an AI agent to investigate potentially fraudulent transactions. A transaction is processed and the bank’s fraud detection system identifies suspicious activity. That event is published to Kafka: fraud.suspected The bank’s Fraud

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access

Greg Peranich

# Announcing Kong AI Gateway 2.0: Built for the Pace of Agentic AI

Since its launch, Kong AI Gateway has shipped as part of Kong API Gateway. Every AI capability we built (LLM routing, prompt guarding, semantic caching, token-based rate limiting) rode the same release train as the API gateway that powers some of th

Alex Drag

# Enterprise-Grade MCP Access Control Is Here. Your Gateway Makes It Real.

Kong makes every MCP client and server work with Enterprise-Managed Authorization, whether they speak the protocol or not.
Standard MCP authorization is user-scoped and interactive: every user authorizes every server, one consent flow at a time. T

Michael Field

# AI Token Cost Management: Why AI Spend Gets Out of Control (and How To Fix It)

This post is based on Kong's webinar on token cost management. Watch it below, or keep reading for the breakdown. Token cost management is how a business tracks, controls, and governs what it spends on AI model usage, the same way it already track

Kong

# AI Governance Tools and Platforms: An Enterprise Guide

AI adoption is outpacing the speed at which many organizations can establish standardized oversight. In fact, Stanford HAI reported that 78% of surveyed organizations were using AI in 2024, a significant jump from 55% the previous year ( AI Index Re

Kong

# Take Control of the Economics of AI with Kong AI Gateway

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

# Event-Driven AI: Agents Shouldn’t Have to Keep Asking if Something Happened

Consider a bank using an AI agent to investigate potentially fraudulent transactions. A transaction is processed and the bank’s fraud detection system identifies suspicious activity. That event is published to Kafka: fraud.suspected The bank’s Fraud

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access

Greg Peranich

# Announcing Kong AI Gateway 2.0: Built for the Pace of Agentic AI

Since its launch, Kong AI Gateway has shipped as part of Kong API Gateway. Every AI capability we built (LLM routing, prompt guarding, semantic caching, token-based rate limiting) rode the same release train as the API gateway that powers some of th

Alex Drag

# Enterprise-Grade MCP Access Control Is Here. Your Gateway Makes It Real.

Kong makes every MCP client and server work with Enterprise-Managed Authorization, whether they speak the protocol or not.
Standard MCP authorization is user-scoped and interactive: every user authorizes every server, one consent flow at a time. T

Michael Field

# AI Token Cost Management: Why AI Spend Gets Out of Control (and How To Fix It)

This post is based on Kong's webinar on token cost management. Watch it below, or keep reading for the breakdown. Token cost management is how a business tracks, controls, and governs what it spends on AI model usage, the same way it already track

Kong

# AI Governance Tools and Platforms: An Enterprise Guide

AI adoption is outpacing the speed at which many organizations can establish standardized oversight. In fact, Stanford HAI reported that 78% of surveyed organizations were using AI in 2024, a significant jump from 55% the previous year ( AI Index Re

Kong

# Take Control of the Economics of AI with Kong AI Gateway

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

# Event-Driven AI: Agents Shouldn’t Have to Keep Asking if Something Happened

Consider a bank using an AI agent to investigate potentially fraudulent transactions. A transaction is processed and the bank’s fraud detection system identifies suspicious activity. That event is published to Kafka: fraud.suspected The bank’s Fraud

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access

Greg Peranich

# Announcing Kong AI Gateway 2.0: Built for the Pace of Agentic AI

Since its launch, Kong AI Gateway has shipped as part of Kong API Gateway. Every AI capability we built (LLM routing, prompt guarding, semantic caching, token-based rate limiting) rode the same release train as the API gateway that powers some of th

Alex Drag

# Enterprise-Grade MCP Access Control Is Here. Your Gateway Makes It Real.

Kong makes every MCP client and server work with Enterprise-Managed Authorization, whether they speak the protocol or not.
Standard MCP authorization is user-scoped and interactive: every user authorizes every server, one consent flow at a time. T

Michael Field

# AI Token Cost Management: Why AI Spend Gets Out of Control (and How To Fix It)

This post is based on Kong's webinar on token cost management. Watch it below, or keep reading for the breakdown. Token cost management is how a business tracks, controls, and governs what it spends on AI model usage, the same way it already track

Kong

# AI Governance Tools and Platforms: An Enterprise Guide

AI adoption is outpacing the speed at which many organizations can establish standardized oversight. In fact, Stanford HAI reported that 78% of surveyed organizations were using AI in 2024, a significant jump from 55% the previous year ( AI Index Re

Kong

## Ready to see Kong in action?

Get a personalized walkthrough of Kong's platform tailored to your architecture, use cases, and scale requirements.
