Over the last several months, Inco AI has been building an inference stack purpose-built for the demands of the agentic era and pushing the efficiency frontier across sustained generation, long-running workloads, and longer sequences—conditions that combine to push inference systems to their limits.
Today's launch of the Inco platform marks an important milestone toward bringing superior agentic inference performance to market and provides a first look at Inco's inference and technology stack. We are releasing high-speed endpoints for Kimi K3, MiniMax M3, GLM 5.3, and GLM 5.3 Flash. Each leads its respective Artificial Analysis provider leaderboard on output speed.
| Model | Output tokens / s | Relative improvement |
|---|---|---|
MiniMax M3GLM 5.3GLM 5.3 FlashInside the Inco Inference Stack
The Artificial Analysis results reflect optimization across the full Inco stack, highlighting our unique approach to bringing peak inference efficiency to the market.
The Inco Platform Is Entering Public Beta We are opening beta access to the Inco platform, starting with Kimi K3, MiniMax M3, GLM 5.3, and GLM 5.3 Flash.
Sign up to try our fastest endpoints on the Inco platform.
Available at launch
- Kimi K3
- MiniMax M3
- GLM 5.3
- GLM 5.3 Flash
If inference speed is on your application's critical path, reach out at
[contact@inco.ai](mailto:contact@inco.ai).
Get updates
One email when we ship something new.
We will never share your email address.