# Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models

> Source: <https://www.marktechpost.com/2026/10/02/prime-intellect-launches-prime-inference-serverless-and-reserved-serving-for-frontier-open-models/>
> Published: 2026-10-03 05:37:18+00:00

**Prime Intellect has launched [Prime Inference](https://www.primeintellect.ai/blog/prime-inference), a serving platform for frontier open-source models.** It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters. Before public release, it processed nearly a trillion tokens per day internally. That traffic came from RL rollouts, synthetic data generation, evaluations and long-running coding agents.

## **What is Prime Inference?**

Prime Inference is the serving layer of Prime Intellect’s open training stack. The company already ships post-training tools such as prime-rl, verifiers and sandboxes. Serving closes that loop: deployed models generate production traces that can feed back into training. Prime reports its [GLM-5.3](https://huggingface.co/zai-org/GLM-5.3) endpoint ranks among the fastest on OpenRouter. It also cites a near-zero tool-call error rate and 100% uptime since launch.

- **Two modes:** serverless endpoints for variable demand, reserved capacity for sustained workloads.
- **OpenAI compatible:** point any OpenAI SDK at`https://api.pinference.ai/api/v1` ([docs](https://docs.primeintellect.ai/inference/overview) ).
- **Uptime:** automatic failover across datacenters routes traffic to healthy deployments.
- **Hardware:** NVIDIA Blackwell today, with Vera Rubin listed as coming soon.
- **Billing:** unified billing with team-level usage tracking. Per-model pricing is not yet fully published in the docs.

## **How the serving stack works**

The stack combines [NVIDIA Dynamo](https://github.com/ai-dynamo/dynamo), [vLLM](https://github.com/vllm-project/vllm), [Mooncake](https://github.com/kvcache-ai/Mooncake) and [FlashInfer](https://github.com/flashinfer-ai/flashinfer). It was built with Inferact and NVIDIA, and fixes are contributed upstream.

The target workload is agentic. A typical agent turn adds about 6K tokens to a 140K-token prompt. Prime benchmarks this mix with SemiAnalysis [AgentX](https://inferencex.semianalysis.com/agentx), and injected cold arrivals.

**Prefill/decode disaggregation**: Prefill and decode run on separate GPU groups. Dynamo handles routing, and vLLM runs the model on each group. Decoders pull computed KV through [NIXL](https://github.com/ai-dynamo/nixl). Prime reports nearly 40% lower p90 inter-token latency in its tests.

**Cache-aware routing**:Dynamo’s KV-aware router weighs cached prefix overlap against queued work. Sessions stay on the same decoder between turns. Mooncake adds a second KV tier in host DRAM.

## **GLM-5.3 on GB200 NVL72: the numbers**

The interactivity target was 100 end-to-end tokens per second per user. At that bar, a 1:4 prefill/decode ratio served the most users. It reached 66 sessions per prefill group at 101 tok/s per user and 100 output tok/s per GPU.

- **DEP8 prefill topology:** roughly 5x more usable prefix-cache capacity than TEP8 on the same hardware.
- **Smaller prefill budget:** halving tokens per step from 8K to 4K per GPU cut median queue wait from 550 ms to 110 ms. Median time to first token fell about 20%.
- **NVFP4 KV compression:** each MLA cache row shrank from 576 to 352 bytes. Cached tokens per decoder rose from 1.09M to 1.63M.
- **Native sparse-MLA kernel:** about 12.0 μs at 15 query tokens, versus 17.7 μs staged and 13.7 μs FP8. Prime notes this is workload specific.
- **BLHNC KV layout:** transfer descriptors fell from 19,559 to about 1,940. Mean transfer time dropped from 146 ms to 78 ms.

## **Reliable tool calls**

Agents fail when tool calls carry wrong names or broken arguments. Prime Intellect’s team contributed a structural-tag builder to Dynamo for GLM’s tool format. vLLM then uses [xgrammar](https://github.com/mlc-ai/xgrammar) to mask tokens that violate the tool schema. The team also fixed parsing bugs, including `<` being decoded into `<` inside code.

## **Interactive explainer**

## **Prime Inference vs closest competitors**

| Feature | Prime Inference | Together AI | Fireworks AI | Baseten | 
|---|---|---|---|---|
| Serverless GLM-5.3 | Yes ( [source](https://www.primeintellect.ai/blog/prime-inference) ) | Yes ( [source](https://docs.together.ai/docs/glm-5.3-quickstart) ) | Yes ( [source](https://artificialanalysis.ai/providers/fireworks) ) | Yes ( [source](https://www.baseten.co/resources/changelog/glm-53-available-on-baseten/) ) | 
| GLM-5.3 price, input / output per 1M tokens | Not yet published in docs | $1.40 / $4.40 ( [source](https://docs.together.ai/docs/serverless-models) ) | $1.40 / $4.40 ( [tracker](https://computeprices.com/providers/fireworks-ai/models/glm-5-3) ) | $1.40 / $4.40 ( [tracker](https://computeprices.com/providers/baseten/models/glm-5-3) ) | 
| Dedicated or reserved capacity | Reserved capacity; 1-click dedicated deploys on roadmap | Dedicated Model endpoints ( [source](https://www.together.ai/models-providers/zai-org) ) | On-demand dedicated GPU deployments ( [source](https://fireworks.ai/models/fireworks/glm-5p3-flash) ) | Dedicated GPU deployments ( [source](https://computeprices.com/providers/baseten/models/glm-5-3) ) | 
| OpenAI-compatible API | Yes | Yes | Yes | Yes | 
| Batch inference | On roadmap | Yes ( [source](https://www.together.ai/models-providers/zai-org) ) | Not compared here | Not compared here | 
| Disclosed serving stack | Open source: Dynamo, vLLM, Mooncake, FlashInfer | Together inference research stack | Fireworks serving stack | Baseten Inference Stack ( [source](https://huggingface.co/docs/inference-providers/providers/baseten) ) | 

*Competitor prices verified October 2, 2026. Tracker figures come from ComputePrices, a third-party price tracker.*

## **Key Takeaways**

- Prime Inference is live with serverless and reserved serving for open models.
- GLM-5.3 runs on GB200 NVL72 with Dynamo, vLLM, Mooncake and FlashInfer.
- 1:4 prefill/decode served 66 sessions per prefill group at 101 tok/s per user.
- NVFP4 KV cache lifted capacity from 1.09M to 1.63M tokens per decoder.
- Batch inference and 1-click dedicated deploys are next on the roadmap.

Check out the [**technical details**](https://www.primeintellect.ai/blog/prime-inference), [** docs**](https://docs.primeintellect.ai/inference/overview) and the [** announcement on X**](https://x.com/PrimeIntellect/status/2106146483003384253). All credit goes to the researcher of this project. Also, feel free to follow us on **[Twitter](https://x.com/intent/follow?screen_name=marktechpost)** and don’t forget to join our **[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)** and Subscribe to **[our Newsletter](https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}})**. Wait! are you on telegram? [now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? [Connect with us](https://forms.gle/MJjjVDPS7whH8Ngs6)

Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.
