{"slug": "inco-ai-launches-day-0-support-for-glm-5-3", "title": "Inco AI launches Day-0 support for GLM 5.3", "summary": "Inco AI launched day-0 support for Z.ai's GLM 5.3 model, releasing DFlash 2 and NVFP4 checkpoints alongside an Inco Engine endpoint that delivers up to 4.4× throughput versus the native FP8 checkpoint with autoregressive decoding at concurrency 1. The NVFP4 checkpoint matches FP8 accuracy across evaluations, and the endpoint is priced at $1.40 per 1M input tokens, $0.26 per 1M cached input tokens, and $4.40 per 1M output tokens.", "body_md": "# Inco AI launches Day-0 support for GLM 5.3\n\n[Overview](#overview)\n\n[GLM 5.3](https://z.ai/blog/glm-5.3) arrives today with day-0 serving support\nfrom Inco AI on\n[TokenRouter](https://www.tokenrouter.com/) compute. The\nrelease pairs [Z.ai's](https://z.ai/) new model with our core inference\noptimization technology:\n\n- A\n[DFlash 2 checkpoint](https://huggingface.co/incoai/GLM-5.3-DFlash2)for speculative decoding. - An\n[NVFP4 checkpoint](https://huggingface.co/incoai/GLM-5.3-NVFP4)for efficient Blackwell inference with native accuracy. [Inco Engine](https://platform.inco.ai/playground?model=glm-5.3)for end-to-end inference performance.\n\nAs a day-0 launch partner, we are releasing both checkpoints alongside the\nmodel endpoint. Together, the checkpoints and Inco\nEngine deliver up to **4.4× throughput** versus the native\n[FP8 checkpoint](https://huggingface.co/zai-org/GLM-5.3) with autoregressive\ndecoding at concurrency 1.\n\n[DFlash 2 for GLM 5.3](#dflash-2-for-glm-53)\n\n[DFlash 2](/blog/dflash2/) is a parallel drafter for speculative decoding: it\npredicts candidate tokens in one pass and lets the target model verify them as\na block.\n\nAcceptance length (AL) measures how many tokens each draft–verify cycle yields, including the verifier's next token. A longer accepted block means fewer full target-model passes for the same output. DFlash 2 improves AL by keeping the parallel draft design while selecting a more coherent path through each position's candidates.\n\nAcceptance length is only half of the serving result: a drafter also adds work to each cycle. For GLM 5.3, we therefore report end-to-end decode throughput separately, comparing the model's native MTP path and DFlash 2 against autoregressive decoding. At concurrency 1, DFlash 2 reaches 383.3 output tok/s on MATH-500, 366.6 on GSM8K, and 363.7 on HumanEval.\n\n[NVFP4 for Blackwell](#nvfp4-for-blackwell)\n\nWe are also releasing an\n[NVFP4 checkpoint](https://huggingface.co/incoai/GLM-5.3-NVFP4) for efficient\nBlackwell deployment, with accuracy matching the native\n[FP8 checkpoint](https://huggingface.co/zai-org/GLM-5.3) across the evaluation\nsuite.\n\n| Precision | GPQA Diamond | AIME 2025 | MATH-500 | HLE | AA-LCR |\n|---|---|---|---|---|---|\n| FP8 | 91.1 | 94.3 | 95.6 | 35.9 | 73.6 |\n| NVFP4 | 91.2 | 95.1 | 95.2 | 35.2 | 73.0 |\n\n[Get access](#get-access)\n\nTry GLM 5.3 in the browser or connect through the Inco API.\n\n### Included in preview\n\n- Browser playground\n- OpenAI and Anthropic API dialects\n- Streaming responses\n- 1M-token context window\n\n### List pricing\n\nper 1M tokens- Input\n- $1.40\n- Cached input\n- $0.26\n- Output\n- $4.40\n\n[Get checkpoints](#get-checkpoints)\n\n[Compute for GLM 5.3](#compute-for-glm-53)\n\nCompute for this launch came from\n[TokenRouter](https://www.tokenrouter.com/). We trained the DFlash 2 drafter\nand produced the NVFP4 checkpoint on its Blackwell clusters, and we serve the\nday-0 endpoint from the same pool. You can reach the endpoint with an existing\nTokenRouter key.\n\n[The bottom line](#the-bottom-line)\n\nTen days ago, we called [DFlash 2](/blog/dflash2/) the first piece of an\nend-to-end serving stack. Launches like this are what the rest of it is for.\n\nGLM 5.3 shipped today, and the stack shipped with it: a DFlash 2 drafter, an NVFP4 checkpoint, and a live endpoint on Inco Engine. Together they serve the model at up to 4.4× the throughput of the native FP8 checkpoint with autoregressive decoding, measured at concurrency 1.\n\nIf you want your model served like this, fine-tunes included, write to us:\n[contact@inco.ai](mailto:contact@inco.ai).\n\n[Get updates](#get-updates)\n\nOne email when we ship something new.\n\nWe will never share your email address.", "url": "https://wpnews.pro/news/inco-ai-launches-day-0-support-for-glm-5-3", "canonical_source": "https://inco.ai/blog/glm-5-3/", "published_at": "2026-08-28 00:00:00+00:00", "updated_at": "2026-08-28 15:51:24.161124+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-products"], "entities": ["Inco AI", "Z.ai", "GLM 5.3", "TokenRouter", "DFlash 2", "NVFP4", "Inco Engine", "FP8"], "alternates": {"html": "https://wpnews.pro/news/inco-ai-launches-day-0-support-for-glm-5-3", "markdown": "https://wpnews.pro/news/inco-ai-launches-day-0-support-for-glm-5-3.md", "text": "https://wpnews.pro/news/inco-ai-launches-day-0-support-for-glm-5-3.txt", "jsonld": "https://wpnews.pro/news/inco-ai-launches-day-0-support-for-glm-5-3.jsonld"}}