# Z.ai adds GLM-5.3 to Amazon Bedrock for enterprise coding agents

> Source: <https://runtimewire.com/article/zai-glm-5-3-amazon-bedrock-enterprise>
> Published: 2026-10-06 01:01:36+00:00

# Z.ai adds GLM-5.3 to Amazon Bedrock for enterprise coding agents

**AWS launched the model on October 5th with managed APIs, prompt caching and cross-Region inference; access is limited to eligible customers.**

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [Z.ai on X](https://x.com/Zai_org/status/2107272032073437185)

## Why it matters

Bedrock gives Z.ai a path into enterprise workloads through AWS-managed APIs, while cross-Region-only routing and restricted eligibility put concrete limits on that convenience.

Z.ai's [GLM-5.3](https://runtimewire.com/models/z-ai/glm-5.3:batch) became available through Amazon Bedrock on October 5th, giving eligible enterprise customers a managed way to run the model for coding and agent workflows. Z.ai [posted the availability on X](https://x.com/Zai_org/status/2107272032073437185) the following day; [AWS's launch post](https://aws.amazon.com/blogs/machine-learning/introducing-glm-5-3-on-amazon-bedrock/) and [Bedrock model documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-zai-glm-5-3.html) describe the service and its restrictions.

The release extends the work of Z.ai co-founder [Tang Jie](https://keg.cs.tsinghua.edu.cn/persons/jietang/), a Tsinghua computer science professor whose research has included knowledge graphs and AMiner, an academic search and mining system. Z.ai grew from a research lineage; this Bedrock launch puts one of its models into a cloud marketplace where enterprise buyers can access it through AWS APIs instead of arranging their own inference deployment. Z.ai's chief executive, Zhang Peng, is also a co-founder, according to the company's [Hong Kong listing prospectus](https://www1.hkexnews.hk/listedco/listconews/sehk/2025/1230/2025123000017.pdf).

GLM-5.3 is aimed at repository-scale coding, software engineering and multi-step agent tasks. Bedrock supports it through OpenAI-compatible Responses and Chat Completions APIs, as well as AWS's Invoke and Converse interfaces. Customers can try it in the Bedrock console or connect it to their applications. AWS also supports implicit and explicit prompt caching, intended to lower repeated-input latency and costs when an agent resends long prompts or code context.

The deployment has a consequential boundary for companies with strict data-location requirements: AWS says GLM-5.3 is available only through cross-Region inference, not in-Region inference. Customers can select a US geographic profile or a global profile. The US option routes requests among supported US regions; the global option can route requests worldwide. That gives buyers a choice of routing scope, but the model is not confined to a single AWS region. Access is also restricted to eligible customers, rather than offered as an unrestricted model choice to every Bedrock account.

The product specifications are inconsistent across the launch materials. AWS's announcement and the [Z.ai model page on Hugging Face](https://huggingface.co/zai-org/GLM-5.3) describe a 753-billion-parameter mixture-of-experts model. AWS's Bedrock model card instead lists 744 billion total parameters, with about 40 billion active per token. The AWS documentation describes a one-million-token context window and up to 128,000 output tokens. Those details matter to customers estimating hardware demands and workload fit, even though Bedrock abstracts away the inference infrastructure.

GLM-5.3 is already available as open weights through Hugging Face, where developers can deploy it with supported software frameworks. Bedrock's commercial proposition is different: AWS handles the serving layer and exposes the model through an existing enterprise cloud account, API setup and service-tier structure. Customers can use standard pay-per-token access, a lower-cost Flex tier for less time-sensitive tasks, or Priority for latency-sensitive workloads. AWS's model card directs buyers to Bedrock's general pricing page rather than listing a GLM-5.3-specific rate.

Z.ai says GLM-5.3 improved 50% over [GLM-5.2](https://runtimewire.com/models/z-ai/glm-5.2) on the company's internal coding benchmark. It also reported an 84.5 score on CyberGym, a vulnerability-discovery benchmark. Those are company-reported results, not independent evaluations. AWS's demonstration pairs the model with Strix, an open-source penetration-testing agent, to test an application the user owns or has permission to assess. The demo is a worked example, not evidence that the model independently secures production software.

For Z.ai, Bedrock offers access to AWS's enterprise distribution channel without requiring customers to build a model-serving stack. For AWS, adding a Chinese-developed open-weight model broadens Bedrock's catalog for teams comparing coding and agent models. The test for both is whether customers choose GLM-5.3 for real workloads once they weigh its performance claims against routing rules, access eligibility and the eventual per-token bill.
