cd /news/large-language-models/glm-5-3-flash-at-1000-tok-s-on-rtx-p… · home topics large-language-models article
[ARTICLE · art-119102] src=localmaxxing.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

GLM-5.3-Flash at 1000 tok/s on RTX PRO 6000

Zhipu AI's GLM-5.3-Flash model achieved 1000 tokens per second output speed on an Nvidia RTX PRO 6000 GPU, according to LocalMaxxing benchmark data. The result highlights the model's high-speed inference capability on professional-grade hardware.

read1 min views1 publishedSep 2, 2026
GLM-5.3-Flash at 1000 tok/s on RTX PRO 6000
Image: source

LocalMaxxing Get started Models Reports Hardware Benchmarks More + Submit Get started Leaderboard Decode calculator Models Reports Hardware Benchmarks Marketplace Rentals Pro API Docs Language English 简体中文 繁體中文 日本語 한국어 Español Français Deutsch Italiano Português (Brasil) Русский Polski Nederlands Türkçe हिन्दी Bahasa Indonesia Tiếng Việt ไทย Sponsor LocalMaxxing Your ad here Reach local AI builders Total runs Highest Best memory ceiling Median Lowest Models zai-org GLM-5.3-Flash Total runs Highest Best memory ceiling Median Lowest Speed Tests Benchmarks Reports Speed Test Results Submit benchmark Advanced filters

#

Hardware Engine Depth tok/s out Prefill tok/s total TTFT ms VRAM GB

── more in #large-language-models 4 stories · sorted by recency
── more on @zhipu ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/glm-5-3-flash-at-100…] indexed:0 read:1min 2026-09-02 ·