GLM-5.3-Flash at 1000 tok/s on RTX PRO 6000 Zhipu AI's GLM-5.3-Flash model achieved 1000 tokens per second output speed on an Nvidia RTX PRO 6000 GPU, according to LocalMaxxing benchmark data. The result highlights the model's high-speed inference capability on professional-grade hardware. LocalMaxxing Get started Models Reports Hardware Benchmarks More + Submit Get started Leaderboard Decode calculator Models Reports Hardware Benchmarks Marketplace Rentals Pro API Docs Language English 简体中文 繁體中文 日本語 한국어 Español Français Deutsch Italiano Português Brasil Русский Polski Nederlands Türkçe हिन्दी Bahasa Indonesia Tiếng Việt ไทย Sponsor LocalMaxxing Your ad here Reach local AI builders Total runs Highest Best memory ceiling Median Lowest Models zai-org GLM-5.3-Flash Total runs Highest Best memory ceiling Median Lowest Speed Tests Benchmarks Reports Speed Test Results Submit benchmark Advanced filters Hardware Engine Depth tok/s out Prefill tok/s total TTFT ms VRAM GB