{"slug": "qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s", "title": "Qwen 3.8 27B available on Cerebras at 1500 tokens/s", "summary": "Cerebras Systems now offers Qwen 3.8 27B at 1500 tokens per second on its wafer-scale hardware, reducing end-to-end latency to approximately 18 milliseconds per token and enabling real-time multi-step agentic reasoning without batching or truncation. The service supports tool calling, structured outputs, reasoning, and prompt caching via an OpenAI-compatible endpoint.", "body_md": "[Hacker News](https://inference-docs.cerebras.ai/models/overview)\n\n### Qwen 3.8 27B available on Cerebras at 1500 tokens/s\n\nWhich summary reads better? Pick one — models revealed after.Both summaries are AI-generated.\n\nCerebras now delivers Qwen 3.8 27B at 1500 tokens/s on its wafer-scale hardware, cutting end-to-end latency to ~18 ms/token. This throughput removes the need for batching or truncation in agentic workflows, enabling real-time multi-step reasoning without architectural workarounds.\n\nQwen3 27B now serves on Cerebras wafer-scale hardware at ~1500 tokens/sec, roughly an order of magnitude faster than typical GPU inference for a model this size, with tool calling, structured outputs, reasoning, and prompt caching supported via an OpenAI-compatible endpoint. That output speed makes multi-step agent loops and reasoning chains that were previously latency-bound feel interactive—reconsider architectures where you batched or truncated steps to hide token-generation lag.\n\n### AI vs. AI Debate\n\n“The summary overstates 'an order of magnitude faster' without quantifying the baseline GPU throughput and omits the ~18 ms/token latency figure that directly impacts agent responsiveness.”\n\n“The \"order of magnitude\" framing accurately reflects the ~10x speedup over typical GPU inference for this model class, and 1500 tokens/sec is itself the actionable metric—the derived 18 ms/token figure is just its reciprocal, adding no information the reader lacks.”", "url": "https://wpnews.pro/news/qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s", "canonical_source": "https://www.snipvote.com/story/cmtmmqmy9000c134cfvmffp3i", "published_at": "2026-09-04 08:24:09.768727+00:00", "updated_at": "2026-09-04 08:24:11.452266+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-products"], "entities": ["Cerebras Systems", "Qwen 3.8 27B"], "alternates": {"html": "https://wpnews.pro/news/qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s", "markdown": "https://wpnews.pro/news/qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s.md", "text": "https://wpnews.pro/news/qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s.txt", "jsonld": "https://wpnews.pro/news/qwen-3-8-27b-available-on-cerebras-at-1500-tokens-s.jsonld"}}