cd /news/artificial-intelligence/qwen-3-8-27b-available-on-cerebras-a… · home topics artificial-intelligence article
[ARTICLE · art-120675] src=inference-docs.cerebras.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Qwen 3.8 27B available on Cerebras at 1500 tok/SEC

Cerebras Systems announced that the Qwen 3.8 27B model is now available on its platform at 1500 tokens per second, with all public models served unpruned and using selective weight-only quantization for storage. The company clarified that it does not alter model architectures via pruning on hosted endpoints, and its REAP pruned models are available only on Hugging Face for research.

read2 min views2 publishedSep 3, 2026
Qwen 3.8 27B available on Cerebras at 1500 tok/SEC
Image: source

rate limitsand pricing. For additional model families, reserved capacity, higher throughput, and production SLAs, see

Dedicated Endpoints.

Available Models #

Model Compression #

This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints. All models served through our public endpoints are the original, unpruned versions. While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through our shared API. You can read more about REAP in ourresearch blog.

All of our public models are unpruned. Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized.

Frequently Asked Questions

Will you change a model's architecture without notice?

Will you change a model's architecture without notice?

No. We are committed to serving the original models for all existing endpoints, without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques (like pruning) in the future, these would be offered as separate endpoints with pruning-specific names, ensuring complete transparency and allowing you to choose which version best fits your needs.

Where can I find your REAP pruned models?

Where can I find your REAP pruned models?

Our REAP pruned models are available on Hugging Face for research and experimentation purposes:

Cerebras REAP Collection. These models demonstrate our pruning research but are not served through our production API.What are compression, quantization, and pruning?

What are compression, quantization, and pruning?

Compression is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include:

Quantization: Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture.Pruning: Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cerebras systems 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen-3-8-27b-availab…] indexed:0 read:2min 2026-09-03 ·