Operating System powered by Qwen 3.8 27B at 1950 tokens/sec! here is what 1,950 tokens/second
@Alibaba_Qwen's 3.8 27b actually looks like on@cerebras: i wrote a minimal python web server that turns cerebras inference into a live operating system. zero apps on disk. when you double click an icon: 1) python proxies a raw SSE stream from qwen 27b at 1,950 tok/s 2) calculator compiles & mounts in 11s 3) full canvas paint studio with brush engine compiles in 10s. at 2,000 tokens/second, software is just an on demand hallucination that runs instantly. the model weights ARE the operating system runtime. what else would you build at 1,950 tokens/second? 00:00
qwen 3.8 27b cerebras inference math: speed: 1,850 tok/s cost: $0.99/M in | $1.49/M out intelligence drops like a database query. standard 60Hz screens refresh every 16.6ms.
- local 3090/4090 (4-bit): 75 tok/s (1.2 tokens/frame)
- cerebras qwen 3.8 27b: 1,850 tok/s (31