Operating System powered by Qwen 3.8 27B at 1950 tokens/SEC A developer built a minimal Python web server that turns Cerebras-hosted inference of Alibaba's Qwen 3.8 27B model into a live operating system with zero apps on disk, streaming tokens at 1,950 tokens per second. In the demonstration, a calculator compiled and mounted in 11 seconds and a full canvas paint studio with a brush engine compiled in 10 seconds, with the developer stating "the model weights ARE the operating system runtime." The post cites Cerebras inference pricing of $0.99 per million input tokens and $1.49 per million output tokens, and compares the 1,850 tok/s Cerebras speed against 75 tok/s for a local 4-bit RTX 3090/4090. Operating System powered by Qwen 3.8 27B at 1950 tokens/sec here is what 1,950 tokens/second @Alibaba Qwen https://x.com/Alibaba Qwen 's 3.8 27b actually looks like on @cerebras https://x.com/cerebras : i wrote a minimal python web server that turns cerebras inference into a live operating system. zero apps on disk. when you double click an icon: 1 python proxies a raw SSE stream from qwen 27b at 1,950 tok/s 2 calculator compiles & mounts in 11s 3 full canvas paint studio with brush engine compiles in 10s. at 2,000 tokens/second, software is just an on demand hallucination that runs instantly. the model weights ARE the operating system runtime. what else would you build at 1,950 tokens/second? 00:00 qwen 3.8 27b cerebras inference math: speed: 1,850 tok/s cost: $0.99/M in | $1.49/M out intelligence drops like a database query. standard 60Hz screens refresh every 16.6ms. - local 3090/4090 4-bit : 75 tok/s 1.2 tokens/frame - cerebras qwen 3.8 27b: 1,850 tok/s 31