Xiaomi's new Xuanjie chips actually deliver on the local AI hype Xiaomi's new Xuanjie O100 chip, demonstrated in two engineering prototypes, delivers on-device AI performance with the foldable prototype achieving 300+ tokens per second (peak 330 tokens/s) running the Xiaomi MiMo 3B model offline, and the Xiaomi AI Cube running a 120B parameter model locally with 80 GB RAM (up to 160 GB). The prototypes showcase a tiered compute model that could reduce reliance on cloud-based LLMs, according to hands-on testing by XRING LAB. Xiaomi's new Xuanjie chips actually deliver on the local AI hype I managed to get hands-on with two engineering prototypes that demonstrate exactly how this silicon translates into a real-world AI workflow. The Fan-Assisted Foldable: 300+ tokens per second on a phone The first prototype is a foldable device powered by the Xuanjie O100. It looks like a standard foldable, but the back is a heavy-duty metal chassis featuring a circular intake/exhaust vent with an active cooling fan built right in. The goal here wasn't just to make a phone, but to create a mobile AI powerhouse. Running the Xiaomi MiMo 3B model entirely offline no Wi-Fi, no cellular , the responsiveness is staggering. Using the XRING LAB testing tool, I clocked a Time to First Token TTFT of roughly 0.45 seconds. The generation speed is where it gets wild: Average Generation Speed: ~303 tokens/s Peak Generation Speed: Up to 330 tokens/s At these speeds, the AI is literally writing faster than you can read. For anyone interested in prompt engineering or real-time translation, this is the holy grail. You get instant, "zero-latency" interaction for document polishing or live transcription, all while maintaining absolute data privacy because the processing happens strictly on-device. The AI Cube: Running a 120B model on your desk If the foldable is about mobility, the "Xiaomi AI Cube" is about brute-force local compute. This thing is a CNC-machined piece of aerospace aluminum, designed to dissipate massive amounts of heat through 33,874 precision-cut cooling holes. This isn't your typical mini-PC. It’s an Android-based AI terminal with 80 GB of RAM supporting up to 160 GB specifically built to deploy 120B parameter models locally. During the demo, we tested its coding capabilities: 1. Task Input: A request to generate a functional piano web application. 2. Processing: The system utilized a dual-model switching strategy 3B + 120B . 3. Output: The 120B model handled the heavy lifting of logic and code structure. 4. Result: Within moments, a fully functional "Web Audio Online Piano" was running on the screen, complete with real-time audio waveforms and interactive keys. The "small model + large model" workflow is the clever part of this deployment. The 3B model handles high-frequency, lightweight commands instantly so the UI feels snappy, while the 120B model is summoned only when the task requires deep reasoning or complex coding. It effectively turns a desktop tool into a local LLM agent that doesn't need a cloud connection to function. Hardware-driven AI differentiation Seeing these two prototypes side-by-side clarifies the future of the AI workflow. We are moving toward a tiered compute model: Mobile/Foldable Tier: Focused on high memory bandwidth thanks to the O100's 1.22 TB/s bandwidth to enable instant, offline interaction for daily tasks. Desktop/Cube Tier: Focused on massive parameter counts and sustained thermal performance to handle heavy-duty reasoning and development. The bottleneck for local AI has always been the trade-off between model intelligence and latency. By optimizing the silicon specifically for these workloads, Xiaomi is showing that we might finally be able to move away from cloud-dependent LLMs toward a truly private, local AI ecosystem. Apple is leaking its own hardware again 5d ago /en/news/6881/ Next Why AI-driven delusions follow a predictable spiral pattern → /en/news/7592/