{"slug": "qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu", "title": "Qualcomm Talks Next-Gen Oryon CPU, Adreno GPU, and Hexagon NPU", "summary": "Qualcomm disclosed details of its next-generation premium Snapdragon platform's Oryon CPU, Adreno GPU, and Hexagon NPU ahead of its Snapdragon Summit later this month. The Oryon CPU's two Prime cores are rated at 5 GHz, which Qualcomm says is the first mobile CPU to reach that frequency, while the Adreno GPU's three slices clock at 1.45 GHz and the Hexagon NPU adds an Element Accelerator, 50 percent more shared memory, and support for Mixture of Experts models up to 30 billion total parameters with roughly 3 billion active parameters per token. Qualcomm reports up to 50 percent prefill uplift for INT4 models and 12 percent power improvement over the Snapdragon 8 Elite Gen 5 baseline, with 40 percent power savings in its internal Dragon Alley demo when Neural Fusion is enabled.", "body_md": "Ahead of its Snapdragon Summit later this month, Qualcomm disclosed the Oryon CPU, Adreno GPU, and Hexagon NPU for its next premium mobile Snapdragon platform. The NPU adds an Element Accelerator and larger shared memory, positioning it alongside the previously detailed Oryon CPU with 5 GHz Prime cores and FlexCache, and the Adreno GPU with Matrix Cores. Since it is the agentic AI era, Qualcomm is talking about the platform in terms of how AI runs on it in a bit of a teaser.\n\n## Qualcomm Oryon CPU\n\nThe Oryon CPU orchestrates agentic work, coordinating planning, tool calls, and accelerator tasks. Our [Computex 2025 coverage](https://www.servethehome.com/qualcomm-computex-2025-live-coverage/) anticipated that newer mobile Oryon designs would move into PC chips, though the current architecture here remains mobile-focused.\n\nThe two Prime cores in the next-generation premium Snapdragon platform are rated at 5 GHz, a figure the company says marks the first mobile CPU to reach that frequency. Qualcomm attributes the design to its custom CPU microarchitecture, implementation, and subsystem.\n\nFlexCache is a dynamically allocated cache pool shared by heterogeneous CPU cores on the new platform.\n\nPrime cores can draw on the entire pool as workloads demand, keeping larger working sets cached and reducing system-memory accesses. Hopefully we will get more CPU details soon.\n\n## Qualcomm Adreno GPU\n\nThe Adreno GPU comprises three slices clocked at 1.45 GHz, a command processor, and one 18 MB Adreno High Performance Memory (HPM) block. HPM serves as local graphics storage for working data such as tiles and frame buffers, reducing traffic to system memory.\n\nQualcomm says this delivers a 12% power improvement over the Snapdragon 8 Elite Gen 5 baseline.\n\nEach of the three GPU slices contains Matrix Cores, which bring dedicated matrix and AI processing into the graphics pipeline. Adreno High Performance Memory (HPM) keeps working data nearby, reducing reliance on external memory and improving efficiency.\n\nNeural Fusion provides AI rendering and super-resolution, integrating with Unity and Unreal Engine upscaling frameworks.\n\nWith Neural Fusion enabled, Qualcomm reports a 40 percent power savings in its internal Dragon Alley demo.\n\n## Qualcomm Hexagon NPU\n\nThe NPU features a new Element Accelerator for transformer operations, plus vector and scalar extensions that handle AI math and agent decision or routing tasks. It supports context lengths up to 32K and includes KV-cache acceleration.\n\nShared memory grows by 50 percent, according to Qualcomm, though the company does not disclose the absolute capacity. Positioned alongside the tensor, vector, scalar, and element compute regions, this block keeps model state, context, and KV-cache near the accelerators, thereby reducing external-memory movement.\n\nFor INT4 models, Qualcomm reports up to 50 percent prefill uplift on its next-generation premium mobile Snapdragon platform versus the Snapdragon 8 Elite Gen 5. Prefill processes the input prompt before the model generates subsequent tokens.\n\nMixture of Experts (MoE) models up to 30 billion total parameters run on the Hexagon NPU, with roughly 3 billion active parameters routed per token generation step. Flash-to-memory expert management and caching handle the model.\n\nThe Sensing Hub captures voice input and routes it to the Oryon CPU, which orchestrates task execution.\n\nInference workloads are dispatched to the GPU or NPU, with a cloud path available for offloading.\n\n## Final Words\n\nThe common thread is keeping data and computation close to the processing resources that use them, reducing reliance on slower system memory. We expect fairly substantial gains when this line finally makes its way into products. Hopefully we get to see them soon and also get a lot more detail on the chips.", "url": "https://wpnews.pro/news/qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu", "canonical_source": "https://www.servethehome.com/qualcomm-details-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu/", "published_at": "2026-09-10 13:05:36+00:00", "updated_at": "2026-09-10 13:38:16.054628+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "ai-products", "large-language-models", "ai-agents"], "entities": ["Qualcomm", "Oryon CPU", "Adreno GPU", "Hexagon NPU", "Snapdragon 8 Elite Gen 5", "FlexCache", "Neural Fusion", "Adreno High Performance Memory"], "alternates": {"html": "https://wpnews.pro/news/qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu", "markdown": "https://wpnews.pro/news/qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu.md", "text": "https://wpnews.pro/news/qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu.txt", "jsonld": "https://wpnews.pro/news/qualcomm-talks-next-gen-oryon-cpu-adreno-gpu-and-hexagon-npu.jsonld"}}