# Qualcomm’s next snapdragon mobile chip comes into focus with on-device AI

> Source: <https://www.computerworld.com/article/4220719/qualcomms-next-snapdragon-mobile-chip-comes-into-focus-with-on-device-ai.html>
> Published: 2026-09-10 13:01:00+00:00

Qualcomm is taking an unusual approach to the launch of its next-generation premium Snapdragon mobile platform this year. Instead of holding everything for its annual Snapdragon Summit in Maui happening later this month, the company has been methodically disclosing the architecture that will power the next wave of flagship Android phones over the past couple of weeks.

First, the company disclosed its new Oryon CPU architecture, followed by a major overhaul of its Adreno GPU. Today, Qualcomm is detailing the next-generation Hexagon NPU that will serve as the platform’s primary AI engine.

Taken together, the disclosures give us a pretty good picture of where Qualcomm is going before we even know the chip’s official name.

There is obviously a strong performance story here, anchored by a 5GHz CPU, but the more interesting theme running through the architecture is memory locality and specialized AI acceleration. Qualcomm is trying to keep more data close to the compute resources that need it, while distributing AI workloads across the CPU, GPU and NPU.

These architectural changes are important as the company continues to drive on-device AI beyond relatively simple generative features toward more persistent, agentic workloads.

Qualcomm disclosed in August that its next Oryon CPU will be the first mobile CPU to reach 5GHz. The eight-core design includes two 5GHz Prime cores and six Performance cores, with Qualcomm attributing the frequency gains to its fully custom Oryon microarchitecture, an in-house Arm-based CPU design that gives the company full control over the cores, cache and memory hierarchy, branch prediction and CPU subsystem.

The 5GHz headline is certainly going to get attention. However, Qualcomm’s new [FlexCache architecture](https://hothardware.com/news/qualcomm-5ghz-mobile-milestone-oryon-cpu-flexcache-architecture) may ultimately have a larger impact on sustained performance. FlexCache allows Oryon’s Prime and Performance cores to tap into the same shared cache pool, with capacity dynamically allocated based on the workload. So rather than a demanding core being limited to a fixed slice of cache, it can draw on more of the available pool when needed. This keeps larger working data sets close to the CPU cores and reduces high-latency trips out to system memory.

This approach should benefit everything from gaming and multitasking to content creation, but it also aligns well with agentic AI workloads, where tasks may move through several stages of processing and across different CPU cores. The basic goal is straightforward: keep the CPU fed and avoid expensive trips to external memory.

Qualcomm’s next disclosure focused on graphics, where the company is making one of the more significant changes to its Adreno GPU engine in recent years. The new GPU adds dedicated Adreno Matrix Cores for running AI models directly inside the graphics pipeline. They’re paired with 18MB of Adreno High-Performance Memory, or HPM, which provides low-latency local storage for tile-based rendering, frame buffers and GPU compute.

Qualcomm claims HPM provides a 12% power improvement compared to its previous-gen Snapdragon 8 Elite Gen 5.

The more visible feature for users, however, may be [Adreno Neural Fusion](https://www.qualcomm.com/news/onq/2026/09/adreno-neural-fusion-ai-rendering), which combines AI super resolution, neural processing and frame generation in a unified graphics pipeline. There are obvious parallels here to what NVIDIA, AMD and others are doing with neural rendering on the PC, though Qualcomm’s next-gen Snapdragon SoC has to operate within a much tighter mobile power envelope.

Qualcomm claims enabling Neural Fusion reduces power consumption by up to 40% in its internal testing using the company’s Dragon Alley graphics demo. That’s an impressive number, though as always, we’ll want to see how the technology behaves across actual shipping games and devices before drawing any definitive conclusions.

Qualcomm has also helped integrate Neural Fusion technology into Unity and Unreal Engine, which may ultimately be just as important as the underlying hardware. Developer adoption determines whether features like this become meaningful platform advantages or just another tick box on a spec sheet.

The most recent disclosure centers on Qualcomm’s Hexagon NPU, where the broader architecture starts to come together. The company’s next-generation Hexagon introduces a new Element Accelerator designed for transformer workloads, alongside its existing vector and scalar processing resources. Qualcomm says the new block is optimized for fast action loops, key-value (KV) cache acceleration and context lengths of up to 32K.

Hexagon is also getting 50% more shared memory, and again, that memory complement is important. Model state, context and KV-cache data can generate substantial memory traffic, particularly with larger language models. Keeping more of that data close to the NPU can reduce latency and power consumption. Qualcomm says the architecture is designed for long-context reasoning, multimodal models, concurrent agents and low-latency action loops.

For INT4 quantized models, the company claims up to a 50% improvement in pre-fill performance versus its previous-generation Snapdragon 8 Elite Gen 5. Qualcomm is also targeting Mixture-of-Experts, or MoE, models as large as 30 billion parameters. In its example, the company points to a 30B model that activates only about 3 billion parameters during a given token-generation step. That selective approach is a natural fit for the tight power and memory constraints of a handset.

Qualcomm isn’t suggesting that a smartphone will simply run a 30B dense model entirely from DRAM, however. By activating only the experts needed for a particular task, MoE architectures can dramatically reduce active compute and memory bandwidth requirements, potentially making much larger classes of AI models practical on mobile hardware.

There is a larger competitive story here as well. For years, flagship smartphone competition largely came down to CPU, GPU, modem performance and power efficiency. AI is now another major area of differentiation, and Qualcomm is clearly designing its next-gen Snapdragon mobile platform around this need.

Apple has the advantage of controlling its silicon, operating system and software stack end-to-end. MediaTek continues to push aggressively into premium Android devices as well. Qualcomm’s response is to make its custom Oryon CPU, Adreno GPU and Hexagon NPU work more like a fully integrated compute platform, tuned for modern AI and agentic workloads.

This is a solid advantage for Android handset OEMs because many don’t have the resources to engineer this level of silicon integration themselves. Qualcomm can effectively provide its Android OEM partners, including Samsung, Xiaomi, Honor and others, with a common hardware foundation for on-device AI, advanced graphics and agentic computing. It also gives Qualcomm another way to defend its premium Snapdragon position.

The company’s custom Oryon architecture now spans smartphones and Windows PCs, and soon servers, while AI acceleration is distributed across its CPU, GPU and NPU. For business users, that could mean more AI processing happens locally, reducing cloud and network dependence, improving response times and keeping more potentially sensitive data on-device. And for IT organizations, Qualcomm now has an opportunity to provide a more consistent AI compute foundation across Windows PCs with Snapdragon X and premium Android handsets.

There is still plenty we don’t know, however. Qualcomm hasn’t disclosed the complete SoC specifications, final performance or power characteristics, or which devices will ship with it initially. Many of these AI experiences will also depend heavily on Android, app developers and the models themselves.

The new 5GHz mobile CPU may generate the biggest headline at Snapdragon Summit, but the more important developments may be what sit around it: more local memory, specialized AI acceleration and tighter integration between the CPU, GPU and NPU.

If Qualcomm can translate this architecture into better sustained performance, enhanced power efficiency and useful on-device AI experiences, it could strengthen its position in premium Android phones while putting more competitive pressure on Apple. For business users, the payoff could be more capable local AI without sacrificing the performance and battery life expected of a flagship device.

That is the part I’ll be watching most closely when the complete silicon picture is unveiled at Snapdragon Summit later this month. And I may even get some hands-on time with it in some benchmarks, as we have in years past, so we shall see.
