On September 8, Qualcomm and Amazon signed a deal worth up to $60 billion over 10 years to co-develop custom AI inference chips and 1.6 Tbps optical interconnects for AWS data centers. Qualcomm stock jumped 9%. The financial press called it a Nvidia challenger. The actual story is more useful: if you run ML workloads on AWS, the economics of your inference stack are starting to shift — and a key piece of infrastructure most analysts ignored is at the center of it.
What Was Agreed #
The deal covers two distinct deliverables. First, custom AI inference silicon — purpose-built accelerators designed specifically for AWS data centers, with multiple chip generations committed. Second, optical interconnects reaching up to 1.6 Tbps, using technology Qualcomm acquired when it bought Alphawave Semi for $2.4 billion in December 2025.
The financial structure is deliberately conditional. Amazon received warrants for up to 25 million Qualcomm shares that vest only if Amazon’s cumulative purchases hit specific milestones, potentially reaching $60 billion through 2036. It is not a guaranteed $60 billion check — it is a decade-long commitment that pays out proportionally to how much Amazon actually buys. Qualcomm already has revenue booked starting in the December 2026 quarter, though that is likely existing networking products, not the next-generation AI300 accelerator.
The Inference Economics Case #
Qualcomm’s Dragonfly AI300, announced at its June 2026 Investor Day, claims 4x to 8x better performance-per-watt compared to GPU-based architectures for inference workloads. The AI300 uses High Bandwidth Compute Gen 2 and delivers a 54x increase in effective memory bandwidth compared to its predecessor. Commercial sampling starts in 2028.
Those numbers matter because inference is where production AI costs live. AWS’s current Inferentia2 already demonstrates what custom inference silicon can do: an inf2.xlarge runs at $0.758/hour, compared to $1.006/hour for a g5.xlarge GPU instance, with up to 40% better cost-per-inference on compatible workloads. The catch is the AWS Neuron SDK, which requires model porting and limits compatibility. If Qualcomm’s chips arrive with broader model support and deliver even half of the claimed perf/watt improvement, the cost curves for decode-heavy serving workloads on AWS shift materially.
That is not a 2026 decision. The AI300 does not sample until 2028. But if you are making multi-year AWS architecture commitments today — reserved instances, long-term contracts, infrastructure lock-in — the 2028 landscape is worth factoring in.
The Interconnect Nobody Talked About #
Most coverage focused on the chip deal. The 1.6 Tbps optical interconnect component deserves more attention.
As AI inference clusters scale from a single rack to hundreds of racks spanning multiple data center zones, the bottleneck shifts. Compute throughput matters less when the accelerators cannot move data between themselves fast enough. Qualcomm’s Alphawave acquisition brought exactly this capability: SerDes and optical DSP technology optimized for 800G to 1.6T connectivity. At 1.6 Tbps per port, the interconnect fabric can keep pace with next-generation accelerator throughput — a constraint that existing 400G optical infrastructure is already straining to meet at scale.
This is not a theoretical problem. It is the reason Nvidia built NVLink and InfiniBand into its data center strategy. Qualcomm is positioning itself to own both the inference compute and the glue between it — a systems-level play, not just a chip play.
Two Hyperscalers Are Now Betting on Qualcomm #
The AWS deal did not come out of nowhere. At Qualcomm’s June 2026 Investor Day, Meta signed a multi-generation agreement to deploy the Dragonfly C1000 CPU — a 250-core server processor using custom Oryon cores, sustained frequencies above 5GHz, and PCIe Gen 7 — in its server infrastructure. The C1000 ships in the second half of 2028.
Meta deploying C1000 for agentic AI orchestration workloads, Amazon deploying custom inference silicon — two of the largest AI infrastructure operators in the world endorsing the same silicon vendor within three months. That is a meaningful signal about where hyperscaler confidence is going, even if the products themselves are still two years out.
The Nvidia Framing Is Wrong #
Headlines called this deal a shot at Nvidia’s dominance. That framing is an oversimplification. Nvidia is still supplying one million GPUs to AWS through 2027. AWS runs Inferentia2, Trainium2, Cerebras, Groq inference chips, and now Qualcomm custom silicon — simultaneously. This is not a replacement strategy; it is a layered inference ecosystem where AWS picks the right silicon for the right workload and uses multi-vendor supply to negotiate better pricing from all of them.
The real implication for developers is that AWS is accelerating its push toward specialized inference silicon away from general-purpose GPUs. If you are using p4d or p5 instances for production inference today because GPU flexibility is worth the cost premium, that calculation will continue to shift as AWS adds inference-optimized alternatives with better price-performance.
What to Watch #
The deal does not require immediate action. The AI300 samples in 2028; the C1000 ships the same year. Watch for three things: whether Qualcomm announces a developer SDK or compatibility layer that avoids the porting friction of AWS Neuron; whether the AI300 delivers independently benchmarked perf/watt results close to its claims; and whether AWS prices the new instances aggressively to push adoption, as it did when Inferentia2 launched.
AWS is building a multi-silicon inference ecosystem. Nvidia still owns training. But inference — the compute that actually serves your users, runs your agents, and shows up on your monthly AWS bill — is becoming a competitive market. For the first time in years, that is good news for anyone paying for it.