Get the best of each week in your inbox, free → Giant AI models keep many near-identical expert copies sitting in memory. Qualcomm's patent keeps one shared copy and stores each expert as a small set of differences.
What Qualcomm's shared-copy AI patent does #
Ever kept ten nearly identical copies of the same document, each with a few edits, and wondered why you didn't just save one and note the changes? Many large AI models have a similar habit. They are built from groups of expert sub-models that share the same shape and differ only in their stored numbers, and every copy takes up room in memory.
Qualcomm's new patent describes keeping one shared set of numbers and, for each expert, storing only the small differences. In the filing's own example, a stored value of 4 becomes a shared 7 plus a difference of minus 3. Small differences take fewer bits to store, so there is less data to shuffle around.
The conversion happens after the model is trained, so the model would not need retraining. The filing says the pieces get added back together at the processor right before the math happens.
How the shared base and delta maps work #
The patent targets Mixture of Experts models, AI systems that split their work across many specialist sub-networks and send each piece of input (called a token) to only a few of them. Every expert has the same structure. Only their weights, the learned numbers inside, differ.
The claimed core has three steps:
- Pick a shared base : one set of base values worked out from the experts' weights.
- For each expert, subtract the base from its weights to get a delta map , a grid of small differences.
- Combine base and deltas to produce each expert's base-delta representation .
The filing also describes optional extras. Experts can be sorted into groups with k-means clustering (a method that piles similar items together) so each group shares a base. Oversized differences can be clamped to a cap, such as negative 2 to 2. The system can also list candidate setups, each defined by how many bases to keep and how many bits each difference gets, then pick the one with the lowest mean absolute error (the average size of the mistakes introduced). In the example, a setup with three bases and 2-bit differences scored best, at about 0.41.
The tradeoff is some extra arithmetic to rebuild the weights in exchange for less memory traffic.
Why smaller experts could help phone AI #
Memory is a major bottleneck for running big AI models. The filing notes that expert models can outgrow a device's main memory, and that jumping between experts strains the data pipe between memory and processor. Storing experts as small differences means less data to move, which is why the filing pitches the idea for phones, laptops and servers alike.
For you, the payoff could be larger AI features running on more modest hardware, if the idea works as described. The filing says no retraining is needed, which would make it easier to apply to models that already exist. The costs are extra work to rebuild weights on the fly and small errors from clamping, so accuracy results on real models would decide whether it holds up. Qualcomm's 57th filing we've tracked in our AI chip wars watchlist since July adds to a pattern that includes one on phones reporting AI limits and one on partial model updates.
Among the ideas in this filing, this one sits close to the shippable end. It is a software-style recipe: take a finished model, run a one-time conversion, and add the pieces back together at the processor. The filing lists ordinary chips (CPUs, GPUs and dedicated AI processors), and no new hardware is described.
What has to exist first is a model built from many experts, plus proof that squeezing it does not hurt answer quality. The only worked example here is a toy table of eight experts with 32 numbers each, and the error figures come from that toy. Nothing in the text shows results on a real chatbot.
The shortest route to a product is a conversion step in the software that prepares a trained model for a device. That is a small step compared with designing a new chip, which makes the idea practical to try.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings #
17 drawing sheets from US 2026/0311068 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover?
[Patentlyze Pro →](https://patentlyze.com/pro/?src=post)
Source. Full patent text and figures from the
official USPTO publication PDF.