{"slug": "qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared", "title": "Qualcomm Files Patent for Storing Giant AI Models as Small Edits to a Shared Copy", "summary": "Qualcomm has filed a patent for storing Mixture of Experts AI models as one shared base plus small per-expert delta maps, converting models after training without retraining, according to the filing. In the patent's example, a stored value of 4 becomes a shared 7 plus a difference of minus 3, and a configuration with three bases and 2-bit differences scored best at about 0.41 mean absolute error. Qualcomm pitches the approach for phones, laptops and servers to cut memory traffic, at the cost of extra arithmetic to rebuild weights and small errors from clamping.", "body_md": "# Qualcomm Files Patent for Storing Giant AI Models as Small Edits to a Shared Copy\n\n[Get the best of each week in your inbox, free →](#get-weekly)\n\nGiant AI models keep many near-identical expert copies sitting in memory. Qualcomm's patent keeps one shared copy and stores each expert as a small set of differences.\n\n## What Qualcomm's shared-copy AI patent does\n\nEver kept ten nearly identical copies of the same document, each with a few edits, and wondered why you didn't just save one and note the changes? Many large AI models have a similar habit. They are built from groups of *expert* sub-models that share the same shape and differ only in their stored numbers, and every copy takes up room in memory.\n\n[Qualcomm](https://patentlyze.com/qualcomm/)'s new patent describes **keeping one shared set of numbers** and, for each expert, storing only the small differences. In the filing's own example, a stored value of 4 becomes a shared 7 plus a difference of minus 3. Small differences take fewer bits to store, so there is less data to shuffle around.\n\nThe conversion happens after the model is trained, so the model would not need retraining. The filing says the pieces get added back together at the processor right before the math happens.\n\n## How the shared base and delta maps work\n\nThe patent targets **Mixture of Experts** models, AI systems that split their work across many specialist sub-networks and send each piece of input (called a token) to only a few of them. Every expert has the same structure. Only their **weights**, the learned numbers inside, differ.\n\nThe claimed core has three steps:\n\n- Pick a **shared base** : one set of base values worked out from the experts' weights.\n- For each expert, subtract the base from its weights to get a **delta map** , a grid of small differences.\n- Combine base and deltas to produce each expert's **base-delta representation** .\n\nThe filing also describes optional extras. Experts can be sorted into groups with **k-means clustering** (a method that piles similar items together) so each group shares a base. Oversized differences can be clamped to a cap, such as negative 2 to 2. The system can also list candidate setups, each defined by how many bases to keep and how many bits each difference gets, then pick the one with the lowest **mean absolute error** (the average size of the mistakes introduced). In the example, a setup with three bases and 2-bit differences scored best, at about 0.41.\n\nThe tradeoff is some extra arithmetic to rebuild the weights in exchange for less memory traffic.\n\n## Why smaller experts could help phone AI\n\nMemory is a major bottleneck for running big AI models. The filing notes that expert models can outgrow a device's main memory, and that jumping between experts strains the data pipe between memory and processor. Storing experts as small differences means less data to move, which is why the filing pitches the idea for phones, laptops and servers alike.\n\nFor you, the payoff could be larger AI features running on more modest hardware, if the idea works as described. The filing says no retraining is needed, which would make it easier to apply to models that already exist. The costs are extra work to rebuild weights on the fly and small errors from clamping, so accuracy results on real models would decide whether it holds up.\n\nQualcomm's 57th filing we've tracked in our [AI chip wars watchlist](https://patentlyze.com/watchlist/the-ai-chip-wars/) since July adds to a pattern that includes [one on phones reporting AI limits](https://patentlyze.com/patent/qualcomm-phones-report-their-ai-beam-limits/) and [one on partial model updates](https://patentlyze.com/patent/qualcomm-updating-ai-models-without-reinstalling-them/).\n\nAmong the ideas in this filing, this one sits close to the shippable end. It is a software-style recipe: take a finished model, run a one-time conversion, and add the pieces back together at the processor. The filing lists ordinary chips (CPUs, GPUs and dedicated AI processors), and no new hardware is described.\n\nWhat has to exist first is a model built from many experts, plus proof that squeezing it does not hurt answer quality. The only worked example here is a toy table of eight experts with 32 numbers each, and the error figures come from that toy. Nothing in the text shows results on a real chatbot.\n\nThe shortest route to a product is a conversion step in the software that prepares a trained model for a device. That is a small step compared with designing a new chip, which makes the idea practical to try.\n\n### There are more where this came from\n\nWe read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.\n\n## The drawings\n\n17 drawing sheets from US 2026/0311068 A1 · click any drawing to enlarge\n\n    Want this weekly breakdown for a company we don't cover?\n    [Patentlyze Pro →](https://patentlyze.com/pro/?src=post)\n\n**Source.** Full patent text and figures from the\n\n[official USPTO publication PDF](https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/20260311068).", "url": "https://wpnews.pro/news/qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared", "canonical_source": "https://patentlyze.com/patent/qualcomm-shrinking-big-ai-models-shared-data/", "published_at": "2026-10-09 04:22:04+00:00", "updated_at": "2026-10-09 04:47:55.148582+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "machine-learning", "large-language-models", "ai-research"], "entities": ["Qualcomm", "Mixture of Experts", "k-means clustering", "mean absolute error", "CPU", "GPU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared", "markdown": "https://wpnews.pro/news/qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared.md", "text": "https://wpnews.pro/news/qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared.txt", "jsonld": "https://wpnews.pro/news/qualcomm-files-patent-for-storing-giant-ai-models-as-small-edits-to-a-shared.jsonld"}}