{"slug": "machines-that-think-part-2-from-cloud-to-edge", "title": "Machines That Think, Part 2: From Cloud to Edge", "summary": "Edge AI inference is arriving on devices but barely denting data-center demand, with Deloitte's 2026 predictions showing inference will account for roughly two-thirds of all AI compute this year, up from half in 2025, and almost all of it still running in large data centers or enterprise on-premises systems. Counterpoint expects GenAI-capable smartphones to reach 45% of global shipments in 2026, up from 36% in 2025, but memory costs keep such devices above roughly $400 wholesale, limiting expansion beyond the high-end segment.", "body_md": "*This is the second post in a 14-part series on how AI is crossing out of software and into the physical world. Each post takes one technology or shift and asks the same three questions: what does it unlock for startups, what does it force on incumbents, and what does it mean for society? Last time we started at the bottom of the stack, with the compute substrate. This time we ask what happens when inference moves off the data center and onto the device in your hand.*\n\nThere is a comforting story about the edge, and it is roughly the opposite of the one I took apart in Part 1. In this story, the data center bottleneck solves itself. Models get smaller, chips in phones and laptops get better, inference quietly decentralizes onto a few billion devices, and the allocation politics of packaging capacity stop mattering quite so much. The pressure escapes through the edges.\n\nI want to believe this story. It is the optimistic one, and parts of it are true. But the numbers for this year say something more awkward: the edge is not draining the data center, and the reason it isn't turns out to be the same reason the data center is constrained in the first place.\n\nStart with what has actually arrived, because it is not nothing. The capability is real and it is shipping. Counterpoint expects [ GenAI-capable smartphones to reach 45% of global shipments in 2026](https://counterpointresearch.com/en/insights/genai-smartphone-share-to-rise-to-45-percent-of-global-shipments-in-2026?ref=janbosch.com), up from 36% the year before, and\n\n[, with neural processing units clearing Microsoft's 40 TOPS bar as a matter of course. The models have come down to meet the hardware. Apple's](https://counterpointresearch.com/en/reports/ai-advanced-pcs-to-surpass-half-of-global-shipments-in-2026?ref=janbosch.com)\n\n__AI-advanced PCs to pass 59% of global shipments__[is three billion parameters, with a larger sparse variant that activates only one to four billion at a time depending on the request. Could you have imagined this five years ago? It is now the default tech stack on a consumer phone.](https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models?ref=janbosch.com)\n\n__third-generation on-device foundation model__And yet. Deloitte's [ 2026 predictions](https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html?ref=janbosch.com) expect inference to account for roughly two-thirds of all AI compute this year, up from half in 2025. And expect almost all of it to still run in large data centers or enterprise on-premises systems. Edge inference remains a small fraction of total demand. Hundreds of millions of NPUs are shipping into devices, and they are barely moving the aggregate needle.\n\nSo we have a paradox: the edge is arriving everywhere and mattering very little. The usual explanation is that the on-device models simply are not good enough yet, and that time will fix it. I think the usual explanation is wrong, or at least badly incomplete.\n\nHere is what I believe is actually going on. The constraint at the edge is not compute. It is memory; the same constraint as in Part 1, wearing different clothes. A model has to fit in RAM to run. Not \"fit on the device,\" not \"fit in flash,\" but sit in memory alongside the operating system and whatever else you have open while working. That is a hard ceiling, and it is set by a bill of materials, not by a benchmark. Counterpoint's analysts are explicit that memory \"will remain a key factor determining how quickly GenAI expands beyond the high-end segment,\" and that the extra DRAM needed to hold model weights is what keeps GenAI devices above roughly $400 wholesale. You can put a 60 TOPS NPU in a mid-range phone. You cannot put an extra eight gigabytes of LPDDR in it and still hit the price point.\n\nWhich brings us to the part I find genuinely uncomfortable. That DRAM and the HBM stacked next to a data-center accelerator are made in the same fabs, on the same wafers, by the same handful of vendors. IDC put it about as plainly as an analyst house ever does: [ \"every wafer allocated to an HBM stack for an Nvidia GPU is a wafer denied to\"](https://www.idc.com/resource-center/blog/global-memory-shortage-crisis-market-analysis-and-the-potential-impact-on-the-smartphone-and-pc-markets-in-2026/?ref=janbosch.com) a consumer device. Memory is 15–20% of the bill of materials on a mid-range phone. IDC's scenarios for 2026 run from a 2.9% smartphone market contraction with a 3–5% rise in average selling price to a 5.2% contraction with a 6–8% rise, and worse for PCs.\n\nSo the edge is not an escape hatch from the substrate. It is a competitor for it. The move from cloud to edge does not route around the bottleneck described in Part 1; it queues up behind the same one. That is the single most important thing to understand about this shift, and it is almost entirely absent from how the edge gets discussed.\n\nFor a startup, this reframes a question that badly needed reframing. \"Cloud or edge\" is not an architectural preference to be settled by taste or by cost per token. It is a question about which parts of your workload can live inside a memory budget you do not control and cannot expand, on a device your customer already owns.\n\nThe trap is to treat the edge as the cheap version of the cloud. If your reason for going on-device is that inference costs too much, you have chosen a strategy that competes for scarce memory in order to escape scarce packaging, and you have added a distribution problem on top. The far better reason is that some workloads are not merely expensive in the cloud but impossible there. Anything where the round trip is the product: low latency, operation without connectivity, data that must never leave the device for legal or commercial reasons. That is a genuine edge workload. Those are defensible. \"The same thing, but we pay less for it\" is not.\n\nThe subtler consequence, and one I keep coming back to, is what running on the edge does to your ability to know whether the thing works. When inference happens in your data center you see everything: every input distribution shift, every failure mode, every regression. When it happens on ten million devices you own none of that by default. You have traded an allocation problem for an observability problem. The durable artifact in this world is the contract the model must honor plus the evidence that it still does, and the edge makes gathering that evidence structurally harder. Any team going on-device should be designing its evaluation telemetry at the same time as its quantization strategy, not two years later.\n\nFor large incumbents the calculus, as usual, inverts. Cristiano Amon has been making the case for years and made it about as bluntly as possible at Davos this January: [ \"Whoever has presence on the edge is going to win. The edge is where the humans are.\"](https://time.com/collections/davos-2026/7339220/qualcomm-ceo-cristiano-amon-ai-edge-computing/?ref=janbosch.com) He sells the silicon that goes in the edge, so he would say that. But the strategic logic holds independently of who is making the argument. If the agent becomes the primary interface, then whoever owns the device owns the default. Defaults, in consumer technology, have historically been worth more than features.\n\nThis is why the companies that control both the silicon and the operating system are in an enviable position, and why everyone else in the device business is in an increasingly awkward one. The awkwardness is again about memory. A vendor with the margin structure to swallow a rising memory bill will ship the AI features and eat the cost. A vendor without it has to choose between raising the price and cutting the RAM, and cutting the RAM means shipping a phone that has the NPU and not the memory to use it. Which is to say a phone that has the marketing and not the capability. The mid-tier squeeze I described in Part 1 for compute buyers has an exact analogue one layer down the stack, and it will thin out the middle of the device market over the next two years.\n\nFor society, the on-device turn is the most encouraging development in this series so far, and it comes with a bill attached. The encouraging part is architectural rather than promissory. Data that never leaves the device cannot be breached at the provider, subpoenaed from the provider, or quietly folded into the next training run. That is a stronger privacy guarantee than any policy, terms-of-service clause or certification, because it is a property of where the computation happens rather than of what someone promises to do with it afterwards. For European enterprises in particular, a meaningful slice of the data sovereignty problem simply dissolves when the inference is local. This is the rare case where the commercially attractive option and the civically desirable one point the same way.\n\nThe bill is that AI capability is becoming a function of what hardware you can afford, at exactly the moment hardware is getting more expensive for reasons that have nothing to do with the buyer. The memory shortage is a direct consequence of data-center demand, and it lands on the price of a phone in a market that has nothing to do with frontier models. We are, in effect, taxing the mid-range device to build the training cluster. If GenAI capability requires a device above a certain price, and that price is being pushed up by AI's own appetite for memory, then the technology is drawing a line through the population and putting the people who would benefit most on the wrong side of it. That is a distributional question, and it is not going to be solved by better quantization.\n\nOne last point, and it is the reason this post sits where it does in the series. Everything from here on, such as robots, factory cells, vehicles, instruments in a lab, is edge inference whether anyone calls it that or not. A robot arm cannot round-trip to a data center to decide whether to stop. A vehicle cannot make its control decisions while relying on a network. For the physical systems that fill the rest of this series, local inference is not an optimization; it is a precondition. Everything I have described here as a constraint on phones and laptops becomes a design constraint on machines that move.\n\nThe models get the headlines. But intelligence only becomes infrastructure when it stops being somewhere you go and starts being something that is simply present, at the point where the work happens. Mark Weiser saw the shape of this in *Scientific American* in 1991, long before there was anything to run: \"The most profound technologies are those that disappear. They weave themselves into the fabric of everyday life until they are indistinguishable from it.\" That is the destination. The memory bill is what we pay to get there.\n\n*Next in the series:* **Embodied Intelligence***— why teaching a machine to act in the physical world is a different problem from teaching one to talk.*\n\n*Want to read more like this? Sign up for my newsletter at jan@janbosch.com or follow me on janbosch.com/blog, LinkedIn (linkedin.com/in/janbosch) or X (@JanBosch).*", "url": "https://wpnews.pro/news/machines-that-think-part-2-from-cloud-to-edge", "canonical_source": "https://janbosch.com/machines-that-think-part-2-from-cloud-to-edge/", "published_at": "2026-08-31 14:21:01+00:00", "updated_at": "2026-08-31 14:26:51.235259+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-products"], "entities": ["Deloitte", "Counterpoint", "Apple", "Microsoft"], "alternates": {"html": "https://wpnews.pro/news/machines-that-think-part-2-from-cloud-to-edge", "markdown": "https://wpnews.pro/news/machines-that-think-part-2-from-cloud-to-edge.md", "text": "https://wpnews.pro/news/machines-that-think-part-2-from-cloud-to-edge.txt", "jsonld": "https://wpnews.pro/news/machines-that-think-part-2-from-cloud-to-edge.jsonld"}}