{"slug": "which-apple-ai-workloads-leave-the-device-and-what-they-cost", "title": "Which Apple AI Workloads Leave the Device and What They Cost", "summary": "Apple's Apple Intelligence is a hybrid system that runs privacy-sensitive tasks on-device but routes heavier workloads to Private Cloud Compute (PCC), with offloads including agentic tool use, complex reasoning on AFM 3 Cloud Pro, and image editing on ADM 3 Cloud. The routing boundary is opaque, as Apple does not publish the thresholds that trigger offloads, and each PCC request is a transmission event that raises data-exposure and cost concerns. The on-device AFM 3 Core Advanced model requires 12GB of unified memory, which excludes the base iPhone 17 with 8GB from the most capable features.", "body_md": "[Apple Intelligence](https://en.wikipedia.org/wiki/Apple_Intelligence) is marketed as on-device AI. In practice it is a hybrid: privacy- and latency-sensitive requests run locally on Apple silicon, and anything heavier routes to [Private Cloud Compute](https://security.apple.com/blog/private-cloud-compute/) (PCC) without asking you first. Where that boundary sits, and what crossing it costs, is opaque.\n\nBy the end you should read every offload as a data-exposure and cost decision. You’ll see what leaves the device, how Apple routes inference, how to assess sovereignty, and how to weigh [iCloud+](https://en.wikipedia.org/wiki/ICloud) costs. It’s one part of the broader [OS 27 routing and cost picture](/ios-27-apples-ai-native-operating-system-play).\n\n## Private Cloud Compute vs on-device AI processing: which workloads leave the device and why?\n\nMost Apple Intelligence requests [stay on-device](https://security.apple.com/blog/expanding-pcc/), but work that exceeds local capability routes to Private Cloud Compute. The main offloads are agentic tool use and complex reasoning on AFM 3 Cloud Pro, heavier general inference on AFM 3 Cloud, and image editing on ADM 3 Cloud, while simpler tasks stay on AFM 3 Core and [AFM 3 Core Advanced](https://machinelearning.apple.com/research/introducing-apple-foundation-models) locally.\n\nApple introduced Private Cloud Compute to extend device-grade security to workloads on-device models can’t handle. How that protection works is covered [here](/private-cloud-compute-and-the-trust-infrastructure-behind-apples-os-27).\n\nAnother way to think about it is four drivers: model size, memory pressure, latency tolerance, and capability. Local inference transmits nothing. Every PCC request is a transmission event, even though Apple says data is [never stored or shared with anyone, including Apple](https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models). That zero-transmission versus transmission distinction applies in every market.\n\n## How does Apple decide where inference runs between on-device and cloud Foundation Models?\n\nApple splits inference across five Foundation Models: two on-device (AFM 3 Core and the 20-billion-parameter AFM 3 Core Advanced) and three server-side (AFM 3 Cloud, ADM 3 Cloud, and AFM 3 Cloud Pro). On-device handles privacy-sensitive or modest tasks; cloud handles heavier inference. The strongest local model needs [12GB of unified memory](https://www.macrumors.com/2026/06/08/most-powerful-on-device-ai-now-requires-iphone-17-pro-or-air/), so the hardware, alongside the request, decides much of the routing.\n\nAFM 3 Core Advanced is the key model. It is a sparse model that activates only 1 to 4 billion parameters per request, with weights in NAND flash and DRAM as a working buffer. That [exotic architecture](https://venturebeat.com/technology/on-device-ai-agents-hit-a-hard-memory-limit-apples-new-architecture-routes-around-it) is what lets it run on a phone at all.\n\nAFM 3 Core Advanced also decides which devices get the most capable on-device features: the base iPhone 17 misses out because it ships with 8GB, below that floor. On lower-RAM hardware, devices fall back to a smaller local model or reach PCC sooner.\n\nApple does not publish the RAM, model-availability, or OS-version thresholds that trigger an offload, so you can see the memory line but not the routing rule. For the 7-billion-parameter Siri model, see [this piece](/siri-ai-and-the-ai-native-operating-system-in-ios-27-and-macos-27). Because the routing rule is unpublished, assessing compliance becomes an evidence problem.\n\n## How should we assess whether Private Cloud Compute meets data sovereignty, residency and compliance requirements?\n\nAssess PCC against four criteria: where processing occurs, what is retained, what residency guarantees exist, and whether routing is verifiable. PCC is stateless and attested, but Apple does not publish the routing heuristic, and the EU and China change availability and the compliance answer. Treat this as an evidence-based assessment.\n\nThose questions are hard to answer from Apple’s public material. PCC’s five guarantees (stateless computation, enforceable properties, no privileged runtime access, non-targetability, verifiable transparency) are privacy engineering, [an architecture enterprises have been told to adopt for years](https://www.fortanix.com/blog/apple-just-made-the-case-for-confidential-ai-for-all-of-us). But they are privacy properties; a legal residency guarantee is a separate requirement. Stateless means data is not kept; it does not say which jurisdiction the silicon sat in.\n\nThe EU makes it concrete. [Siri AI](https://www.apple.com/apple-intelligence/) is unavailable on iOS, iPadOS, and watchOS there, and Apple says [there is no current timeline for Siri AI in the EU](https://www.pcmag.com/news/eu-apple-refused-to-follow-rules-meant-to-keep-siri-ai-in-check-wwdc-2026). China is the second boundary, with [Apple Intelligence excluded while regulatory questions stay unresolved](https://www.apple.com/newsroom/2026/06/apple-intelligence-brings-powerful-ai-capabilities-into-everyday-experiences/). In both cases the answer is decided by distribution rules.\n\nThe lesson for your team: privacy engineering buys you a lot, but it does not fill in the residency and routing evidence a sign-off needs. See how PCC protects cloud inference here, and where this sits among the five OS 27 decision areas. Once you’ve assessed the exposure, the next question is what it costs.\n\n## How should we weigh iCloud+ costs for heavy AI use against on-device limits and server-side fallback?\n\nHeavy AI use is not free. On-device inference avoids cloud token costs but is capped by the 12GB unified-memory threshold and weaker local models. Server-dependent features that route to PCC sit behind iCloud+ tiers and daily usage limits, and Apple has not published per-request pricing. Weigh capability, cost, and privacy together, then convert the trade-off into an allow/block decision.\n\nOn-device inference is the cost-avoidance path. Organisations with IP-sensitive workloads use it to avoid cloud token costs. Tim Cook cites Disney’s creative teams doing this on Mac to [reduce cloud token costs and keep IP secure](https://www.cnbc.com/2026/07/30/tim-cook-sees-apples-hybrid-ai-strategy-as-a-competitive-weapon-.html), and has flagged upgrade options on iCloud+ for heavy users.\n\nThe server side is where Apple’s monetisation shows up. Features running on powerful server models, image generation, for example, carry daily usage limits, and [iCloud+ subscribers get higher allowances](https://www.macrumors.com/2026/06/11/icloud-plus-apple-ai-subscription-model/). What is missing is the number: [Apple has not announced AI-specific prices or numerical caps](https://mlq.ai/news/apple-eyes-icloud-upgrades-for-heavier-ai-use-but-pricing-remains-unset/). Treat the daily caps as a soft cost signal: heavy use is throttled by daily caps, [Apple does not bill per request](https://blakecrosley.com/blog/foundation-models-private-cloud-compute), and the ceiling stays opaque.\n\nThe practical framing is a triangle of capability, cost, and privacy. On-device is private and free but limited; PCC is more capable but carries usage limits and data exposure. There is no switch to force everything on-device, so the lever is policy, covered [here](/apple-intelligence-third-party-models-and-enterprise-ai-governance).\n\nSo where does that leave you? Apple’s on-device promise is genuine but bounded: the 12GB memory threshold and unpublished routing rules push work to PCC. PCC’s privacy engineering is real, and a legal residency guarantee is a separate requirement, so compliance has to be assessed against four criteria. Every offload is a cost event Apple monetises through iCloud+. The useful next step, within the broader OS 27 picture, is turning that triangle into an allow/block decision, where each offload is a deliberate decision to accept data exposure and cost.\n\n## Frequently Asked Questions\n\n### Can I turn off Private Cloud Compute and force everything on-device?\n\nNo, there is no user-facing switch that forces all Apple Intelligence work to stay on-device. Routing between local models and Private Cloud Compute is automatic and gated by Apple’s unpublished hardware thresholds, so heavier requests offload without asking. Your practical lever is policy: restrict server-dependent features through device management, or avoid the features that trigger cloud fallback in the first place.\n\n### Is it true that Apple Intelligence never sends my data to Apple’s servers?\n\nNo. That claim only holds for the on-device path. Apple Intelligence is a hybrid, and workloads that exceed local capability are sent to Private Cloud Compute. The difference is what happens next: Apple says PCC processing is stateless and does not retain your data, but the request still leaves the device. Treat every offload as a transmission event, even when no retention occurs.\n\n### How do I know if my iPhone can run the most capable on-device model?\n\nThe strongest on-device model, AFM 3 Core Advanced, requires 12GB of unified memory, so devices below that floor fall back to a smaller local model or reach Private Cloud Compute more often. Apple does not publish a simple per-device table of offload thresholds. Check your device’s unified memory as the practical line, then assume heavier features will hit the cloud sooner on lower-RAM hardware.\n\n### What happens if I exceed the daily usage limit on server-dependent features?\n\nApple applies daily usage limits to some server-dependent features such as Image Playground. When you hit a limit, the feature stops serving until the quota resets rather than silently routing elsewhere. Apple does not publish the exact quotas or a per-request price, so treat these caps as a soft cost signal: heavy use is throttled, not billed, and the ceiling stays opaque.\n\n### Does Private Cloud Compute store my data after it answers?\n\nNo, not in Apple’s stated design. PCC is described as stateless: your request is processed, the result is returned, and the data is not retained. However, stateless processing is a privacy property, not a legal residency guarantee. Compliance teams still need evidence of where processing occurred and how routing was decided before they can sign off.\n\n### Can Apple read what I send to Private Cloud Compute?\n\nApple’s architecture is designed so that even privileged operators cannot target your request, and PCC nodes run without the usual remote-access surfaces. In practice, no human is reading your prompts as a matter of course. The caveat is verifiability: because Apple does not publish the routing heuristic, you cannot independently confirm which requests left the device in the first place.\n\n### Is Apple Intelligence available in Australia, and how do the EU and China limits affect it?\n\nAvailability follows Apple’s regional rollout rather than the routing rules. The bigger signal is distribution: Siri AI was initially unavailable on iOS, iPadOS and watchOS in the European Union, and Apple Intelligence is excluded in China while regulatory questions stay unresolved. In Australia the features generally roll out with the global release, but the compliance picture differs by market, not by device capability alone.\n\n### Do I have to pay extra for Apple Intelligence, or is it free?\n\nBasic Apple Intelligence features are included with the device, and on-device inference avoids cloud token costs. The catch arrives with heavy use: server-dependent features sit behind iCloud+ tiers and daily usage limits, and Apple has not published per-request pricing. In practice, light local use is effectively free, while heavy cloud use is monetised through iCloud+.\n\n### Does on-device AI drain my battery faster than cloud processing?\n\nOn-device inference runs locally on Apple silicon, which avoids cloud token costs, but it draws memory and power on the device itself. Heavier local work can increase battery and memory pressure, while offloading to Private Cloud Compute shifts the load to Apple’s servers. Apple does not publish power figures for either path, so the practical trade is capability and privacy, not a measured battery number.\n\n### What does “stateless” mean for my data?\n\nStateless means PCC holds your request only long enough to compute an answer, then discards it, with no persistent storage tied to you. That limits what Apple could hand over or leak later, and it is a genuine privacy property. It does not mean the data never left your device, and it does not by itself satisfy a residency or retention audit, which needs documented evidence of location and routing.\n\n### Can I see a log of which requests went to Private Cloud Compute?\n\nNot easily. Apple does not expose a user-facing ledger of which prompts ran on-device and which offloaded to Private Cloud Compute, and the routing heuristic itself is unpublished. That opacity is the core problem for regulated teams: you cannot document what you cannot see. Treat any request that might exceed local capability as a potential offload until Apple offers verifiable routing.\n\n### Is on-device AI always more private than Private Cloud Compute?\n\nIn one narrow sense, yes: on-device processing transmits nothing, so no data leaves the device. But privacy is not a single switch. A weaker local model may handle sensitive content differently, and PCC’s stateless, attested design is engineered to minimise exposure. The honest comparison is a triangle of capability, cost and privacy, not a simple ranking where local always wins.", "url": "https://wpnews.pro/news/which-apple-ai-workloads-leave-the-device-and-what-they-cost", "canonical_source": "https://www.softwareseni.com/which-apple-ai-workloads-leave-the-device-and-what-they-cost/", "published_at": "2026-08-26 16:00:00+00:00", "updated_at": "2026-08-27 02:49:48.838828+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-infrastructure", "ai-policy"], "entities": ["Apple", "Apple Intelligence", "Private Cloud Compute", "AFM 3 Core Advanced", "AFM 3 Cloud Pro", "ADM 3 Cloud", "iPhone 17"], "alternates": {"html": "https://wpnews.pro/news/which-apple-ai-workloads-leave-the-device-and-what-they-cost", "markdown": "https://wpnews.pro/news/which-apple-ai-workloads-leave-the-device-and-what-they-cost.md", "text": "https://wpnews.pro/news/which-apple-ai-workloads-leave-the-device-and-what-they-cost.txt", "jsonld": "https://wpnews.pro/news/which-apple-ai-workloads-leave-the-device-and-what-they-cost.jsonld"}}