{"slug": "agent-muse-compute-demand", "title": "Agent (Muse) Compute Demand", "summary": "An analysis estimates that serving 100 million daily active users of Meta's Muse agent would require roughly 1–2 GW of average total power in the base case, with only about 0.1 GW coming from the CPU/VM sandbox layer, and 3–4 GW if reasoning-equivalent model calls run higher. The sandbox layer is projected at under $1B of CPU content and about $2B of DRAM, based on roughly 25 million provisioned live VMs at 0.5 physical cores and ~3GB RAM each, benchmarked against DeepSeek's DSec agent sandbox platform.", "body_md": "There are obviously a lot of moving assumptions: how long agents stay active, how aggressively CPUs and memory can be oversubscribed, how many model calls an agent generates, and how efficiently those models are served.\n\nRough conclusion is:\n\n- **~1–2 GW of average total power** to serve 100M DAU in the base case, of which only**~0.1 GW comes from the CPU/VM layer** . Depending on the # of reasoning-equivalent model calls one Muse DAU generates per day,**3-4GW** is entirely plausible.\n- The sandbox layer = **sub $1B****of CPU content** and**~$2B of DRAM content, which is smaller than many expected.**\n\nWelcome all feedbacks/ pushbacks.\n\nTwo very different pieces of infrastructure behind Muse.\n\n1. Muse VM / sandbox infrastructure\n\n**2 vCPUs, ~8 GB of RAM and ~100 GB of persistent logical storage per user**. [Observed Muse VM configuration](https://www.starkinsider.com/2026/09/meta-muse-specs-what-it-runs-on.html)\n\n2. Muse Spark inference\n\nModel inference goes out through Meta's external inference infrastructure. [Meta’s description of Muse architecture](https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse)\n\n# 1/ Sandbox infrastructure\n\n## A. CPU\n\nThe first mistake is assuming that **100M DAU means 100M VMs are actively consuming compute at the same time**.\n\nSuppose the average Muse DAU has an agent actively working for two hours per day.\n\n100M users * 2 hours / 24 hours = ~8M average simultaneous active VMs\n\nMeta obviously cannot provision only for the daily average. Usage will be concentrated during waking hours and bursty.\n\nAssume a 2.5x peak-to-average ratio:\n\n8M * 2.5 = ~20M peak active VMs\n\nThen add roughly 20% capacity headroom: **~25M provisioned live VMs.** So the base assumption is effectively that Meta needs enough infrastructure to support roughly **25% of DAU being live simultaneously**.\n\nThe next important distinction is between **virtual CPU allocation** and **physical CPU demand**. Agent sandboxes are particularly well suited to CPU oversubscription. They spend a lot of time waiting. During those periods, the VM may still be alive, but it is barely using CPU.\n\nDeepSeek’s recently published DSec infrastructure provides a useful benchmark. Its production agent sandbox platform runs approximately **30,000 physical CPU cores and 250TB of DRAM across ~160 nodes**, with peak concurrency above **380,000 sandboxes**. [DeepSeek DSec paper](https://arxiv.org/abs/2609.22978). DSec also demonstrates stable operation at around: **800 microVMs per node.** With roughly 188 physical cores per node: 188 physical cores / 800 microVMs = ~0.23 physical cores per live VM.\n\nDeepSeek is obviously the King of efficiency. The number for Muse might be at 0.3-0.75 physical cores per live VM, or assume **0.5 physical cores per live VM as the base case.** That is equivalent to roughly **two simultaneously live Muse VMs per physical CPU core**.\n\nUsing the base assumptions: **25M live VMs * 0.5 physical cores per VM = 12.5M physical CPU cores.**\n\nOn a 256-core CPU: 12.5M cores / 256 cores per CPU = ~50K CPUs; Or on a 192-core CPU that would be 65K CPUs.\n\n**At the current public pricing, that is ~$800M.**\n\n## B. DRAM\n\nCPU can be aggressively oversubscribed because a VM that is waiting may consume almost no CPU. Memory is harder to oversubscribe because a live VM still needs to retain its working state.\n\nMuse exposes roughly **8GB of RAM** to the user environment, but one observed instance was actually using only around **3GB** at the time of measurement.\n\n25M live VMs * 3GB = 75PB of physical DRAM, call it ~75-100PB of physical DRAM feels like a reasonable base range.\n\n**At the current public pricing, that is ~$2B.**\n\n## **C. Sandbox power**\n\n**~0.1 GW for the entire Muse sandbox / VM layer at 100M DAU.**\n\n# 2/ Inference\n\nMuse’s personal computer executes tools and stores state locally, but the actual model runs on separate inference infrastructure. [Meta’s Muse architecture](https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse)\n\n## Energy per inference event\n\nMicrosoft’s 2026 study estimates that optimized frontier-scale inference consumes a median of approximately: **0.31Wh per normal query**\n\nBut a long reasoning query with roughly 15x the token count consumes approximately 13x as much energy, or around: **4Wh per long reasoning query**\n\nThe study specifically highlights reasoning and agentic workloads as significantly more energy intensive. [Microsoft Research on AI inference energy](https://www.microsoft.com/en-us/research/publication/energy-use-of-ai-inference-efficiency-pathways-and-test-time-scaling/?utm_source=chatgpt.com)\n\nFor a Muse-like workload, let’s assume **5-10Wh per heavy reasoning-equivalent inference event** as a rough sensitivity range.\n\n## Sensitivity analysis on # reasoning-equivalent events per DAU per day\n\nAssume 50 reasoning-equivalent events per DAU per day\n\nSuppose each active Muse user generates the equivalent of 50 heavy inference events per day.\n\nAt 5Wh each:\n\n100M users * 50 events/day * 5Wh = 25GWh/day\n\n25GWh/day / 24 hours = ~1.0GW average power\n\nSo inference alone could require: **~1-2GW of average power**\n\nSo a **3-4GW Muse** is entirely plausible. Maybe that’s why Meta is rumored to be adding 7-10GW of compute next year. \n\n# Bottom line\n\nThe popular framing around Muse is that giving every user **2 vCPUs and 8GB of RAM** creates an enormous CPU requirement. But the naive calculation materially exaggerates the CPU requirement because it treats logical VM allocation as dedicated physical infrastructure.\n\nThe more interesting conclusion is: **Consumer agents may be a meaningful new demand driver for CPUs and conventional DRAM, but inference remains the real compute bottleneck.** And as agents do more work, run longer trajectories and increasingly spawn other agents, **inference demand can scale much faster than the number of users itself.**", "url": "https://wpnews.pro/news/agent-muse-compute-demand", "canonical_source": "https://robonomics.substack.com/p/agent-muse-compute-demand", "published_at": "2026-09-30 15:38:59+00:00", "updated_at": "2026-09-30 15:49:27.843420+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "ai-chips", "large-language-models"], "entities": ["Meta", "Muse", "DeepSeek", "DSec"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agent-muse-compute-demand", "markdown": "https://wpnews.pro/news/agent-muse-compute-demand.md", "text": "https://wpnews.pro/news/agent-muse-compute-demand.txt", "jsonld": "https://wpnews.pro/news/agent-muse-compute-demand.jsonld"}}