{"slug": "private-inference-for-coding-agents", "title": "Private Inference for Coding Agents", "summary": "MoonMath AI launched Zro, a private inference endpoint for coding agents that serves open-weight models from EU infrastructure with zero request retention and no training on customer data. The service, starting at $20/month, supports models including MiniMax M3 and GLM-5.2 and offers OpenAI- and Anthropic-compatible APIs for tools like Claude and Codex.", "body_md": "# Private inference for coding agents.\n\nFast, EU-based endpoint for open-weight models. Zero data retention, zero training, and optimized for long-context workloads.\n\n- CodePrivate\n- ModelsOpen\n- SetupMinutes\n\n- API\n- AgentsCLI + IDE\n- InfraEU regions\n\n## Private inference on EU infrastructure.\n\nZro serves coding-agent workloads from EU infrastructure with zero request retention and no training on customer data.\n\nLocation\n\nEU regions\n\nRetention\n\nZero by default\n\nTraining\n\nNever\n\nZro is tuned for long-context, multi-turn coding sessions. Under the endpoint, MoonMath applies HyperQuant compression, custom kernels, and hardware-aware deployment.\n\nCompression\n\nHyperQuant\n\nKernels\n\nCustom attention\n\nHardware\n\nAMD · NVIDIA · TPU\n\n## $ zro launch claude, codex, opencode, hermes, openclaw, pi\n\nInstall the npm package, log in once, then launch supported coding tools with temporary session config.\n\n```\nnpm install -g @moonmath-ai/zro\nzro login\nzro launch claude\nzro launch codex\n```\n\n## Open coding models.\n\nOne endpoint.\n\nOpen-source models are becoming competitive with closed-source systems for coding tasks. [i](https://artificialanalysis.ai/models/capabilities/coding)Zro starts with MiniMax M3 and GLM-5.2, with more open coding models coming soon.\n\n## Built for private inference.\n\nZro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure, with zero request retention, no training on customer data, and setup paths for the tools developers already use.\n\nYes. Zro exposes OpenAI-compatible access for chat completions, so existing clients and agent tools can point at the Zro base URL.\n\nYes. Zro also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.\n\nUse the `@moonmath-ai/zro`\n\nnpm package for launch-supported tools: run `zro login`\n\nonce, then launch your harness with `zro launch`\n\n. The [integrations page](/integrations) also covers manual setup for Cursor and Cline.\n\nNo. Prompt and completion bodies are not retained by default after inference is processed.\n\nYes. Zro is built for responsive, streaming inference, so developer tools, agents, and production apps do not have to trade speed for privacy.\n\nMiniMax M3 and GLM-5.2 are available now, with more open-model options being added across regions.\n\nNo. Customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.\n\nZro runs on privacy-forward EU infrastructure. Current regions include Finland and France.\n\nPlans start at $20/month for $60 of inference spend. Usage packs are available without a subscription and expire after 90 days. Plan spend resets monthly.", "url": "https://wpnews.pro/news/private-inference-for-coding-agents", "canonical_source": "https://zro.moonmath.ai/", "published_at": "2026-07-21 09:25:36+00:00", "updated_at": "2026-07-21 09:53:08.207458+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-infrastructure", "ai-startups", "ai-tools"], "entities": ["MoonMath AI", "Zro", "MiniMax M3", "GLM-5.2", "Claude", "Codex", "AMD", "NVIDIA"], "alternates": {"html": "https://wpnews.pro/news/private-inference-for-coding-agents", "markdown": "https://wpnews.pro/news/private-inference-for-coding-agents.md", "text": "https://wpnews.pro/news/private-inference-for-coding-agents.txt", "jsonld": "https://wpnews.pro/news/private-inference-for-coding-agents.jsonld"}}