{"slug": "launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents", "title": "Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents", "summary": "Magnitude, a Y Combinator S25 company, launched an open source inference engine for AI agents that compiles and tunes its kernels on the user's own device, claiming up to 2x faster performance than llama.cpp — 92% faster decode on Metal and 19% on CUDA. The Apache 2.0 desktop app for macOS, Windows, and Linux connects to existing agents including Pi, OpenCode, Hermes, Codex, Claude Code, and Cline, and uses 27% less memory per agent, with all prompts, files, and models staying on the machine.", "body_md": "**Run open models as fast as your hardware allows**\n\nMagnitude is an open source inference engine for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the agent you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.\n\n[Download Magnitude for macOS, Windows, or Linux](https://magnitude.dev/download)\n\n⭐ Help us reach more developers and grow the Magnitude community. Star this repo!\n\n## demo-9-29.mp4\n\n1. [Download Magnitude](https://magnitude.dev/download) , install it, and open the app.\n2. Choose a recommended model in **Discover** and download it.\n3. Connect your agent in **Connections** and start using it.\n\nThe desktop app includes the `magnitude` CLI. No separate installation is needed.\n\n- **Up to 2x faster than llama.cpp:** 92% faster decode on Metal, 19% on CUDA\n- **Tuned on your device:** kernels are tuned on your hardware before a model runs\n- **Built for the best models:** hand-optimized kernels for popular open-weight families\n- **Memory that flexes:** 27% less memory per agent, freed when agents stop\n- **Fast concurrent sessions:** sessions share prefix caches to prevent slowdown\n- **Works with your agent:** one click to connect Pi, OpenCode, Hermes, Codex, and more\n- **Free, private, open source:** no token costs, nothing leaves your machine, Apache 2.0\n\nAn open source inference engine that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the agent you already use.\n\nThey ship kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip. [See the benchmarks against llama.cpp.](#up-to-2x-faster-than-llamacpp)\n\nAny Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and more memory lets you run larger ones.\n\nmacOS, Linux, and Windows.\n\nSee the full list at [magnitude.dev/models](https://magnitude.dev/models). We write optimized kernels for the most popular open-weight families, which is how we beat generalist engines.\n\nOne click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works through the OpenAI-compatible API.\n\nYes. Prompts, files, and models stay on your machine. No internet needed once a model is downloaded.", "url": "https://wpnews.pro/news/launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents", "canonical_source": "https://github.com/magnitudedev/magnitude", "published_at": "2026-09-30 17:37:40+00:00", "updated_at": "2026-09-30 17:48:49.681670+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["Magnitude", "Y Combinator", "llama.cpp", "Pi", "OpenCode", "Hermes", "Codex", "Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents", "markdown": "https://wpnews.pro/news/launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents.md", "text": "https://wpnews.pro/news/launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents.txt", "jsonld": "https://wpnews.pro/news/launch-hn-magnitude-yc-s25-self-optimizing-inference-engine-for-agents.jsonld"}}