{"slug": "magnitudes-2x-speed-claim-has-a-catch", "title": "Magnitude’s 2x Speed Claim Has a Catch", "summary": "Magnitude v0.2.5, an open-source local inference engine for AI agents, claims a 2x decode-speed advantage over llama.cpp by micro-benchmarking each Mac on download and caching the fastest kernel configuration, but its own headline test on an M4 Pro running Qwen 3.6 35B A3B (4-bit, 64K context) showed 57 tokens per second decode — a 92% gain over llama.cpp's 30 t/s — while prefill improved only 9%. The comparison is not universal: an Nvidia DGX Spark saw decode gains drop to 19%, and the M4 Pro test ran with differing KV-cache configurations because llama.cpp's compressed-cache path failed and fell back to a full 16-bit cache. Independent reports found MLX or llama.cpp faster in some cases, including MLX hitting 175 t/s versus Magnitude's 161 t/s on an M5 Max, and the developers note Magnitude does not yet leverage the M5's new Metal Matrix hardware.", "body_md": "## The clever trick: benchmark the Mac, not the average machine\n\nEngines like `llama.cpp` and [Ollama](https://www.stork.ai/en/ollama-2) ship **precompiled kernels**. This maximizes portability; one build works across diverse chips. However, this convenience often leaves performance on the table for any specific device.\n\n[Magnitude](https://www.stork.ai/en/magnitude) takes a different tack. When you download a model, it runs a series of micro-benchmarks directly on your Mac. It tests various kernel configurations, times them, and caches the fastest one. This on-device tuning process reportedly takes about a minute.\n\nYou’ll encounter two key performance metrics: **prefill** and **decode**. Prefill measures how fast the engine processes your prompt. Decode, on the other hand, tracks the speed of token generation, one by one.\n\nDecode speed is often the more noticeable factor in interactive agent workflows, particularly when you’re waiting for a coding agent to stream its output. Magnitude’s claim of 2x speed specifically targets this decode performance.\n\n## Built for agents—and impressively easy to connect\n\nPractical setup is straightforward. Choose a recommended model from the Discover tab, download it, and Magnitude tunes it for about a minute. Then, connect Claude Code, Codex, OpenCode, Cline, or [Pi](https://www.stork.ai/en/inflection-pi) via its local compatible APIs—both OpenAI and Anthropic API endpoints expose on localhost.\n\nMagnitude’s agent-focused design differentiates it. Engine optimizes for multi-session workflows:\n\n- Shared prompt caches across sessions\n- Dynamically managed memory\n- Compressed KV cache (Keys at 8-bit, Values at 4-bit)\n- Models unload when idle\n\nThis positions Magnitude uniquely. llama.cpp and Ollama prioritize broad hardware reach, while MLX specializes in Apple Silicon. Server engines like [vLLM](https://www.stork.ai/en/vllm-runtime) target batching for datacenter inference. Magnitude, however, focuses on local tuning and **agent workflows**. It measures your exact machine, tunes for it, and builds around long-running agents. This approach yields specific advantages for developers running local AI assistants.\n\n## The almost-2x result doesn’t tell the whole story\n\nHeadline test results for Magnitude’s 2x claim show nuance. On an M4 Pro, Qwen 3.6 35B A3B (4-bit, 64K context) hit 57 tokens per second for **decode**, a 92% gain over llama.cpp’s 30 t/s. Prefill, however, improved by only 9%. The \"2x\" is largely a decode-centric metric.\n\nThis comparison isn’t universal. An Nvidia DGX Spark running the same test saw decode gains drop to 19%. Critically, the M4 Pro test ran with differing KV-cache configurations; llama.cpp’s compressed-cache path failed, forcing it to use a full 16-bit cache while Magnitude used its default compressed cache.\n\nIndependent reports further complicate the picture. Some users found MLX or llama.cpp faster. Reports include llama.cpp leading Magnitude on M5 Max and RTX 5070 Ti, and MLX outperforming Magnitude on an M5 Max (175 t/s vs. 161 t/s). Hardware, model, and specific settings clearly dictate performance. For deeper technical dives, check the [GitHub - magnitudedev/magnitude: Open source inference engine for agents](https://github.com/magnitudedev/magnitude) repository.\n\nMagnitude is still early (v0.2.5), and its developers note it doesn't yet leverage the M5’s new Metal Matrix hardware. Expect more independent MLX comparisons as the project matures. Raw speed claims, while attention-grabbing, rarely tell the whole story across diverse hardware and workloads.\n\nEnjoying this? Get one like it in your inbox each morning.\n\none email a day · unsubscribe in two clicks · no third-party tracking\n\n## Who should install it—and who should wait\n\nMagnitude, at version 0.2.5, is an early experiment. M1–M4 Mac owners running local coding agents should install it. The one-click connections to Claude Code, Codex, OpenCode, Cline, or Pi are compelling, often outweighing raw throughput gains for agent workflows.\n\nExpect rough edges. The curated catalog offers roughly 15 models, all 4-bit or higher. Guidance for custom GGUFs is absent, downloaded models hide in a user’s home directory, and idle unloads mean slower first responses. These are minor frictions for tinkerers.\n\nM5 users chasing maximum speed should stick with **MLX** for now; Magnitude doesn't yet leverage the M5's new Metal Matrix hardware. Production deployments require patience. Wait for broader independent comparisons before committing.\n\nThe enduring idea is **device-measured tuning**, not a guaranteed 2x win. Magnitude’s approach to optimizing kernels for specific hardware is a smarter way to run local LLMs. Even if the 92% decode gain on an M4 Pro with Qwen 3.6 35B A3B at 64K context isn't universal, the principle is sound.\n\n## Frequently Asked Questions\n\n### What is Magnitude?\n\nMagnitude is an open-source local inference engine designed around coding agents and hardware-specific kernel tuning.\n\n### Is Magnitude really twice as fast as llama.cpp?\n\nIt recorded nearly twice the decode speed in one M4 Pro test, but independent results vary by hardware, model, and configuration.\n\n### How does Magnitude tune itself to a Mac?\n\nIt benchmarks different kernel configurations on the device, selects the fastest options, and caches them for later use.\n\n### Can Magnitude connect to Claude Code?\n\nYes. Its local OpenAI- and Anthropic-compatible APIs support a one-click connection to Claude Code and other agent tools.\n\n### Who should try Magnitude now?\n\nMac users on M1 through M4 who want a simple local backend for coding agents may find it useful; production users should wait for more comparisons.", "url": "https://wpnews.pro/news/magnitudes-2x-speed-claim-has-a-catch", "canonical_source": "https://www.stork.ai/blog/magnitudes-2x-speed-claim-has-a-catch", "published_at": "2026-10-09 12:51:55+00:00", "updated_at": "2026-10-09 13:51:42.936974+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "ai-infrastructure", "developer-tools", "large-language-models"], "entities": ["Magnitude", "llama.cpp", "Ollama", "MLX", "vLLM", "Qwen 3.6 35B A3B", "Claude Code", "Nvidia DGX Spark"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/magnitudes-2x-speed-claim-has-a-catch", "markdown": "https://wpnews.pro/news/magnitudes-2x-speed-claim-has-a-catch.md", "text": "https://wpnews.pro/news/magnitudes-2x-speed-claim-has-a-catch.txt", "jsonld": "https://wpnews.pro/news/magnitudes-2x-speed-claim-has-a-catch.jsonld"}}