cd /news/ai-tools/magnitudes-2x-speed-claim-has-a-catc… · home › topics › ai-tools › article
[ARTICLE · art-148292] src=stork.ai ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Magnitude’s 2x Speed Claim Has a Catch

Magnitude v0.2.5, an open-source local inference engine for AI agents, claims a 2x decode-speed advantage over llama.cpp by micro-benchmarking each Mac on download and caching the fastest kernel configuration, but its own headline test on an M4 Pro running Qwen 3.6 35B A3B (4-bit, 64K context) showed 57 tokens per second decode — a 92% gain over llama.cpp's 30 t/s — while prefill improved only 9%. The comparison is not universal: an Nvidia DGX Spark saw decode gains drop to 19%, and the M4 Pro test ran with differing KV-cache configurations because llama.cpp's compressed-cache path failed and fell back to a full 16-bit cache. Independent reports found MLX or llama.cpp faster in some cases, including MLX hitting 175 t/s versus Magnitude's 161 t/s on an M5 Max, and the developers note Magnitude does not yet leverage the M5's new Metal Matrix hardware.

by read4 min views5 publishedOct 9, 2026
Magnitude’s 2x Speed Claim Has a Catch
Image: Stork (auto-discovered)

The clever trick: benchmark the Mac, not the average machine #

Engines like llama.cpp and Ollama ship precompiled kernels. This maximizes portability; one build works across diverse chips. However, this convenience often leaves performance on the table for any specific device.

Magnitude takes a different tack. When you download a model, it runs a series of micro-benchmarks directly on your Mac. It tests various kernel configurations, times them, and caches the fastest one. This on-device tuning process reportedly takes about a minute.

You’ll encounter two key performance metrics: prefill and decode. Prefill measures how fast the engine processes your prompt. Decode, on the other hand, tracks the speed of token generation, one by one.

Decode speed is often the more noticeable factor in interactive agent workflows, particularly when you’re waiting for a coding agent to stream its output. Magnitude’s claim of 2x speed specifically targets this decode performance.

Built for agents—and impressively easy to connect #

Practical setup is straightforward. Choose a recommended model from the Discover tab, download it, and Magnitude tunes it for about a minute. Then, connect Claude Code, Codex, OpenCode, Cline, or Pi via its local compatible APIs—both OpenAI and Anthropic API endpoints expose on localhost.

Magnitude’s agent-focused design differentiates it. Engine optimizes for multi-session workflows:

  • Shared prompt caches across sessions
  • Dynamically managed memory
  • Compressed KV cache (Keys at 8-bit, Values at 4-bit)
  • Models unload when idle

This positions Magnitude uniquely. llama.cpp and Ollama prioritize broad hardware reach, while MLX specializes in Apple Silicon. Server engines like vLLM target batching for datacenter inference. Magnitude, however, focuses on local tuning and agent workflows. It measures your exact machine, tunes for it, and builds around long-running agents. This approach yields specific advantages for developers running local AI assistants.

The almost-2x result doesn’t tell the whole story #

Headline test results for Magnitude’s 2x claim show nuance. On an M4 Pro, Qwen 3.6 35B A3B (4-bit, 64K context) hit 57 tokens per second for decode, a 92% gain over llama.cpp’s 30 t/s. Prefill, however, improved by only 9%. The "2x" is largely a decode-centric metric.

This comparison isn’t universal. An Nvidia DGX Spark running the same test saw decode gains drop to 19%. Critically, the M4 Pro test ran with differing KV-cache configurations; llama.cpp’s compressed-cache path failed, forcing it to use a full 16-bit cache while Magnitude used its default compressed cache.

Independent reports further complicate the picture. Some users found MLX or llama.cpp faster. Reports include llama.cpp leading Magnitude on M5 Max and RTX 5070 Ti, and MLX outperforming Magnitude on an M5 Max (175 t/s vs. 161 t/s). Hardware, model, and specific settings clearly dictate performance. For deeper technical dives, check the GitHub - magnitudedev/magnitude: Open source inference engine for agents repository.

Magnitude is still early (v0.2.5), and its developers note it doesn't yet leverage the M5’s new Metal Matrix hardware. Expect more independent MLX comparisons as the project matures. Raw speed claims, while attention-grabbing, rarely tell the whole story across diverse hardware and workloads.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Who should install it—and who should wait #

Magnitude, at version 0.2.5, is an early experiment. M1–M4 Mac owners running local coding agents should install it. The one-click connections to Claude Code, Codex, OpenCode, Cline, or Pi are compelling, often outweighing raw throughput gains for agent workflows.

Expect rough edges. The curated catalog offers roughly 15 models, all 4-bit or higher. Guidance for custom GGUFs is absent, downloaded models hide in a user’s home directory, and idle unloads mean slower first responses. These are minor frictions for tinkerers.

M5 users chasing maximum speed should stick with MLX for now; Magnitude doesn't yet leverage the M5's new Metal Matrix hardware. Production deployments require patience. Wait for broader independent comparisons before committing.

The enduring idea is device-measured tuning, not a guaranteed 2x win. Magnitude’s approach to optimizing kernels for specific hardware is a smarter way to run local LLMs. Even if the 92% decode gain on an M4 Pro with Qwen 3.6 35B A3B at 64K context isn't universal, the principle is sound.

Frequently Asked Questions #

What is Magnitude?

Magnitude is an open-source local inference engine designed around coding agents and hardware-specific kernel tuning.

Is Magnitude really twice as fast as llama.cpp?

It recorded nearly twice the decode speed in one M4 Pro test, but independent results vary by hardware, model, and configuration.

How does Magnitude tune itself to a Mac?

It benchmarks different kernel configurations on the device, selects the fastest options, and caches them for later use.

Can Magnitude connect to Claude Code?

Yes. Its local OpenAI- and Anthropic-compatible APIs support a one-click connection to Claude Code and other agent tools.

Who should try Magnitude now?

Mac users on M1 through M4 who want a simple local backend for coding agents may find it useful; production users should wait for more comparisons.

── more in #ai-tools 4 stories · sorted by recency
── more on @magnitude 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/magnitudes-2x-speed-…] indexed:0 read:4min 2026-10-09 · —