Skip to content
ai tools

Magnitude’s 2x Speed Claim Has a Catch

A near-2x benchmark sounds like a breakthrough—until you look at what was measured, on which Mac, and against what setup. The bigger story may be less about a universal speed win and more about a new way to build local AI tools.

Theo Brandt
Magnitude’s 2x Speed Claim Has a Catch

The clever trick: benchmark the Mac, not the average machine

Engines like llama.cpp and Ollama ship precompiled kernels. This maximizes portability; one build works across diverse chips. However, this convenience often leaves performance on the table for any specific device.

Magnitude takes a different tack. When you download a model, it runs a series of micro-benchmarks directly on your Mac. It tests various kernel configurations, times them, and caches the fastest one. This on-device tuning process reportedly takes about a minute.

You’ll encounter two key performance metrics: prefill and decode. Prefill measures how fast the engine processes your prompt. Decode, on the other hand, tracks the speed of token generation, one by one.

Decode speed is often the more noticeable factor in interactive agent workflows, particularly when you’re waiting for a coding agent to stream its output. Magnitude’s claim of 2x speed specifically targets this decode performance.

Built for agents—and impressively easy to connect

Practical setup is straightforward. Choose a recommended model from the Discover tab, download it, and Magnitude tunes it for about a minute. Then, connect Claude Code, Codex, OpenCode, Cline, or Pi via its local compatible APIs—both OpenAI and Anthropic API endpoints expose on localhost.

Magnitude’s agent-focused design differentiates it. Engine optimizes for multi-session workflows:

  • Shared prompt caches across sessions
  • Dynamically managed memory
  • Compressed KV cache (Keys at 8-bit, Values at 4-bit)
  • Models unload when idle

This positions Magnitude uniquely. llama.cpp and Ollama prioritize broad hardware reach, while MLX specializes in Apple Silicon. Server engines like vLLM target batching for datacenter inference. Magnitude, however, focuses on local tuning and agent workflows. It measures your exact machine, tunes for it, and builds around long-running agents. This approach yields specific advantages for developers running local AI assistants.

The almost-2x result doesn’t tell the whole story

Headline test results for Magnitude’s 2x claim show nuance. On an M4 Pro, Qwen 3.6 35B A3B (4-bit, 64K context) hit 57 tokens per second for decode, a 92% gain over llama.cpp’s 30 t/s. Prefill, however, improved by only 9%. The "2x" is largely a decode-centric metric.

This comparison isn’t universal. An Nvidia DGX Spark running the same test saw decode gains drop to 19%. Critically, the M4 Pro test ran with differing KV-cache configurations; llama.cpp’s compressed-cache path failed, forcing it to use a full 16-bit cache while Magnitude used its default compressed cache.

Independent reports further complicate the picture. Some users found MLX or llama.cpp faster. Reports include llama.cpp leading Magnitude on M5 Max and RTX 5070 Ti, and MLX outperforming Magnitude on an M5 Max (175 t/s vs. 161 t/s). Hardware, model, and specific settings clearly dictate performance. For deeper technical dives, check the GitHub - magnitudedev/magnitude: Open source inference engine for agents repository.

Magnitude is still early (v0.2.5), and its developers note it doesn't yet leverage the M5’s new Metal Matrix hardware. Expect more independent MLX comparisons as the project matures. Raw speed claims, while attention-grabbing, rarely tell the whole story across diverse hardware and workloads.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Who should install it—and who should wait

Magnitude, at version 0.2.5, is an early experiment. M1–M4 Mac owners running local coding agents should install it. The one-click connections to Claude Code, Codex, OpenCode, Cline, or Pi are compelling, often outweighing raw throughput gains for agent workflows.

Expect rough edges. The curated catalog offers roughly 15 models, all 4-bit or higher. Guidance for custom GGUFs is absent, downloaded models hide in a user’s home directory, and idle unloads mean slower first responses. These are minor frictions for tinkerers.

M5 users chasing maximum speed should stick with MLX for now; Magnitude doesn't yet leverage the M5's new Metal Matrix hardware. Production deployments require patience. Wait for broader independent comparisons before committing.

The enduring idea is device-measured tuning, not a guaranteed 2x win. Magnitude’s approach to optimizing kernels for specific hardware is a smarter way to run local LLMs. Even if the 92% decode gain on an M4 Pro with Qwen 3.6 35B A3B at 64K context isn't universal, the principle is sound.

Frequently Asked Questions

What is Magnitude?

Magnitude is an open-source local inference engine designed around coding agents and hardware-specific kernel tuning.

Is Magnitude really twice as fast as llama.cpp?

It recorded nearly twice the decode speed in one M4 Pro test, but independent results vary by hardware, model, and configuration.

How does Magnitude tune itself to a Mac?

It benchmarks different kernel configurations on the device, selects the fastest options, and caches them for later use.

Can Magnitude connect to Claude Code?

Yes. Its local OpenAI- and Anthropic-compatible APIs support a one-click connection to Claude Code and other agent tools.

Who should try Magnitude now?

Mac users on M1 through M4 who want a simple local backend for coding agents may find it useful; production users should wait for more comparisons.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$199 · AI tools & software only

For builders

This page is doing a job for someone else’s tool.

AI agents read it. Buyers land on it. It answers in eight languages and over MCP. Your tool can have one like it — live in 24 hours.