Magnitude Review 2026: The Inference Engine That Profiles Your Mac Before You Download a Single Model

magnitudeapple-siliconlocal-llminference-enginemac

TL;DR: Magnitude (YC S25, Apache 2.0) profiles your chip, memory, and bandwidth, then tells you the tokens/sec of every model in its catalog before you download anything — the exact step Ollama and LM Studio skip. Its headline “up to 2x faster than llama.cpp” number is real but narrower than it sounds: 57 vs 30 tok/s decode on an M4 Pro running a MoE model, a gap MLX-based engines had already closed. The profiling workflow, not the speed, is the feature worth installing.

MagnitudeOllama 0.30+ (MLX)llama.cpp (vanilla Metal)
Best forPicking the right model for your exact hardware, agent backendsSet-and-forget daily driver on 32GB+ MacsMaximum control, every tuning flag
Speed on Apple Silicon57 tok/s decode, Qwen3.6-35B-A3B Q4 on M4 Pro (vendor)~56 tok/s equivalent via MLX (bandwidth-scaled estimate)30 tok/s, same model, same Mac
The catchDay-one project, vendor-run benchmarks, Windows means WSLMLX path needs a 32GB+ Mac to engage by defaultLeaves half your Mac’s bandwidth unused on MoE models

Honest take: Install Magnitude for the profiler — knowing a model’s real tok/s on your machine before a 20GB download is genuinely new. But if you’re already on Ollama 0.30+ with a 32GB Mac, the MLX path gets you most of the speed, and none of the switching cost.

Magnitude launched on Hacker News this morning (Launch HN, YC S25) with a pitch aimed squarely at the most common failure mode in local AI: you pull a model because a Reddit thread liked it, wait out a 20GB download, and discover it either doesn’t fit your memory or decodes at 3 tok/s. The GitHub repo — 5.8k stars within the day, Apache 2.0, 1,061 commits — describes an inference engine that “profiles your machine, recommends the best open models for it, and tunes them for your exact hardware.” It runs on Apple Silicon, NVIDIA, AMD, or plain CPU.

We’ve covered five tools that benchmark your hardware after the fact. Magnitude inverts the order: measure first, download second. That ordering is the interesting part. The “2x faster than llama.cpp” claim needs more scrutiny, and the bandwidth math below shows why.

What Magnitude actually is

Three layers, per the project’s documentation and launch coverage:

  1. A TypeScript CLI and an Electron desktop app — the desktop app bundles the magnitude CLI, so there’s no separate install step.
  2. A Rust “Inference Control Node” (ICN) that profiles hardware, manages model lifecycle, and serves an OpenAI-compatible HTTP API on localhost:8080.
  3. A GPU runtime with a kernel autotuner — the team describes a custom Rust kernel runtime; launch coverage also notes llama.cpp-derived bindings over Metal. In practice: tuned kernels compiled against your specific chip before the model runs, rather than one generic binary for every Mac.

The setup flow profiles your chip, memory capacity, and memory bandwidth, then estimates fit and tokens/sec for every model in its catalog and ranks them on four axes: speed, accuracy, intelligence, and memory footprint. Only after you pick does it download, tune (including speculative decoding setup), and serve.

Memory handling is agent-shaped: Magnitude reserves only enough to hold model weights up front, then grows the heap as agent sessions accumulate KV cache and frees it when sessions stop. The README claims 27% less memory per agent versus its baseline. That matters if you run Cline or Claude Code against a local backend, where several concurrent sessions each drag their own context.

Once running, anything that speaks the OpenAI API can use it:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3.6-35b-a3b", "messages": [{"role": "user", "content": "Say hi"}]}'

Expected shape of the response (illustrative):

{
  "object": "chat.completion",
  "model": "qwen3.6-35b-a3b",
  "choices": [{"message": {"role": "assistant", "content": "Hi!"}, "finish_reason": "stop"}],
  "usage": {"prompt_tokens": 9, "completion_tokens": 3}
}

A one-click “Connections” flow wires that endpoint into Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. For Cline or any BYOK tool, it’s the usual drill: OpenAI-compatible provider, base URL http://localhost:8080/v1, any string as the API key.

The 2x claim, decomposed

Magnitude’s published numbers, from its own benchmark page: Qwen3.6 35B-A3B, 4-bit, 64K context, no speculative decoding, on an M4 Pro Mac with 48GB:

MetricMagnitudellama.cpp (Metal)Delta
Prefill507 tok/s466 tok/s+9%
Decode57 tok/s30 tok/s+90%

The README generalizes this to 92% faster decode on Metal, 19% on CUDA, with prefill +9% (Metal) and +23% (CUDA). Three things to notice before repeating the headline.

First, the numbers pass the physics check. Decode speed is capped by memory bandwidth divided by bytes read per token. An M4 Pro moves 273 GB/s, and Qwen3.6-35B-A3B — a MoE that activates ~3B parameters per token — reads roughly 1.9 GB per token at 4-bit. Ceiling: 273 ÷ 1.9 ≈ 143 tok/s. Magnitude’s 57 tok/s is 40% of that ceiling; at 64K context, with attention overhead growing, that’s a credible number, not marketing fiction. (We run this check on every benchmark we cite; plenty of sites fail it.)

Second, the baseline is llama.cpp’s weakest path. Vanilla llama.cpp on Metal at 30 tok/s is using about a fifth of the M4 Pro’s bandwidth on this MoE model — a known soft spot. It’s the same gap we documented when Ollama 0.30 made MLX the default engine on 32GB+ Macs: the same model family ran ~112 tok/s on MLX versus ~58 tok/s on Metal llama.cpp on an M4 Max. Scale that M4 Max MLX number down by the bandwidth ratio (273/546 GB/s) and you land at ~56 tok/s on an M4 Pro — flag that as a bandwidth-scaled estimate, not a measurement, but it sits right on top of Magnitude’s 57.

In other words: Magnitude’s 2x is real, and it’s the same 2x that MLX already delivered. The honest comparison isn’t Magnitude vs llama.cpp; it’s Magnitude vs Ollama-running-MLX, and on that matchup nobody has published numbers yet. If you benchmark this pairing on your own machine, that’s the result worth posting.

Third, the CUDA story is modest. +19% decode is worth having, but it tells you llama.cpp’s CUDA kernels were already near the bandwidth ceiling. On a used RTX 3090 ($1,150–$1,350 on the used market, September 2026) this model already decodes at ~107 tok/s under llama.cpp — a 19% bump is nice, not transformative. Nobody should buy different silicon because of it.

The problem Magnitude actually fixes

Here’s the failure it targets, which we hit constantly in reader mail: you have an M4 Pro, you heard 35B-class MoE models are the sweet spot, you pull one through a llama.cpp wrapper, and you get 30 tok/s — on a machine whose bandwidth supports over 100 for that workload. Nothing errors. Nothing warns you. You just silently get a third of your hardware.

The fixes, in order of effort:

  1. If you’re on Ollama: upgrade to 0.30+ and confirm the MLX engine engaged — ollama ps should show 100% GPU, and 32GB+ Macs get MLX by default. Our 0.30 upgrade guide covers the cases where it doesn’t.
  2. If you want the choice made for you: Magnitude’s profiler ranks the catalog by projected tok/s on your exact chip before anything downloads, and its tuned Metal path hits the same ~2x-over-vanilla number.
  3. If you’re on vanilla llama.cpp by choice: you’re tuning flags by hand anyway; test both and keep your flags.

That pre-download projection is the part with no real equivalent elsewhere. OMLX tunes Apple Silicon inference but doesn’t rank a catalog against your chip; hardware-checker sites estimate fit but not speed, and several of them publish numbers that fail the bandwidth-ceiling check. Measuring the machine first is the correct design.

Does this change what hardware you should buy?

No — and it’s worth being precise about why, because “2x faster” headlines reliably produce bad purchase decisions.

Kernel tuning changes how much of your bandwidth gets used. It cannot change the bandwidth, and it cannot change memory capacity. The ceilings stay where they were:

MachineBandwidthRealistic 35B-A3B Q4 decodeWhat it can’t do
M4 Pro (48GB)273 GB/s~57 tok/s (Magnitude, measured)Dense 70B at usable speed
M4 Max / M5 Max546 / 614 GB/s~110+ tok/s (MLX, measured on M4 Max)200B+ MoE without 128GB config
Used RTX 3090 24GB936 GB/s~107 tok/s (llama.cpp, measured)Anything over ~22GB VRAM resident
M5 Ultra1,200 GB/snot yet benchmarked with Magnitude—

A software layer that gets you from 21% to 40% of ceiling is worth installing. It is not worth paying for twice the silicon to chase — the M5 Ultra math only makes sense for model classes that physically can’t fit smaller machines, with or without Magnitude. Check your own model-plus-context case against the VRAM calculator before any purchase; if the model fits your current machine, try the free software fix first.

One genuine buying-decision shift: Magnitude raises the floor of what a Mac you already own delivers, which weakens the “upgrade because it’s slow” case. If your 36GB M4 Pro MacBook was feeling inadequate at 30 tok/s, it may be a 57 tok/s machine after a free install. Spend the $0 before the $2,499.

What to actually buy

Prices as of October 2026, all verified in the comparison above or against our September street-price tracking:

Your situationThe movePriceWhere
Own any 16GB+ Apple Silicon Mac, models feel slowInstall Magnitude (or Ollama 0.30+ MLX), re-test$0GitHub
Best decode-per-dollar for a desktop rigUsed RTX 3090 24GB$1,150–$1,350Check price
Buying a Mac for 100B-class MoE at homeMac Studio M5 Max (from $2,499; 128GB config $5,099)$2,499+Check price
Want to test a model class before buying anythingRented 3090, billed hourlyfrom $0.07/hrVast.ai

FAQ

Is Magnitude free and open source? Yes — Apache 2.0, full source on GitHub, no account required, and inference runs entirely locally after download. It’s a YC S25 company, so expect a commercial layer eventually; the engine itself carries no restrictions today.

Does Magnitude replace Ollama? Functionally it can — it downloads, serves, and exposes an OpenAI-compatible API like Ollama does. Whether you should switch depends on whether the profiler and agent-session memory handling matter to you. On a 32GB+ Mac already running Ollama’s MLX path, the raw speed difference is likely small; below 32GB, or on mixed NVIDIA/AMD fleets, Magnitude’s tuning has more room to matter.

Does it run on Windows? Via WSL. macOS and Linux run natively. If your rig is a Windows gaming PC, factor in the WSL GPU passthrough setup before counting on it.

Does it work with NVIDIA and AMD GPUs? Yes — the vendor’s own numbers show +19% decode and +23% prefill over llama.cpp on CUDA. That’s a smaller win than on Metal because llama.cpp’s CUDA path was already efficient. AMD numbers haven’t been published in comparable detail yet; treat ROCm performance as unverified.

Which Mac runs Magnitude best? Same answer as every inference engine, because decode is bandwidth-bound: M5 Ultra (1,200 GB/s) > M5 Max (614) > M4 Max (546) > M4 Pro (273). Software can move you closer to your ceiling; only hardware moves the ceiling. Our 128GB unified memory guide maps model classes to those tiers.

Products linked in this review:

  • Used RTX 3090 24GB — $1,150–$1,350 used; still the decode-per-dollar benchmark at 936 GB/s.
  • Mac Studio M5 Max — from $2,499; the 128GB config ($5,099) is the entry point for 100B-class MoE on one quiet box.

Sources

Last updated October 1, 2026. Prices and specs change; verify current rates before purchasing. Magnitude launched today — expect its numbers and catalog to move fast.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.