Apple Neural Engine for Local LLMs in 2026: The 1 MiB Bug That Halved Token Speed, and Why the GPU Still Wins
TL;DR: Two reverse-engineering posts about the Apple Neural Engine hit the Hacker News front page in mid-September, and the headline number — 50 GB/s recovered, token speed up 2.4× — is real. It’s also a bug fix, not a breakthrough: even after the fix, an 8B model decodes at 2.97 tok/s on the M3’s ANE while the same chip’s GPU has 30% more measured DRAM bandwidth and none of the ANE’s context limits. The ANE is a tokens-per-watt engine, not a tokens-per-second engine, and nothing in these posts changes what hardware you should buy.
| Measured on the same base M3 | Neural Engine | GPU (Metal) |
|---|---|---|
| Sustained DRAM read bandwidth | 38 GB/s weights + 59 GB/s activations, serialized | 77.7 GB/s |
| Qwen3-8B decode | 2.97 tok/s (after the fix; 1.36 before) | MLX runs 8B-class models several times faster |
| Practical context window | 512–2048 tokens | 32K–128K routine |
Honest take: Read these posts because they’re the best public documentation of Apple’s NPU that has ever existed — then keep running your models on the GPU with MLX, and keep sizing Mac purchases by memory bandwidth, not TOPS.
What actually happened
On August 10, 2026, Eileen Yoon — the engineer behind the reverse-engineered open-source Linux driver for the Apple Neural Engine (eiln/ane) — published two posts that later spent a weekend on the Hacker News front page: Retrospectively Reverse-Engineering Apple’s Neural Engine and Getting 50 GB/s Back Out of the ANE.
The first maps the ANE’s full internal architecture on the M1 — compute cores, scheduler, DMA engines, memory hierarchy — three years after Yoon shelved the driver project with the blunt assessment that the ANE block “is just not that useful.” The second documents a hardware performance bug she found in the M3’s Neural Engine while profiling, plus a software workaround that more than doubles LLM decode speed for affected models.
A separate, concurrent effort deserves mention because the trending coverage kept blending the two: Spencer Bryngelson’s ANEForge (a pip install aneforge Python runtime that dispatches directly to the ANE, bypassing Core ML entirely) and its companion papers, ANEForge: Python for direct computation on the Apple Neural Engine and Apple Neural Engine: Architecture, Programming, and Performance, the latter a 2026 reverse-engineered reference covering A11 through A18 and M1 through M5. Two independent groups spent 2026 prying open the same undocumented accelerator. What they found is consistent — and consistently sobering for anyone hoping the NPU would rescue local LLM inference.
The bug: a DMA prefetcher that chokes on round numbers
Yoon was profiling weight-streaming throughput for single-token decode on a fanless M3 MacBook Air when she noticed something absurd: a matrix with inner dimension D=1536 streamed nearly 3× faster than D=2048 — the exact dimension Llama 3.2 uses. At D=2016, throughput was 44.5 GB/s. At D=2048, it collapsed to 16.9 GB/s.
Sweeping the full range showed the pattern: whenever the total kernel (weight) transfer per core lands on an exact multiple of 1 MiB, throughput drops from the nominal 45–60 GB/s to a hard floor of 17–19 GB/s, recovering to normal just 16 KiB away on either side. The cause, after ruling out DRAM bank contention and core contention, is almost certainly a register-transfer-level (RTL) bug in the kernel DMA engine’s speculative prefetch ring — pointer arithmetic done in 14 bits:
// Suspected bug: prefetch distance computed modulo 0x4000 lines (1 MiB)
distance = end_ptr - rd_ptr; // 14-bit wraparound
// A transfer of exactly k * 0x4000 lines aliases
// "one full lap remaining" to "empty":
// the prefetcher starves and the transfer crawls at 17-19 GB/s.
// The fix Apple would need in silicon:
distance = transfer_lines - issued_lines; // 32-bit
A transfer sized at an exact multiple of 16,384 64-byte lines ends with the ring pointer exactly where it started, so the prefetcher believes there is nothing left to fetch and stops issuing lookahead requests. The transfer still completes correctly — it’s a performance erratum, not a correctness bug — but a streaming transfer degrades into stop-and-go.
The software workaround is almost comically simple: don’t request 1 MiB transfers. Split them into two 512 KiB chunks. Yoon’s measurements on the M3:
| Weights per core | Unsplit | Split into 512 KiB chunks | Speedup |
|---|---|---|---|
| 1 MiB | 17.3 GB/s | 43.5 GB/s | 2.51× |
| 2 MiB | 18.4 GB/s | 52.1 GB/s | 2.84× |
| 4 MiB | 18.8 GB/s | 57.8 GB/s | 3.07× |
| 8 MiB | 19.1 GB/s | 60.5 GB/s | 3.16× |
The control case is what makes this rigorous: splitting a transfer that was not a 1 MiB multiple produces no speedup at all. The gain comes entirely from dodging the prefetch bug.
Why “round” model dimensions made it worse
Here’s the cruel part: transformer weight matrices love powers of two. A 2048×8192 projection at FP16 is exactly 2 MiB per core lane on a 16-core ANE. Yoon audited ANEMLL — the open-source (MIT) pipeline that converts Llama, Qwen, Gemma, and DeepSeek-distill models to Core ML for ANE execution — and found 7 of its 15 models hit the erratum, including Llama 3.2 1B (gate/up/down projections), Llama 3.1 8B (q and o projections plus MLP), Qwen3-8B, and Gemma 3 4B’s lm_head shard.
Patching the models to emit split transfers produced the numbers that made the HN headline:
- Llama 3.2 1B: 10.0 → 24.3 tok/s (DRAM usage 24.7 → 60.0 GB/s)
- Qwen3-8B: 1.36 → 2.97 tok/s (DRAM usage 22.4 → 48.7 GB/s)
That’s a 2.2–2.4× speedup from a one-line-of-thinking compiler change. If you run ANEMLL (v0.3.5 beta as of September 2026; macOS Sequoia, 16GB+ RAM, 32GB for 8B models), this fix is the difference between an unusable 8B and a barely-usable one.
Now the cold water: 2.97 tok/s is the fixed number
Celebrate the detective work, then look at the absolute values. After recovering the full 50–60 GB/s streaming path, Qwen3-8B decodes at under 3 tokens per second. Comfortable reading speed is 7–10 tok/s. InsiderLLM’s testing of the ANE path on newer, higher-bandwidth M-series chips pegs 1B models at 47–62 tok/s and 8B models around 9 tok/s, while MLX on the same machine’s GPU exceeds 93 tok/s on the same 8B model — a 2–5× GPU advantage before you even mention context length, where ANE conversions cap at 512–4096 tokens (ANEMLL recommends 512–1024) versus the 32K–128K that’s routine in llama.cpp and MLX.
Yoon’s retrospective explains why the ANE loses decode even when both engines share the same unified memory. She measured sustained DRAM read bandwidth on the base M3, whose LPDDR5-6400 ceiling is 102.4 GB/s:
- ANE KernelDMA (weights): 37.99 GB/s
- ANE TileDMA (activations): 59.08 GB/s
- GPU via Metal: 77.70 GB/s
Worse, the kernel and tile DMA engines don’t overlap — their combined runtime equals the sum of the isolated runtimes. Single-token decode is the worst-case workload for this design: you stream every weight in the model to produce one token, there is no reuse, and whichever engine reads DRAM faster wins. On the M1, the ANE’s roofline ridge point sits at roughly 162 operations per byte (11 TOPS against 68 GB/s) — autoregressive decode delivers a tiny fraction of that arithmetic intensity, so the ANE’s 2,048 parallel MAC lanes mostly sit idle waiting on memory.
The architecture makes the point even more sharply: each of the 16 cores has a private 64 KiB kernel memory that can only be filled from DRAM — there is no path to stream weights from the shared 2 MiB L2. That was a perfectly sensible decision for 2017-era CNNs, where a small kernel is loaded once and reused across an entire image. Transformers broke the assumption. As Yoon puts it, Apple “simply never expected L2-resident tensors to become kernels.”
What the ANE is actually good for
None of this means the ANE is junk. It means it’s specialized. Bryngelson’s measurements make the case from the other direction: on compute-dense workloads, the ANE beats the GPU handily — a 256-channel 3×3 convolution runs 3.8× faster and 9× more energy-efficient than the GPU, and across ANEForge’s test workloads the ANE delivers 8–16× better energy efficiency. ResNet-18 inference on an M5 Pro: 0.33 ms on the ANE versus 2.03 ms on the GPU.
ANEForge itself is the most interesting practical artifact here. It’s a real pip install aneforge package (MIT license, M1–M5, macOS 14+, Python 3.10+) with 58 fused operators that compiles tensor graphs straight to ANE programs — no Core ML. It runs decoder LLMs with a resident KV cache: Qwen3-0.6B at ~75 tok/s, with speculative decoding adding 2.28×. The catch list is honest: a ~2 GB single-program ceiling that forces auto-segmentation of larger models, and a dependency on private Apple framework symbols that could break in any macOS update.
So the realistic ANE use cases in 2026:
- Small models where battery is the constraint. A 1B–4B model at 25–60 tok/s at a few watts, on a fanless machine or an iPhone, with the GPU left free for everything else. This is what ANEMLL was built for.
- Vision and convolution workloads — the thing the silicon was actually designed to do, where it wins on both speed and efficiency.
- Research. For the first time, the ANE has public documentation: ane-guide.readthedocs.io covers the datapath, compiler format, driver, and firmware, and there’s a community effort mapping the roofline across every Apple Silicon chip (ANEForge issue #137).
There’s also a fourth answer, and it’s Apple’s: the M5 folded neural accelerators into the GPU cores. Yoon calls it “the beginning of the end for the standalone NPU” — the compute units survive, but inside the GPU’s dataflow, with the GPU’s memory paths. That is exactly the fix for everything documented above, and it’s why the M5 generation’s headline feature was LLM performance. We covered what that means for buyers in our MacBook Pro M5 Max guide and the M5 Ultra Mac Studio breakdown.
Does any of this change what you should buy?
No — and that’s the useful conclusion. The whole saga is a 3,000-word proof of the rule this site keeps repeating: LLM decode speed is memory bandwidth, full stop. A 50 GB/s NPU decodes worse than a 78 GB/s integrated GPU, which decodes worse than a 936 GB/s used RTX 3090 (~$1,343 average on the used market, September 2026). TOPS numbers, NPU marketing, and clever DMA fixes don’t move that hierarchy; they only decide how efficiently you sit inside it. It’s the same conclusion our NPU vs GPU comparison reached from Windows-laptop data: NPUs win tokens-per-watt, GPUs win tokens-per-second, and decode belongs to whoever reads DRAM fastest.
If you’re on a Mac, run your models on the GPU via MLX — see our Ollama MLX guide — and buy by bandwidth tier: base chips (~100–150 GB/s) for 8B-class models, Max chips (410–614 GB/s) for 70B-class, as laid out in our 128GB unified memory guide. If you’re deciding between a Mac and a discrete-GPU tower, the $3K/$6K/$10K build comparison runs the numbers. And if you just want to experiment with what a given card can do before spending anything, renting a 3090 on Vast.ai from $0.07/hr costs less than a sandwich.
For the open-source inference stacks that sit on top of all this hardware — MLX, Core ML tooling, llama.cpp — our sister site aifoss.dev tracks the framework side, and aicoderscope.com covers wiring local models into coding tools.
FAQ
Can I get the 50 GB/s fix on my Mac today? Only if you run models through ANEMLL, where the affected models are being patched to emit split transfers. Ollama, LM Studio, and MLX don’t touch the ANE for LLM inference, so they were never affected. Core ML apps that ship their own models remain at the mercy of whether the transfer sizes land on 1 MiB multiples.
Does the bug affect M1, M2, or M4? Yoon’s measurements are on the M3. The Bryngelson project has documented other silicon-dependent quirks appearing and disappearing across generations (one DMA clamp present on M1/M2 Pro is gone by M5), so assume nothing transfers between generations until measured — that’s precisely what the community roofline-mapping effort is for.
Is the ANE ever faster than the GPU for LLMs? For decode throughput, no — its measured DRAM streaming bandwidth is lower than the GPU’s on the same chip, and its two DMA engines serialize. Where it wins is efficiency: 8–16× better energy per inference on compute-dense workloads, which matters for battery-powered, always-on, or background inference with small models.
Why did Apple design it this way? The ANE was architected around 2017 CNN workloads: small kernels loaded once into per-core memory and reused across an image. Weight streaming for autoregressive transformers — reading gigabytes of weights per token — is the exact opposite access pattern. The M5’s move of neural accelerators into the GPU cores is Apple’s own admission of that mismatch.
Should I buy Apple Silicon or a discrete GPU for local AI? Depends on the model size and your power budget. For 70B+ models, high-memory Macs are the practical path for most people. For maximum tok/s per dollar under 24GB of VRAM, a used RTX 3090 remains the value pick. Full reasoning in our GPU buying guide.
Sources
- Getting 50 GB/s Back Out of the ANE — Eileen Yoon
- Retrospectively Reverse-Engineering Apple’s Neural Engine — Eileen Yoon
- eiln/ane: reverse-engineered Linux ANE driver — GitHub
- ANEMLL: LLMs on the Apple Neural Engine via Core ML — GitHub
- ANEForge: Pythonic binding to the Apple Neural Engine — GitHub
- Apple Neural Engine: Architecture, Programming, and Performance — arXiv:2606.22283
- ANEForge: Python for direct computation on the Apple Neural Engine — arXiv:2606.17090
- Apple Neural Engine: A Complete Guide — ane-guide.readthedocs.io
- Apple Neural Engine for LLM Inference: What Actually Works — InsiderLLM
- Deploying Transformers on the Apple Neural Engine — Apple Machine Learning Research
- Used RTX 3090 price tracking — ResalePrices
- Help map the ANE roofline across every Apple Silicon chip — ANEForge issue #137
Last updated September 20, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.