Intelligence per Watt: What Stanford's Local AI Efficiency Metric Means for Your Home Lab in 2026
TL;DR: Stanford’s intelligence-per-watt (IPW) study measured 20+ local models on 8 accelerators across 1M real queries and found local AI efficiency improved 5.3× in two years — but a local box still delivers at least 1.4× less intelligence per watt than the same model on cloud hardware. For home labs, the metric reranks nothing at the top: it explains why Apple Silicon and Strix Halo win always-on duty while a power-capped used RTX 3090 wins interactive speed.
| Used RTX 3090 tower | Strix Halo mini PC (128GB) | Mac Studio M5 Max (128GB) | |
|---|---|---|---|
| Best for | Fastest tokens under 24GB | Always-on 120B-class serving | Same capacity, lowest idle |
| Tokens per joule (measured class) | ~0.33 (7B), ~0.13 (27B dense) | ~0.34 (gpt-oss-120b) | ~0.7–1.5 (gpt-oss-120b) |
| Whole-system idle | 80–120W | ~30–35W | ~10–15W |
| Street price | ~$1,150–$1,350 card / ~$2,700 build | ~$2,200 (GMKtec EVO-X2) | $5,099 |
| The catch | 3× the idle power of the others | ~31 tok/s ceiling on 120B MoE | Efficiency never repays the premium |
Honest take: Buy on VRAM and tokens/sec first, exactly as before — then let watts break the tie. The one genuinely new takeaway from the paper is that “local is the green option” is false today: per unit of intelligence, your basement rig burns more power than the cloud. Run local for privacy, control, and flat costs, not for the planet.
A Stanford Hazy Research paper that quietly shipped last November hit the Hacker News front page on September 17 (161 points) and put a name on something home labbers have been arguing about in comment threads for years: how do you compare a 35W NPU, a 90W Mac, and a 350W GPU when they’re all “running local AI”? The paper’s answer is intelligence per watt — and unlike the TOPS numbers on marketing slides, it’s measured on real workloads with a methodology you can run on your own hardware tonight.
This piece unpacks what the metric actually measures, applies it to the hardware this site already benchmarks, and answers the buying question: should intelligence per watt change what you order? (Spoiler: it changes when certain machines win, not which machines are good. Check your own model-and-context math against the VRAM calculator before any of it matters.)
What the metric actually measures — and what it deliberately isn’t
Intelligence per watt is task accuracy divided by mean power draw, measured per query. The numerator is not tokens: the Stanford team (Hazy Research and the Scaling Intelligence Lab, with Together AI) scored each local model’s answers as a win rate against frontier cloud models on 1M real single-turn chat and reasoning queries. The denominator is measured electrical power during that query, captured through NVIDIA NVML, AMD ROCm, Apple’s powermetrics, or Linux RAPL depending on the box. The companion metric, intelligence per joule (IPJ), swaps mean power for total energy consumed.
That numerator is the whole point, and it’s what separates IPW from the tokens-per-watt arithmetic you see on r/LocalLLaMA. A Raspberry Pi running a 1B model at 3W might post a respectable tokens-per-joule figure while producing answers that lose to a frontier model 95% of the time. IPW scores that configuration near zero, because it delivers almost no intelligence per unit of power — it just delivers cheap tokens. Tokens per joule tells you what your electricity buys; intelligence per watt tells you whether it bought anything useful.
The study covered 20+ local models — Qwen3, gpt-oss, Gemma 3, and IBM Granite 4 on the current end, Mixtral-8x7B and Llama-3.1-8B representing 2023–24 — across 8 accelerators spanning NVIDIA, AMD, Apple, and SambaNova silicon.
The three numbers worth remembering
Out of the paper’s 1M-query sweep, three findings matter for anyone speccing home hardware:
1. Local models can now handle 88.7% of single-turn chat and reasoning queries. Accuracy varies by domain, but the coverage number — queries where some local model matches frontier-quality output — rose from 23.2% in 2023 to 71.3% in 2025. The gap between “toy” and “daily driver” closed fast, which matches what our September open-model roundup found from the leaderboard side.
2. Efficiency improved 5.3× in two years — and models did most of the work. The paper decomposes the gain: 3.1× from model improvements (MoE sparsity, better training, quantization-aware releases) and 1.7× from hardware, measured across the span from a 2023-era Quadro RTX 6000 to an Apple M4 Max. That ratio is a buying signal in itself: the efficiency of the software you’ll run in 2027 will improve faster than the hardware you’d wait for. It’s the same conclusion our mid-2026 deep dive on MoE and speculative decoding reached from the throughput side.
3. Local accelerators deliver at least 1.4× less IPW than cloud accelerators running the identical model. Qwen3-32B on an M4 Max posts 1.5× lower intelligence per watt than the same model on an NVIDIA B200; SambaNova’s SN40L reaches up to 1.78× the M4 Max’s IPW. Batched serving, better cooling, and datacenter-grade silicon are simply more thermodynamically efficient per answer. So the honest framing for local AI in 2026: you run it for privacy, latency, flat marginal cost, and offline capability — the energy argument belongs to the cloud, at least until local accelerators close what the authors call significant optimization headroom.
Tokens per joule on hardware you can actually buy
The paper’s accuracy-scored metric needs their harness to reproduce, but the site has measured enough (speed, wall power) pairs that we can compute the simpler cousin — tokens per joule — for the machines readers actually ask about. Every number below is from a source already cited in a linked article; none are new estimates.
| Hardware | Model class | Measured speed | Inference draw | Tokens per joule |
|---|---|---|---|---|
| Laptop NPU (Snapdragon X Elite / Lunar Lake) | 8B Q4 | ~10–20 tok/s | ~35W | ~0.3–0.6 |
| Ryzen AI Max+ 395 iGPU | 8B-class Q4 | ~48–61 tok/s | ~80W | ~0.6–0.8 |
| Ryzen AI Max+ 395 (EVO-X2 / N5 Max) | gpt-oss-120b MoE | ~31 tok/s | ~90W | ~0.34 |
| Mac Studio M5 Max 128GB | gpt-oss-120b MoE | 65–88 tok/s | ~60–90W | ~0.7–1.5 |
| Used RTX 3090, stock 350W | 7B Q4 | ~95 tok/s | ~285W | ~0.33 |
| Used RTX 3090, 250W cap | Qwen3.6-27B dense Q4 | 31.7 tok/s | ≤250W | ~0.13 |
Two readings of that table, one naive and one correct.
The naive reading: Apple Silicon embarrasses everything, the NPU is fine, and the 3090 is a space heater. That’s the tokens-per-joule trap the paper was written to correct. The NPU’s 0.3–0.6 tok/J comes from an 8B model — a different intelligence class than the 120B MoE the Mac and Strix Halo rows are running. Weight the rows by answer quality, IPW-style, and the NPU row collapses while the Mac and Strix Halo rows strengthen: they’re producing frontier-adjacent answers at laptop wattage. Our NPU vs GPU analysis reached the same verdict from the bandwidth side — the NPU’s real win is tokens per watt within its small-model class, not intelligence.
The correct reading: unified-memory architectures (Apple, Strix Halo) dominate efficiency for big-model inference, and discrete GPUs dominate raw speed while paying roughly 3× the energy per token. Both facts were already in the site’s benchmarks; IPW just gives them one axis.
Also visible in the table: the 3090’s 27B-dense row is poor because dense models read every weight for every token. Swap in a comparably capable MoE (the paper’s 3.1× model-side gain, in action) and the same silicon roughly doubles its efficiency. Your model choice moves IPW more than your hardware choice.
The duty-cycle blind spot
IPW is a per-query metric, and your electricity meter doesn’t bill per query. A personal AI box spends the overwhelming majority of wall-clock time idle — and idle is where the architectures really diverge. From our power-bill breakdown: a complete used-3090 tower idles at 80–120W, a Strix Halo machine at ~30–35W, a Mac Studio at 8–15W.
Run the year on the US national average of 17.65¢/kWh (EIA, February 2026): a ~90W idle-draw gap between a 3090 tower and a Mac Studio is roughly 788 kWh, or about $139/year — $236+ in California at 30¢/kWh. Real money, but hold it against the hardware gap: a $5,099 Mac Studio M5 Max versus a ~$2,700 used-3090 build is ~$2,400 of premium, or 17 years of break-even at the national average and about a decade in California. The efficiency argument never pays for the Mac. Capacity (128GB unified memory), silence, and form factor pay for the Mac; the power bill is a rounding error on the decision, as our MacBook vs Mac Studio comparison found when it priced the same silicon in three chassis.
The exception where efficiency is the argument: Strix Halo. A GMKtec EVO-X2 at ~$3,649 undercuts the 3090 build on price and idles at a third the wattage and runs the 120B class the 3090 can’t hold. That’s why it keeps winning the always-on rows in our decision tables, most recently the N5 Max AI NAS review.
Measure your own box tonight
The Stanford team open-sourced the whole harness (Apache 2.0), with adapters for Ollama, vLLM, MLX, and OpenAI-compatible servers, and a Rust energy-telemetry service underneath:
git clone https://github.com/HazyResearch/intelligence-per-watt.git
cd intelligence-per-watt
bash intelligence-per-watt/scripts/setup.sh
ipw profile --client ollama --model llama3.2:1b --client-base-url http://localhost:11434
ipw analyze ./runs/profile_*
It captures per-query energy, power, TTFT, throughput, and memory, then computes IPW and IPJ. Note the requirements are current-gen strict (Python ≥3.13, a Rust toolchain), so on a lazy Sunday the ten-second version on any NVIDIA box is:
$ nvidia-smi --query-gpu=power.draw --format=csv -l 1
power.draw [W]
283.47 W
284.12 W
Run it during generation, divide your tokens/sec by that draw, and you have tokens per joule to compare against the table above. On a Mac, sudo powermetrics --samplers gpu_power -i 1000 plays the same role. And if your 3090 is drawing 340W+ during decode, you’re donating watts for nothing: cap it to 280W and keep ~97% of your speed — the single cheapest IPW upgrade in local AI.
Quantization is the second cheap one. Decode is bandwidth-bound, so Q4 reads half the bytes of Q8 per token — nearly double the tokens per joule on identical silicon at identical watts. IPW’s accuracy numerator keeps that honest: Q4_K_M’s measured quality loss is negligible for most tasks, so its IPW is genuinely higher, while below Q4 the accuracy term starts eating the efficiency gain. The sweet spot the site has recommended all year is also the intelligence-per-watt sweet spot.
What to actually buy
Intelligence per watt is a tiebreaker, not a headline spec. Prices as of September 2026, all verified in the linked coverage above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Interactive speed on ≤32B models; box sleeps when you do | Used RTX 3090 tower, power-capped to 280W | ~$1,150–$1,350 card / ~$2,700 build | Check price |
| Always-on 24/7 server running 100B-class MoE | GMKtec EVO-X2 128GB | ~$3,649 | Check price |
| Same capacity, lowest idle and noise, macOS ecosystem | Mac Studio M5 Max 128GB | $5,099 | Check price |
| Power-constrained apartment, 8–14B models are enough | Ryzen AI Max mini PC (64GB) or current MacBook | varies | see mini PC guide |
| Want to test the workload before buying anything | Rented 3090, by the hour | from $0.07/hr | Vast.ai |
If the box will feed a coding agent all day — the highest-duty-cycle workload most readers have — the efficiency rows matter more than usual; our sister site covers local BYOK backends on aicoderscope.com, and aifoss.dev covers the Ollama/vLLM serving stack the IPW harness plugs into.
FAQ
Is intelligence per watt the same as tokens per watt? No. Tokens per watt counts output volume; IPW counts answer quality (win rate against frontier models) per watt. A tiny model can post great tokens-per-watt numbers while scoring terribly on IPW because its answers aren’t competitive.
Does running local AI save energy versus using ChatGPT or Claude? Per the Stanford measurements, no — local accelerators deliver at least 1.4× less intelligence per watt than cloud hardware running the same models. Local wins on privacy, flat cost, and offline capability, not on joules.
What’s the most power-efficient local AI hardware you can buy in September 2026? For big-MoE inference, Apple Silicon (M5 Max/Ultra class) posts the best tokens per joule we can compute from published measurements, with Strix Halo machines close behind at half the price. For small models, NPUs sip the least power but deliver the least intelligence.
Can I improve the efficiency of a GPU I already own? Yes, meaningfully: power-cap it (280W on an RTX 3090 keeps ~97% of decode speed), run Q4_K_M instead of Q8, and prefer MoE models over dense ones. Together those roughly double tokens per joule on the same card.
Will the 5.3× efficiency trend continue? The paper attributes 3.1× of it to model-side advances, and those show no sign of stopping (the 2025–26 MoE wave is exactly this). The hardware side moves slower — another reason not to delay a purchase waiting for efficient silicon.
Recommended Gear
- Used RTX 3090 24GB — still the interactive-speed pick; one
nvidia-smi -pl 280away from respectable efficiency - GMKtec EVO-X2 128GB — the efficiency argument that also wins on price for always-on 120B-class serving
- Mac Studio M5 Max 128GB — best measured tokens per joule, bought for capacity and silence rather than payback math
Sources
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI (arXiv 2511.07885) — Stanford / Together AI
- Intelligence Per Watt: A Study of Local Intelligence Efficiency — Hazy Research, Stanford
- intelligence-per-watt profiling harness (Apache 2.0) — HazyResearch, GitHub
- Intelligence per Watt paper page — Hugging Face
- Intelligence per Watt discussion, Sep 17 2026 — Hacker News
- Intelligence Per Watt: Measuring AI Efficiency — Snorkel AI
- Intelligence Per Watt: Measuring Local Inference Viability (paper summary) — SemiEngineering
- RTX 3090 Power Limit: Finding the Sweet Spot for Local LLM Inference — Jean Brito
- vLLM Performance Benchmarks 4x RTX 3090, Power Limits and NVLink — Himesh Prasad
- LLM Inference Speed Comparison on Ryzen AI Max+ 395 (GPU > CPU > NPU) — Zenn
- Minisforum N5 Max Review with AMD Ryzen AI Max+ 395 (measured idle and inference wattage) — ServeTheHome
- Performance of llama.cpp on Snapdragon X Elite/Plus — GitHub Discussion #8273
- llama.cpp gpt-oss benchmark thread #15396 (M5 Max and RTX decode numbers) — GitHub
- Electric Power Monthly, average residential price — U.S. Energy Information Administration
Last updated September 22, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.