GPT-6 Sol and Luna Cut API Prices in Half: The Buy-a-GPU Math, Rerun for Late 2026

gpt-6cloud-vs-localgpu-roirtx-3090cost-analysislocal-llm

TL;DR: OpenAI launched GPT-6 Sol ($2/$10 per million tokens) and Luna ($0.10/$0.50) on September 22 — half of GPT-5.6 pricing, permanent. That undercuts DeepSeek five weeks after DeepSeek’s own hike briefly made local hardware look competitive. A used RTX 3090 now needs ~156M output tokens a month to beat Luna on cost. Buy the card for privacy and control, not savings.

Monthly cost, ~10M output tokensGPT-6 Luna APIDeepSeek V4.1-Flash (US off-peak)Used RTX 3090, amortized 24 mo
Token bill / hardware~$5.00~$6.00$52.08 + electricity
Electricity$0$0~$1.67
Total~$5~$6~$54
The catchYour data leaves the buildingPeak-hour billing traps night owlsCosts 10× more until volume gets extreme

Honest take: The five-week window where buying a GPU almost penciled out on cost alone just closed. Buy a used RTX 3090 because you want your tokens generated in your own house — privacy, no rate limits, no vendor mood swings — and treat every API price cut as irrelevant to that decision.

In August we wrote that DeepSeek’s planned price hike was the first 2026 API price move that went the wrong way for cloud — and that if it landed hard enough, a used GPU could finally beat the API on raw cost. The hike landed. Then, on September 22, OpenAI walked into the same market and cut prices in half.

This article does three things: documents what actually changed on both rate cards, reruns this site’s standard break-even formula with verified September 2026 numbers, and gives a straight answer on what a $1,150–$1,350 used RTX 3090 is still for when the cheapest capable API costs fifty cents per million output tokens. If you want to plug in your own workload, our local vs. cloud cost calculator runs this exact math interactively.

What OpenAI shipped on September 22

GPT-6 Sol and Luna launched on Monday, September 22, 2026, with API pricing 50% or more below the GPT-5.6 family it replaces. The rate card, cross-checked across Techzine, gHacks, and VentureBeat:

ModelInput /MCached input /MOutput /MPrevious (GPT-5.6)
GPT-6 Sol (flagship)$2.00$0.20$10.00$4.00 / $20.00
GPT-6 Luna (small/fast)$0.10$0.01$0.50$0.20 / $1.20

Three details matter more than the headline halving:

  1. The prices are permanent. An OpenAI spokesperson confirmed to VentureBeat that the new rates carry no expiration date. This is not a promo like DeepSeek’s launch pricing was.
  2. Luna’s output price fell 58%, not 50% — $1.20 to $0.50. The small-model tier, the one that actually competes with what you can run on a consumer GPU, got the deepest cut.
  3. Caching got smarter, not just cheaper. Per Kingy AI’s spec roundup, GPT-6 preserves the prompt cache when an agent changes reasoning effort or toggles tools between turns — the two things that invalidated cached prefixes on GPT-5.6. Both models keep the 1,050,000-token context window (922K input cap, 128K max output). For agentic workloads that re-read the same repository context hundreds of times a day, cached input at $0.01/M on Luna is the number that ends arguments.

OpenAI attributes the cut to caching and inference improvements. The competitive read is simpler: the open-weight ecosystem and DeepSeek had made $1.20/M output look expensive, and OpenAI decided to own the floor instead.

Meanwhile, the DeepSeek floor everyone was standing on moved up

Every break-even we published from May through August used DeepSeek V4-Flash at $0.14/$0.28 per million tokens as the API floor — the price a GPU had to beat. That floor no longer exists.

The timeline, per PricePerToken’s model history: the flat $0.14/$0.28 launch rate was retired on August 16, 2026 — ten days after the “significant increase” notice we covered. On September 10, the deepseek-v4-flash model ID itself was retired, and requests now route to DeepSeek-V4.1-Flash on a time-of-day rate card:

DeepSeek V4.1-FlashInput /MOutput /MCache-hit input /M
Peak (01:00–04:00, 06:00–10:00 UTC, weekdays)$0.30$1.20$0.006
Off-peak (everything else)$0.15$0.60—

Note what happened to the rankings: for a US user paying off-peak rates, DeepSeek output now costs $0.60/M — and GPT-6 Luna costs $0.50/M. OpenAI is now cheaper than DeepSeek on the output side for American daytime workloads. That sentence would have been absurd in May.

One trap worth a paragraph, because it will show up on real invoices. DeepSeek’s peak window is defined in UTC, and 01:00–04:00 + 06:00–10:00 UTC is roughly 9 p.m.–midnight and 2–6 a.m. Eastern. The classic money-saving move — “schedule the big batch job overnight” — lands squarely inside DeepSeek’s peak window and pays $1.20/M instead of $0.60/M, a clean 2× overcharge for doing what used to be the frugal thing. The fix is to run US batch jobs during the US workday (late morning through evening Eastern is off-peak), or route them to Luna, which doesn’t care what time it is. We flagged the earlier version of this billing structure in our peak-pricing analysis; it has since graduated from surcharge experiment to the permanent rate card.

The break-even, rerun with September numbers

Same formula as August’s DeepSeek piece, same measured throughput, updated inputs: a used RTX 3090 at $1,250 (the midpoint of the $1,150–$1,350 September range from BestValueGPU and ResalePrices tracking), amortized over 24 months; US average residential electricity at 18.34¢/kWh (September 2026); and the card generating with Qwen3.6-35B-A3B at the 107 tok/s we benchmarked at ~350W.

# breakeven.py — output tokens/month where a used 3090 beats the API
KWH = 0.1834  # US residential average, Sep 2026

def breakeven_m_tokens(card_price, months, tok_s, watts, api_price_per_m):
    amort = card_price / months                              # $/month
    elec_per_m = (1e6 / tok_s) / 3600 * (watts / 1000) * KWH # $/M tokens
    margin = api_price_per_m - elec_per_m
    return amort / margin if margin > 0 else float("inf")

for name, api in [("GPT-6 Luna", 0.50), ("DeepSeek off-peak", 0.60),
                  ("DeepSeek peak", 1.20), ("GPT-6 Sol", 10.00)]:
    m = breakeven_m_tokens(1250, 24, 107, 350, api)
    print(f"{name} at ${api:.2f}/M: {m:,.0f}M output tokens/month")

Output:

GPT-6 Luna at $0.50/M: 156M output tokens/month
DeepSeek off-peak at $0.60/M: 120M output tokens/month
DeepSeek peak at $1.20/M: 50M output tokens/month
GPT-6 Sol at $10.00/M: 5M output tokens/month

For scale: electricity alone costs this rig about $0.17 per million output tokens — a third of Luna’s all-in price before the card itself costs a cent. And a 3090 generating flat-out, 24/7, produces at most ~277M tokens a month, so the Luna break-even of 156M means running your card at full tilt 13.5 hours a day, every day, just to tie. In August, against a hypothetical 5× DeepSeek hike, the break-even had fallen to 36M — a volume real agentic workflows hit, which is why that article read like the door was finally opening. GPT-6 Luna slammed it: the effective floor for US users went from a threatened $1.40/M back down to $0.50/M in five weeks.

The Sol line needs its own honesty: 5M tokens a month is nothing — two hours of generation — but nothing you run on a $1,250 card is Sol. A 35B-A3B MoE is a Luna-tier workhorse, not a frontier flagship. Comparing your 3090 to Sol pricing flatters the card; the fair fight is against Luna and V4.1-Flash, and the card loses that fight on cost at any sane volume.

The cache twist: input-heavy agents don’t save you either

The standard rebuttal is that agentic coding is input-dominated — 30:1 input-to-output is normal when an agent re-reads a codebase every turn — and local prefill is free. With GPT-6’s cache pricing, run the actual numbers on a genuinely heavy month: 300M input tokens (90% cache-hit, which the effort-change cache preservation makes realistic) plus 10M output on Luna:

  • Fresh input: 30M × $0.10 = $3.00
  • Cached input: 270M × $0.01 = $2.70
  • Output: 10M × $0.50 = $5.00
  • Total: ~$10.70/month — versus $52.08/month in card amortization before your first kilowatt-hour.

A workload big enough to feel like a full-time AI habit now costs about what two coffees do. If your monthly token bill on the small-model tier is under ~$50, hardware cannot pay for itself, full stop. The margin-collapse thesis we wrote in July — you buy the card for sovereignty, not savings — is now true at half the price it was true at then.

What a used 3090 is still for

Here’s the part the API rate card can’t touch, and why this site still runs local hardware daily:

  • Privacy is binary. Tokens generated at home never leave home. If the work involves client code, medical records, or anything under NDA, the API price is irrelevant because the API is unavailable. Our privacy audit covers exactly what stays local.
  • No rate limits, no deprecations, no repricing risk. DeepSeek just demonstrated that a rate card is a promise with a ten-day fuse. GPT-6’s cut is permanent — until it isn’t. A GGUF on your SSD runs identically in 2029.
  • Unlimited experimentation. Fine-tuning runs, 50-attempt sampling sweeps, embedding jobs over your whole archive — volume-insensitive work is where owning the meter wins psychologically, even when renting wins arithmetically.
  • The asset has been holding its value. The same used 3090 sold for ~$1,010 in March and $1,150–$1,350 in September — the DRAM crunch has been appreciating these cards while API prices fell. Nobody should buy a GPU as an investment, but the resale floor makes the real cost of trying local AI much lower than the sticker.

What the card is not for, as of this week: beating OpenAI on cost per token. That case had five weeks of plausibility in August and it’s gone.

What to actually buy

Prices as of September 2026, all verified above:

Your situationThe movePriceWhere
Privacy-bound work, or you want AI that can’t be repricedUsed RTX 3090 24GB$1,150–$1,350Check price
Cost is the only criterion and the data can leave the buildingGPT-6 Luna API$0.10/$0.50 per MOpenAI platform
Want to feel out local models before spending four figuresRented 3090, hourlyfrom $0.07/hrVast.ai

If you’re in the first row, the used RTX 3090 deep dive covers what to check before buying, and the 24/7 power-bill math covers what it costs to keep running. If your local models will back a coding agent, our sister site aicoderscope.com tracks which assistants support local BYOK backends.

FAQ

Did GPT-6 Sol and Luna actually get cheaper, or is this a promo? Permanent, per OpenAI’s statement to VentureBeat on September 22, 2026: Sol at $2/$10 per million tokens and Luna at $0.10/$0.50, with no expiration date. GPT-5.6 was $4/$20 and $0.20/$1.20 respectively.

Is DeepSeek still the cheapest API for local-AI-class models? For US daytime users, no. V4.1-Flash off-peak is $0.15/$0.60; GPT-6 Luna is $0.10/$0.50. DeepSeek’s cache-hit input at $0.006/M peak is still the cheapest single line item, but on output — the number that dominates generation-heavy bills — OpenAI now owns the floor.

How many tokens do I need to generate before a used RTX 3090 beats the API? Against GPT-6 Luna at $0.50/M output: about 156 million output tokens per month, using a $1,250 card amortized over 24 months, 107 tok/s measured throughput, 350W, and 18.34¢/kWh electricity. That’s 13.5 hours of flat-out generation daily. Against GPT-6 Sol it’s only ~5M, but no consumer-GPU model matches Sol’s tier, so that comparison is decorative.

Does the electricity cost alone ever beat the API? No — and this is new. At 18.34¢/kWh, a 3090 generating at 107 tok/s spends ~$0.17 per million output tokens on power. That’s 33% of Luna’s output price before amortizing a single dollar of hardware. In 2025, when small-tier APIs charged $1–2/M, electricity-only local generation was a clear win. The 2026 price war ended that.

Should I wait for GPU prices to drop before going local? Used 3090s went from ~$1,010 in March to $1,150–$1,350 in September, and nobody credible forecasts DRAM-crunch relief before late 2027 — GPU price drops are not on the menu. If you need local AI, buy for the use case now; if you don’t, the API price cuts mean waiting costs you almost nothing.

  • Used RTX 3090 24GB — $1,150–$1,350 used, September 2026. The only hardware this article links, and only for the privacy/control row of the decision table — not as a way to out-price GPT-6 Luna, which it cannot do.

Sources

Last updated September 24, 2026. API rate cards and used-GPU prices change frequently; verify current rates before committing to either path.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.