FlashAttention Only Supports Ampere GPUs or Newer? Every Fix for Turing, Volta, and Pascal Cards (2026)

flashattentionvllmtransformerscudatroubleshootingrtx-2080tesla-t4local-llm

TL;DR: RuntimeError: FlashAttention only supports Ampere GPUs or newer means your GPU’s compute capability is below 8.0 — Turing (RTX 20-series, GTX 16-series, T4), Volta (V100), or Pascal (GTX 10-series, P40). The flash-attn 2 package has never shipped kernels for these cards and, after three years of “support coming soon” in its README, it never will. The fix is to stop requesting FlashAttention: --dtype=half plus the xformers backend in vLLM, attn_implementation="sdpa" in Transformers, or switch to llama.cpp/Ollama, which has its own flash attention that runs fine on Turing.

What you’ll be able to do after this fix:

  • Start vLLM on a Turing card (RTX 2080 Ti, T4, Quadro RTX 5000–8000) without the FlashAttention or bfloat16 crash
  • Load any Hugging Face model on pre-Ampere hardware by swapping one attn_implementation argument
  • Know which errors on old cards are fixable with flags — and which ones (FP8 models, Pascal in vLLM) are hard walls

Honest take: every fix below works today, but pre-Ampere cards are losing software support piece by piece — first flash-attn 2, then bf16, then FP8, now individual model architectures. If you’re hitting this error monthly, the durable fix is the cheapest Ampere card that fits your models, not another environment variable.

What the error actually means

The flash-attn package (Dao-AILab’s FlashAttention-2) ships CUDA kernels only for compute capability 8.0 and up — Ampere, Ada, and Hopper. The project’s own requirements say it plainly: “Ampere, Ada, or Hopper GPUs (e.g., A100, RTX 3090, RTX 4090, H100). Support for Turing GPUs (T4, RTX 2080) is coming soon, please use FlashAttention 1.x for Turing GPUs for now.” That sentence has been in the README since 2023. Turing support never landed, and the project has since moved on to FlashAttention-3, which targets Hopper.

So when vLLM, Transformers, Axolotl, or a fine-tuning script asks flash-attn to run on your card, the package checks your compute capability and raises:

RuntimeError: FlashAttention only supports Ampere GPUs or newer.

First, confirm which side of the line your card is on:

nvidia-smi --query-gpu=name,compute_cap --format=csv

Expected output on an affected card:

name, compute_cap
NVIDIA GeForce RTX 2080 Ti, 7.5

Anything below 8.0 lands in this article. Here’s the map:

ArchitectureCompute capabilityConsumer/homelab cardsflash-attn 2bf16Hardware FP8
Pascal (2016)6.1GTX 1060–1080 Ti, Tesla P40NoNoNo
Volta (2017)7.0V100, Titan VNoNoNo
Turing (2018)7.5GTX 16-series, RTX 2060–2080 Ti, Tesla T4NoNoNo
Ampere (2020)8.0 / 8.6RTX 3060–3090 Ti, A100YesYesNo
Ada (2022)8.9RTX 4060–4090YesYesYes (native)

Compute capability is fixed in silicon. No driver, CUDA toolkit, or PyTorch upgrade changes it — which is also why the related “no kernel image available” error keeps appearing on old cards as projects raise their minimum architecture.

Fix it in vLLM

vLLM’s hard floor is compute capability 7.0. Turing and Volta can run it with the right flags; Pascal (6.1) — including the beloved 24GB Tesla P40 — is below the floor and cannot run modern vLLM at all. P40 owners, skip straight to the llama.cpp section.

Step 1: force float16

On pre-Ampere cards you’ll usually hit the sibling error first, because most 2025–2026 models ship bf16 weights:

ValueError: Bfloat16 is only supported on GPUs with compute capability
of at least 8.0. Your Tesla T4 GPU has compute capability 7.5.
You can use float16 instead by explicitly setting the `dtype` flag in CLI,
for example: --dtype=half

Do what it says:

vllm serve Qwen/Qwen2.5-7B-Instruct --dtype=half

This came up as early as vLLM issue #1284 (Tesla P100) and still trips people on Quadro RTX 5000s today. One honest caveat from the V100 community: converting a model trained and saved in bf16 down to fp16 can shift outputs slightly. For chat and coding use you’re unlikely to notice; for evals, rerun your benchmarks.

Step 2: let the attention backend fall back

Modern vLLM detects pre-Ampere hardware and selects a non-FlashAttention backend (xformers) on its own — a T4 on vLLM 0.8.4+ logs that it’s running the older V0 engine path for exactly this reason. If your version still tries to load FlashAttention, pin the backend explicitly:

VLLM_ATTENTION_BACKEND=XFORMERS vllm serve <model> --dtype=half

xformers’ memory-efficient attention runs on compute capability 7.x and delivers most of FlashAttention’s benefit at single-user batch sizes.

Step 3: know the two walls that flags can’t fix

  • FP8-quantized models (anything tagged -FP8 or FP8-dynamic): there is no workaround on pre-Ampere or even Ampere hardware for weights that require hardware FP8 — the vLLM maintainers’ answer is to use a different quantization or a newer GPU. On Turing, pick AWQ or GPTQ versions instead; a 2025 Marlin kernel update for sm75 made AWQ models genuinely fast again on T4s and 2080 Tis.
  • Per-model architecture drops: newer multimodal stacks have started requiring Ampere outright — Qwen3-VL shipped without Turing support in late 2025. The model card, not vLLM, decides.

If vLLM won’t start for reasons beyond attention, our vLLM engine startup errors guide covers the other six common crashes.

Fix it in Hugging Face Transformers

If your script or a downloaded example sets attn_implementation="flash_attention_2", that explicit request is why it crashes: Transformers hard-errors when you name an implementation whose dependency can’t run, and only falls back silently when you leave it unset (that behavior was formalized in transformers PR #27940). The fix is one argument:

model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-7B-Instruct",
    torch_dtype=torch.float16,   # not bfloat16, same reason as vLLM
    attn_implementation="sdpa",  # was: "flash_attention_2"
    device_map="auto",
)

sdpa routes through PyTorch’s built-in scaled_dot_product_attention, which picks a memory-efficient kernel that runs on Pascal, Volta, and Turing with zero extra dependencies. It’s the right default for every pre-Ampere card.

Two edge cases:

  • Models that SDPA can’t serve correctly: a few architectures with attention soft-capping (Gemma 2 era) need attn_implementation="eager" instead. Slower, always correct.
  • Frameworks that hard-code flash_attention_2: if you can’t edit the call, there’s a third-party drop-in — flash-attention-legacy — that reimplements the FA2 interface for Pascal and Volta. It works per its README, but it’s an unaudited community project, not Dao-AILab code; prefer SDPA when you have the choice.

For fine-tuning on Turing, the same logic applies in Axolotl and friends: their docs state flash attention needs Ampere+, so set the attention option to SDPA and accept the longer step time. QLoRA itself runs fine on a 2080 Ti — it’s only the attention kernel that’s gated.

The escape hatch: llama.cpp, Ollama, and LM Studio

Here’s the part most threads bury: llama.cpp does not use the flash-attn package at all. It has its own flash-attention CUDA kernels, including a path that doesn’t require tensor cores, and users confirm -fa working on Turing hardware. Recent builds enable flash attention automatically when the device supports it — no flag needed.

./llama-server -m qwen2.5-7b-instruct-q4_k_m.gguf -ngl 99 -fa on

Pascal is shakier: self-compiled builds that omit arch 610 die with CUDA kernel flash_attn_ext_f16 has no device code compatible with CUDA arch 610, so P40 and GTX 10-series owners should use the official prebuilt binaries (which include Pascal) or compile with -DCMAKE_CUDA_ARCHITECTURES=61.

Ollama and LM Studio bundle llama.cpp internally, so they inherit all of this. The practical conclusion: if your goal is running models locally — chat, coding assistant, home server — rather than serving concurrent users with vLLM or training, the zero-effort fix for this entire error class is switching runtime. Our sister site’s vLLM review draws the serve-vs-local line in more detail; for a single user on a single old GPU, llama.cpp is almost always the right tool.

When the fix stops being software

Count the walls a Turing card has hit by October 2026: no flash-attn 2 (since 2023), no bf16, no FP8 models, no Qwen3-VL, V0-engine fallbacks in vLLM. Each one has a workaround today, and each workaround is a little slower and a little further off the beaten path. Pascal has it worse — it’s already below vLLM’s floor entirely.

If this error is what finally revealed that your GPU, not your software, is the bottleneck, the upgrade math in late 2026 looks like this: the cheapest Ampere entry is a used RTX 3060 12GB at $260–$296 (eBay sold listings, September 2026) — it clears every compatibility wall above, though 12GB limits you to ~13B Q4 models. The card this site actually recommends for the tier is the used RTX 3090 24GB at $1,399–$1,450 (median of eBay sold listings, October 2026 — up 15% in 11 days, the DRAM crisis reprices the used market too). Check what your target models need against either card with the VRAM calculator, and the GPU buying guide covers the full ladder. If you own a P40, our Tesla P40 reality check weighs how long the llama.cpp-only life stays viable.

FAQ

Can I just install FlashAttention 1.x on my Turing card like the README says? You can (pip install flash-attn==1.0.9), but almost nothing will use it — vLLM, Transformers, and every 2025+ training framework call the FA2 API, which 1.x doesn’t expose. SDPA is the supported answer.

Does this error mean my GPU can’t run local LLMs? No. It means one attention kernel package doesn’t support your GPU. A 2080 Ti runs 7B–13B Q4 models comfortably in llama.cpp or Ollama; even a GTX 1080 runs 7B models. The error only gates specific serving/training stacks.

How much slower is xformers/SDPA than FlashAttention? At home-lab batch sizes (one or two concurrent requests), the gap is modest — memory-efficient attention closes most of it, and decode speed is bound by memory bandwidth anyway. FlashAttention’s big wins are long contexts and high batch counts, which is why datacenter users care more than you need to.

I have an RTX 50-series card and I’m getting a FlashAttention architecture error too — same problem? Opposite problem, same symptom. Early Blackwell (sm_120) adopters hit “no supported architecture” failures because flash-attn wheels hadn’t caught up (flash-attention issue #1987). That one is fixed by upgrading flash-attn and PyTorch, not by fallback flags — see our no-kernel-image fix.

What about AMD cards? Different package (the ROCm flash-attention fork), different error strings, different arch gates (RDNA3/CDNA2+). Start with our ROCm setup guide.

Sources

Last updated October 10, 2026. Prices and software support change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.