Skip to content

Best GPU for gpt-oss-120b Locally

Real-time prices and hardware recommendations updated for August 2026.

117B
5.1B
moe

gpt-oss-120b is a mixture-of-experts model: all 117B parameters must sit in VRAM, but only 5.1B activate per token — so it generates far faster than a dense model of the same size, while still demanding the memory of one.

To run gpt-oss-120b locally you need roughly 67.9 GB of VRAM at Q4_K_M quantization with a 32k token context. The best-value card that fits is the RTX PRO 6000 Blackwell (96 GB), which should generate around 190 tokens per second.

Adjust Context Length

Slide to popular sizes or type any value manually.

k tokens
8k
16k
32k
64k
128k
256k
Budget EntryQ3_K_M quant
50.5 GB
0.7 GB
Total VRAM:51.1 GB

Recommended Hardware

RTX PRO 6000 Blackwell
5392 tok/sprefill
230 tok/sgeneration
EUR 19916.01·96GB VRAM
View Card

Q3 weights are ~50 GB — past every consumer card, so the 96 GB RTX PRO 6000 Blackwell is the single-GPU entry point. With only 5.1B params active per token, llama.cpp expert-offload (--n-cpu-moe) is the cheaper route: keep the hot path on a 24-48 GB card and page the experts from system RAM.

Balanced Sweet SpotQ4_K_M quant
67.3 GB
0.7 GB
Total VRAM:67.9 GB

Recommended Hardware

RTX PRO 6000 Blackwell
5392 tok/sprefill
190 tok/sgeneration
EUR 19916.01·96GB VRAM
View Card

Q4 weights are ~67 GB, close to the model's native MXFP4 (~63 GB) format. A single 96 GB RTX PRO 6000 Blackwell (or an 80 GB data-center card) runs it at full quality and context; only 5.1B active parameters keep decode fast despite the 117B total.

Near LosslessQ8_0 quant
134.5 GB
0.7 GB
Total VRAM:135.2 GB

Recommended Hardware

No fitting GPUs found in database. Browse all GPUs

Q8 needs ~135 GB — a multi-GPU or unified-memory setup. gpt-oss ships in MXFP4 and is already near-lossless at that ~4-bit format, so Q8 is rarely worth the extra hardware.

Other models with the same VRAM requirement

Because gpt-oss-120b's weights fit a 96 GB card, other models of a similar size run on the same GPU. Generation speed varies — mixture-of-experts models are faster, dense models slower — but any of these load in the same VRAM:

Qwen3.5-122B-A10B122BDBRX132BMixtral 8x22B141B

Optimizing Setup for gpt-oss-120b

Quantization Recommendations

For daily coding and reasoning tasks, Q4_K_M (4-bit quantization) offers the best balance of quality and memory efficiency — it reduces memory requirements by over 70% with minimal quality loss compared to FP16. Q8 and higher presets preserve more fidelity at the cost of significantly higher VRAM usage, which may force layer offloading and hurt throughput.

Recommended Local Software

We recommend using Ollama as the primary runner for local inference due to its automated GPU model splitting and context cache optimizations. For advanced fine-tuning or quantization splits, llama.cpp with Flash Attention compiled natively provides the best granular control.

Running gpt-oss-120b locally — FAQ

How much VRAM do I need to run gpt-oss-120b?

At a 32k context with KV cache quantization on, gpt-oss-120b needs about 67.9 GB of VRAM at Q4_K_M — the quantization most people should use. Dropping to Q3_K_M brings that down to roughly 51.1 GB at some quality cost, while Q8_0 needs about 135.2 GB for the best quality this model can give.

What size graphics card does gpt-oss-120b fit on?

gpt-oss-120b needs about 67.9 GB at Q4_K_M, which is more than a single 24 GB consumer card provides. You need a workstation card, a multi-GPU setup, or a more aggressive quantization — otherwise layers spill into system RAM and generation slows dramatically.

Which quantization should I use for gpt-oss-120b?

Use Q4_K_M unless you have VRAM to spare. It needs about 67.9 GB and loses very little quality against full precision. Q8_0 needs about 135.2 GB for a quality gain most people cannot detect in everyday coding and chat. Spend spare VRAM on a longer context instead.

How does context length affect the VRAM gpt-oss-120b needs?

Model weights are fixed, but the KV cache grows linearly with context. For gpt-oss-120b at a 32k context the cache is about 0.7 GB; doubling to 64k takes it to roughly 1.3 GB. Turning KV cache quantization off doubles those figures again.

Why is gpt-oss-120b faster than its parameter count suggests?

gpt-oss-120b is a mixture-of-experts model. Its 117B parameters all have to be held in VRAM, but only 5.1B are used to produce each token. Generation speed is bound by streaming those 5.1B active parameters, so it feels much closer to a 5.1B model than a 117B one — while still needing memory for the full 117B.

Can the same GPU run other models similar to gpt-oss-120b?

Yes. gpt-oss-120b needs about 67.9 GB at Q4_K_M, and any model whose weights fit the same card runs on it — including Qwen3.5-122B-A10B (122B), DBRX (132B), Mixtral 8x22B (141B). The weights fit the same GPU; generation speed varies (mixture-of-experts models are faster, dense models slower).

How token speeds are estimated

Two metrics are shown per GPU: Read tok/s (how fast the model ingests your prompt) and Decode tok/s (how fast it streams tokens back). They model fundamentally different bottlenecks.

📖

Read (Prefill)

The prompt is processed in one parallel pass. This is compute-bound: it saturates the GPU's tensor cores.

read tok/s ≈ TFLOPS × readFactor × 400 ÷ activeParams

Decode (Generation)

Each new token requires loading the entire model's active weights from VRAM. This is memory-bandwidth-bound: the GPU stalls waiting for data, not computing.

decode tok/s ≈ bandwidth × decodeFactor ÷ (weights + kv_cache)

Weights = (activeParams × bits ÷ 8) × 1.15 overhead. KV cache per step = activeParams × multiplier × contextK.

Architecture utilization factors

Architecture
Decode
Read
Blackwell, Xe2
0.45
0.55
Ada Lovelace, RDNA 4, Battlemage
0.38
0.48
Ampere, Turing, RDNA 3, Xe-HPG
0.28
0.38
Volta, RDNA 1/2
0.2
0.25
Pre-tensor-core (Pascal, Maxwell, Kepler, GCN, Alchemist)
0.12
0.15

Left: decode factor — Right: read factor

Data sources

TFLOPS and memory bandwidth are read from the GPU database. When missing, bandwidth falls back to a hardcoded dictionary.

Limitations

These are analytical estimates, not benchmark results. Use them as a relative comparison, not an absolute performance guarantee.