What LLM can an RTX 5080 run?
At the default reserve, the generic one-card floor reaches a 20B INT4, 8B INT8, or 3B FP16 tier. Named models still need exact checkpoint, context, cache, and runtime checks.
Hardware guide / 16 GB VRAM
A 16 GB card reaches the generic 20B INT4 weight tier. Exact Qwen3 14B and short-context gpt-oss-20b profiles land so close to the budget that runtime allocation decides the result.
Open the 16 GB model finder →Fast answer
INT420B listed tier · 25.77B arithmetic ceiling
INT88B listed tier · 12.88B arithmetic ceiling
FP16 / BF163B listed tier · 6.44B arithmetic ceiling
Default reserve16 × 0.90 ÷ 1.20 = 12 GiB for weights
Editable evidence
Every model card opens an editable architecture-aware calculator rather than hiding the assumptions.
Decision boundary
Two named profiles sit within 0.13 GiB of the default budget, so graph capture, allocator fragmentation, or runtime workspace can reverse the result.
| Profile | Planning target | 16 GB result |
|---|---|---|
| Qwen3 8B INT4 · 32K | 9.98 GiB | Fits with 4.42 GiB margin |
| Qwen3 14B INT4 · 32K | 14.27 GiB | Paper fit; only 0.13 GiB margin |
| gpt-oss-20b · 4K | 14.29 GiB | Paper fit; only 0.11 GiB margin |
| Qwen3 8B INT8 · 32K | 14.56 GiB | Misses by 0.16 GiB |
Capacity is only gate one
0.11 GiB is not operational headroom.
Treat the gpt-oss-20b 4K result as a paper boundary, not a promise.
Context consumes VRAM.
The conservative gpt-oss 128K profile rises to 22.56 GiB.
Fit is not throughput.
Bandwidth, kernels, offload, and batching decide delivered speed.
Primary sources
NVIDIA specifies 16 GB of GDDR7 memory for the RTX 5080. Model profiles reuse the primary-source architecture and checkpoint inputs documented in the Qwen3 and gpt-oss worksheets.
Official NVIDIA RTX 5080 specifications → Open the Qwen3 GPU worksheet → Open the gpt-oss hardware worksheet →At the default reserve, the generic one-card floor reaches a 20B INT4, 8B INT8, or 3B FP16 tier. Named models still need exact checkpoint, context, cache, and runtime checks.
The 32K INT4 planning profile reaches 14.27 GiB, about 0.13 GiB below the 14.4 GiB budget. Treat that as a paper fit requiring runtime validation.
The 4K profile reaches 14.29 GiB and clears the budget by only 0.11 GiB. The conservative 128K profile reaches 22.56 GiB and does not fit fully on the card.
Not under this fully GPU-resident 32K INT4 profile. Its 27.93 GiB planning target exceeds the one-card allowance.
Hardware claims, checked
ResearchAudio traces exact checkpoints, architecture, context, runtime constraints, and the gap between “fits” and “works.”
Prefer the hosted signup page?Join ResearchAudio free →