What LLM can an RTX 4060 8GB run?
At the default reserve, the generic one-card floor reaches an 8B INT4, 3B INT8, or 3B FP16 tier. Named models still need artifact, context, cache, and runtime inputs.
Hardware guide / 8 GB VRAM
Eight gigabytes is a context-sensitive 8B INT4 tier. Qwen3 8B clears this planning model at 4K and 8K, then the full 16K profile crosses the one-card boundary.
Open the 8 GB model finder →Fast answer
INT48B listed tier · 12.88B arithmetic ceiling
INT83B listed tier · 6.44B arithmetic ceiling
FP16 / BF163B listed tier · 3.22B arithmetic ceiling
Default reserve8 × 0.90 ÷ 1.20 = 6 GiB for weights
Editable evidence
The same model and precision can fit or miss depending on context. Each card opens the calculator with the architecture, reserve, and headroom already filled in.
Decision boundary
The four Qwen3 profiles hold weights and runtime headroom constant. Only the BF16 KV-cache context changes.
| Profile | Planning target | 8GB result |
|---|---|---|
| Qwen3 8B INT4 · 4K | 5.26 GiB | Fits with 1.94 GiB margin |
| Qwen3 8B INT4 · 8K | 5.93 GiB | Fits with 1.27 GiB margin |
| Qwen3 8B INT4 · 16K | 7.28 GiB | Narrow paper miss |
| Qwen3 8B INT4 · 32K | 9.98 GiB | Does not fit fully on GPU |
Capacity is only gate one
Context is a budget.
In this profile the BF16 KV cache grows from 0.56 GiB at 4K to 4.5 GiB at 32K.
Small changes move a tight fit.
KV-cache quantization or a smaller reserve can change the 16K result, but must be benchmarked.
Fit is not throughput.
Bandwidth, kernels, laptop power limits, offload, and batching decide delivered speed.
Primary sources
NVIDIA lists the RTX 4060 with 8GB GDDR6. The Qwen3 profiles reuse the architecture inputs and primary references documented in the ResearchAudio Qwen3 worksheet; this page adds explicit 4K-to-32K comparisons for the 8GB card.
Official NVIDIA RTX 4060 family specifications → Open the Qwen3 GPU worksheet → Compare the 12GB RTX 3060 boundary →At the default reserve, the generic one-card floor reaches an 8B INT4, 3B INT8, or 3B FP16 tier. Named models still need artifact, context, cache, and runtime inputs.
The INT4 planning profile reaches 5.26 GiB at 4K and 5.93 GiB at 8K. Both clear the 7.2 GiB budget, but the exact artifact and runtime still need validation.
Not under this conservative fully GPU-resident plan. The target reaches 9.98 GiB with a BF16 KV cache and 20 percent headroom.
The generic 13B INT4 weight-only target reaches 7.26 GiB, narrowly above this budget before context. Lower-bit artifacts or offload can move the line, with quality or speed tradeoffs.
Hardware claims, checked
ResearchAudio traces exact checkpoints, architecture, context, runtime constraints, and the gap between “fits” and “works.”
Prefer the hosted signup page?Join ResearchAudio free →