GPUzV2 · Experimental

Recipe Planner

Choose your GPU and the context size you need. GPUz compares verified Ollama model recipes and shows which ones are estimated below or above your GPU's nominal VRAM.

How does this work?

• GPU memory comes from source-verified NVIDIA specifications.

• Model recipes use exact curated GGUF artifacts from verified sources.

• Memory estimates come from source-backed, versioned estimators selected for each model architecture.

• Some standard models use the versioned Ollama estimator; hybrid architectures may use a separately validated analytical model.

• The result is an estimate, not proof that a model will run.

GPU
NVIDIA GeForce RTX 3090
24 GB GDDR6X
Source: NVIDIA
Requested context
2,048 tokens
NVIDIA GeForce RTX 3090: 24 GB GDDR6X.
Memory estimate covers model weights and modeled context memory. Runtime compute/graph overhead is not included, so this is an analytical estimate — not a guarantee of device residency.

Estimated below nominal capacity

Mistral Small 22B
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Related evidence available
· Curated observation
Independence of sources not established
Estimated memory
12.9 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~9.5 GiB before runtime overhead.
Gemma 2 27B
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
16.3 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~6.0 GiB before runtime overhead.
Qwen3 8B
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
5.1 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~17.2 GiB before runtime overhead.
Qwen3 14B
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
9.0 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~13.4 GiB before runtime overhead.
Qwen3 32B
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
19.1 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~3.2 GiB before runtime overhead.
Qwen3 30B (A3B)
q4_K_M · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
17.4 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~5.0 GiB before runtime overhead.
Qwen3.6 35B A3B (UD-IQ4_NL)
UD-IQ4_NL · Runtime: Ollama
Estimated to fit
No direct evidence yet
Estimated memory
16.9 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Nominal headroom: ~5.4 GiB before runtime overhead.

Estimated above nominal capacity

Llama 3.3 70B
q4_K_M · Runtime: Ollama
Estimated over capacity
No direct evidence yet
Estimated memory
40.2 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Other options that fit this GPU
Mistral Small 22B · q4_K_M · ~12.9 GiB
Gemma 2 27B · q4_K_M · ~16.3 GiB
Qwen3 8B · q4_K_M · ~5.1 GiB
Qwen3 14B · q4_K_M · ~9.0 GiB
Qwen3 32B · q4_K_M · ~19.1 GiB
Qwen3 30B (A3B) · q4_K_M · ~17.4 GiB
Qwen3.6 35B A3B (UD-IQ4_NL) · UD-IQ4_NL · ~16.9 GiB
Can't run it locally? Run on a cloud GPU:

Affiliate links — we may earn a commission at no extra cost to you.

DeepSeek V3 671B
q4_K_M · Runtime: Ollama
Estimated over capacity
No direct evidence yet
Estimated memory
380.0 GiB
GPU capacity
22.4 GiB
(24 GB advertised)
Other options that fit this GPU
Mistral Small 22B · q4_K_M · ~12.9 GiB
Gemma 2 27B · q4_K_M · ~16.3 GiB
Qwen3 8B · q4_K_M · ~5.1 GiB
Qwen3 14B · q4_K_M · ~9.0 GiB
Qwen3 32B · q4_K_M · ~19.1 GiB
Qwen3 30B (A3B) · q4_K_M · ~17.4 GiB
Qwen3.6 35B A3B (UD-IQ4_NL) · UD-IQ4_NL · ~16.9 GiB
Can't run it locally? Run on a cloud GPU:

Affiliate links — we may earn a commission at no extra cost to you.