VRAM CalculatorCan I Run
🔴 No — OOM Risk

DeepSeek V3 685B (MoE)onRTX 4090

Affiliate disclosure: We earn commissions when you shop through the links below at no additional cost to you.

MHA model · 8 attention heads · 8k context · llama.cpp

355.2 GiB
24 GiB available

estimated

Model
DeepSeek V3 685B (MoE)
671B params · MoE · 61 layers
MoE Architecture
GPU
RTX 4090
24GB VRAM · NVIDIA
Enthusiast GPU

VRAM by Quantization

Calculated estimate using: Weights + KV Cache (GQA-aware) + Activations + Engine Overhead. These are theoretical calculations — actual VRAM usage varies by runtime and batch size.

FP16 (Full precision)2 bpw
1253.5 GiB— OOM
Q8_0 (High quality)1 bpw
628.6 GiB— OOM
Q5_K_M (Balanced)0.6875 bpw
433.3 GiB— OOM
Q4_K_M (Efficient)0.5625 bpw
355.2 GiB— OOM
IQ4_XS (Ultra efficient)0.53 bpw
334.9 GiB— OOM

KV cache always stored in FP16 (2 bytes). GQA reduces KV cache size by 1.0× vs standard MHA.

Confidence

Medium

Standard inference workload estimated with llama.cpp engine overhead at 0.8GB. Variance depends on batch size, exact context used, and runtime implementation. This is a calculated estimate, not a measured runtime value.

Rent a Cloud GPU

Need 331.2 GiB more. Don't compromise your model with heavier quantization.

What can I run instead?

DeepSeek V3 685B (MoE) Q4_K_M doesn't fit on RTX 4090 at 8k context. Here are viable alternatives:

Rent a cloud GPUrun DeepSeek V3 685B (MoE) Q4_K_M at full contextVast.ai →

Affiliate Disclosure

Some links on this page are affiliate links. If you purchase through them, we may earn a small commission at no additional cost to you. We only recommend tools we've thoroughly researched and believe add real value.

Our reviews and comparisons are based on objective analysis and are not influenced by affiliate partnerships.

Other combinations: