DeepSeek V3 685B (MoE)onA100 80GB
Affiliate disclosure: We earn commissions when you shop through the links below at no additional cost to you.
MHA model · 8 attention heads · 8k context · llama.cpp
estimated
VRAM by Quantization
Calculated estimate using: Weights + KV Cache (GQA-aware) + Activations + Engine Overhead. These are theoretical calculations — actual VRAM usage varies by runtime and batch size.
KV cache always stored in FP16 (2 bytes). GQA reduces KV cache size by 1.0× vs standard MHA.
Confidence
MediumStandard inference workload estimated with llama.cpp engine overhead at 0.8GB. Variance depends on batch size, exact context used, and runtime implementation. This is a calculated estimate, not a measured runtime value.
Need 275.2 GiB more. Don't compromise your model with heavier quantization.
What can I run instead?
DeepSeek V3 685B (MoE) Q4_K_M doesn't fit on A100 80GB at 8k context. Here are viable alternatives:
Affiliate Disclosure
Some links on this page are affiliate links. If you purchase through them, we may earn a small commission at no additional cost to you. We only recommend tools we've thoroughly researched and believe add real value.
Our reviews and comparisons are based on objective analysis and are not influenced by affiliate partnerships.