The Complete Guide to Local AI Hardware, VRAM Sizing & Quantization in 2026
Everything you need to know about running open-weight LLMs locally: GPU VRAM vs Apple Unified RAM, Quantization (FP16 down to Q2), KV Cache scaling, Layer Offloading, and Inference Speed (tok/s).
Speed ≈ Bandwidth / Weights
1,000 GB/s bandwidth streaming 25 GB weights yields ~40 tok/s.
Params × B/Param + KV Cache
70B in Q4 needs 43.4 GB weights + 4–8 GB KV cache ≈ 48–52 GB.
Q4_K_M (97.4% Accuracy)
Standard 4-bit cuts memory by 68% with virtually imperceptible loss.
Test any GPU, Mac Mini, or cloud cluster with live KV cache math.
Memory Architectures: Dedicated VRAM vs Apple Unified vs Host RAM
During auto-regressive generation, modern LLMs must stream all parameter weights through memory for every single token produced. Therefore, memory bandwidth—not raw compute TFLOPS—is the primary bottleneck for text streaming speed.
Best for real-time coding assistants, sub-second TTFT, and high-concurrency production deployments.
Best value for 70B & 235B models on a single workstation. Low power draw (35–90W) with zero multi-GPU setup hassle.
Usable for offline batch jobs, testing layer offload overflow, but too slow for interactive chat or real-time coding copilot loops.
Quantization Precision Matrix: FP16 down to Q2
Quantization compresses the model weights from 16-bit floating point representations down to integer representations (8-bit, 5-bit, 4-bit, 3-bit, 2-bit), drastically reducing VRAM size and boosting memory bandwidth throughput.
| Format / Quant | Bytes / Param | Accuracy Retention | 70B VRAM Footprint | Target Use Case |
|---|---|---|---|---|
| FP16 / BF16 (Reference) | 2.00 B | 100.0% (Baseline) | ~140 GB | Lab research & fine-tuning source |
| Q8_0 (8-bit) | 1.15 B | 99.7% | ~80 GB | Enterprise mission-critical zero degradation |
| Q5_K_M (5-bit) | 0.78 B | 98.9% | ~54 GB | Sweet spot for complex coding & multi-step logic |
| Q4_K_M (4-bit)STANDARD | 0.62 B | 97.4% | ~43.4 GB | Global consumer standard (Ollama / vLLM / llama.cpp) |
| Q3_K_M (3-bit) | 0.49 B | 92.5% | ~34 GB | Enables 70B on 32GB GPUs or 235B MoE on 128GB |
| Q2_K (2-bit) | 0.35 B | 88.0% | ~25 GB | Extreme compression for 671B MoE exploration |
VRAM Sizing Cheat Sheet by Model Class
Quick lookup table for exact weight sizing across popular open-weight model parameter classes, with 4k–8k KV cache headroom included:
Runs effortlessly on RTX 3060 12GB, RTX 4060, or standard M4 Mac Mini (16GB).
Sweet spot for single RTX 4090 24GB or Apple Mac Mini 32GB/64GB.
The flagship single-consumer GPU tier. Maximum reasoning density on a single RTX 4090.
Standard for local workstation AI. Requires Dual RTX 3090/4090 or Mac Mini M4 Pro (64GB).
High speed due to sparse 21B routing, but requires large unified memory pool to store weights.
Full frontier open-weight flagship. Deployed on enterprise 8× H100/H200 nodes or multi-Mac clusters.
KV Cache Memory Formula & Context Scaling
Model weights are only the base memory requirement. As conversation context expands, the attention layers store past tokens in the Key-Value (KV) Cache.
KV_Bytes = 2 × n_layers × n_kv_heads × head_dim × n_ctx × bytes_per_element × batch_sizeThanks to GQA (Grouped Query Attention), modern models reduce `n_kv_heads` by 4x to 8x compared to legacy Multi-Head Attention, cutting KV cache size by ~75%.
~0.6–1.2 GB VRAM
Fits easily in standard GPU buffer
~2.5–5.0 GB VRAM
Requires allocating dedicated GPU headroom
~10–22 GB VRAM (FP16)
Use FP8/Q4 KV cache quantization (`--ctk q4_0`)
CPU/GPU Layer Offloading (`-ngl` in llama.cpp)
When a quantized model exceeds your GPU's dedicated VRAM, layer offloading enables you to load a portion of transformer layers directly into GPU VRAM while streaming the remainder from system RAM.
./llama-server -m models/llama-3.3-70b-q4_k_m.gguf -ngl 35 -c 8192`-ngl 35` offloads 35 of 80 layers into a 16GB GPU VRAM, keeping remaining 45 layers in System RAM.
Hardware Buyer's Guide & Cloud Break-Even Economics
Hardware: Apple Mac Mini M4 (24GB/32GB) or NVIDIA RTX 4070 Ti Super 16GB.
Capabilities: 8B to 14B models at 40–80 tok/s. Fast coding autocompletion with Qwen 2.5 Coder 14B.
Hardware: Mac Mini M4 Pro (64GB) or Dual RTX 3090 / 4090 48GB Rig.
Capabilities: Llama 3.3 70B & Qwen 2.5 72B Q4_K_M running at 18–35 tok/s natively with zero offload penalties.
If your team spends >$150/month on cloud API token billing (approx. 50M tokens/month), a dedicated 64GB workstation amortizes its cost within 9 to 14 months while providing complete offline data privacy.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
Related guides
Same topic, next level of detail.
Open-weight vs closed API models
When to self-host or buy open weights versus calling a frontier API. Catalog flags and trade-offs.
Reasoning Effort, Dynamic CoT, and Multi-Tier Model Derivatives in 2026
How test-time compute, reasoning effort controls, fast routes, and derivative tiers (Sol/Terra/Luna, Opus/Fable/Sonnet, DeepSeek V4, Qwen 3) affect benchmarks and production bills.
Explore More AI Intelligence Tools
Pairwise Model Comparisons
Compare any two models on our category ratings, speed classes and token pricing.
Curated 'Best-Of' Indexes
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
AI Hardware Calculator
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live Market Intelligence
Live audit trail of benchmark updates, new model releases, and API price cuts.
Interactive Model Finder
Answer 5 quick questions to compute deterministic model recommendations for your use case.
Methodology & Standards
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.
