Skip to main content
CompareLLM Technical Deep-Dive
Updated 2026-08-19

The Complete Guide to Local AI Hardware, VRAM Sizing & Quantization in 2026

Everything you need to know about running open-weight LLMs locally: GPU VRAM vs Apple Unified RAM, Quantization (FP16 down to Q2), KV Cache scaling, Layer Offloading, and Inference Speed (tok/s).

Bandwidth Formula

Speed ≈ Bandwidth / Weights

1,000 GB/s bandwidth streaming 25 GB weights yields ~40 tok/s.

Graphics Memory (VRAM) Needed

Params × B/Param + KV Cache

70B in Q4 needs 43.4 GB weights + 4–8 GB KV cache ≈ 48–52 GB.

Best Compression Level

Q4_K_M (97.4% Accuracy)

Standard 4-bit cuts memory by 68% with virtually imperceptible loss.

Interactive SizerLaunch Hardware Sizer

Test any GPU, Mac Mini, or cloud cluster with live KV cache math.

1

Memory Architectures: Dedicated VRAM vs Apple Unified vs Host RAM

During auto-regressive generation, modern LLMs must stream all parameter weights through memory for every single token produced. Therefore, memory bandwidth—not raw compute TFLOPS—is the primary bottleneck for text streaming speed.

Dedicated GPU VRAM1,000–3,800 GB/s
Hardware:NVIDIA RTX 4090 / 5090, A100, H100
Streaming Speed:30–140+ tok/s
Max Single GPU:24 GB (Consumer) / 80–192 GB (Data Center)

Best for real-time coding assistants, sub-second TTFT, and high-concurrency production deployments.

Apple Unified Memory150–800 GB/s
Hardware:Mac Mini M4 Pro, Mac Studio M2/M3 Ultra
Streaming Speed:15–45 tok/s
Max Capacity:64 GB to 192 GB Unified RAM

Best value for 70B & 235B models on a single workstation. Low power draw (35–90W) with zero multi-GPU setup hassle.

System Host RAM (CPU)50–100 GB/s
Hardware:Standard PC DDR4 / DDR5 Dual-Channel
Streaming Speed:1–5 tok/s
Max Capacity:64 GB to 256 GB+

Usable for offline batch jobs, testing layer offload overflow, but too slow for interactive chat or real-time coding copilot loops.

2

Quantization Precision Matrix: FP16 down to Q2

Quantization compresses the model weights from 16-bit floating point representations down to integer representations (8-bit, 5-bit, 4-bit, 3-bit, 2-bit), drastically reducing VRAM size and boosting memory bandwidth throughput.

Format / QuantBytes / ParamAccuracy Retention70B VRAM FootprintTarget Use Case
FP16 / BF16 (Reference)2.00 B100.0% (Baseline)~140 GBLab research & fine-tuning source
Q8_0 (8-bit)1.15 B99.7%~80 GBEnterprise mission-critical zero degradation
Q5_K_M (5-bit)0.78 B98.9%~54 GBSweet spot for complex coding & multi-step logic
Q4_K_M (4-bit)STANDARD0.62 B97.4%~43.4 GBGlobal consumer standard (Ollama / vLLM / llama.cpp)
Q3_K_M (3-bit)0.49 B92.5%~34 GBEnables 70B on 32GB GPUs or 235B MoE on 128GB
Q2_K (2-bit)0.35 B88.0%~25 GBExtreme compression for 671B MoE exploration
3

VRAM Sizing Cheat Sheet by Model Class

Live Sizer

Quick lookup table for exact weight sizing across popular open-weight model parameter classes, with 4k–8k KV cache headroom included:

8B Class (Llama 3.1, DeepSeek 8B)8 Billion
Q4_K_M (Min VRAM):~5.5 GB (8GB GPU OK)
Q8_0:~9.5 GB (12GB GPU)
FP16 Reference:~16.5 GB (24GB GPU)

Runs effortlessly on RTX 3060 12GB, RTX 4060, or standard M4 Mac Mini (16GB).

14B–27B (Qwen 2.5 14B, Qwen 3.8 27B)14B–27B
Q4_K_M (Min VRAM):~10–18 GB (16GB–24GB GPU)
Q8_0:~18–32 GB
FP16 Reference:~30–56 GB

Sweet spot for single RTX 4090 24GB or Apple Mac Mini 32GB/64GB.

32B Class (Qwen 2.5 32B, QwQ 32B)32 Billion
Q4_K_M (Min VRAM):~20.5 GB (Fits 24GB VRAM)
Q8_0:~37 GB
FP16 Reference:~65 GB

The flagship single-consumer GPU tier. Maximum reasoning density on a single RTX 4090.

70B Class (Llama 3.3 70B, Qwen 72B)70 Billion
Q4_K_M (Min VRAM):~43.4 GB (Dual 24GB or 64GB Mac)
Q8_0:~80 GB
FP16 Reference:~140 GB

Standard for local workstation AI. Requires Dual RTX 3090/4090 or Mac Mini M4 Pro (64GB).

235B MoE (Qwen 3 235B, 21B Active)235B MoE
Q4_K_M (Min VRAM):~135 GB (Mac Studio 192GB)
Q8_0:~250 GB
FP16 Reference:~470 GB

High speed due to sparse 21B routing, but requires large unified memory pool to store weights.

671B MoE (DeepSeek R1 / V3)671B MoE
Q4_K_M (Min VRAM):~385 GB (8× 80GB DGX or Cluster)
Q2_K (Extreme):~220 GB
FP16 Reference:~1.34 TB

Full frontier open-weight flagship. Deployed on enterprise 8× H100/H200 nodes or multi-Mac clusters.

4

KV Cache Memory Formula & Context Scaling

Model weights are only the base memory requirement. As conversation context expands, the attention layers store past tokens in the Key-Value (KV) Cache.

Mathematical FormulaKV_Bytes = 2 × n_layers × n_kv_heads × head_dim × n_ctx × bytes_per_element × batch_size

Thanks to GQA (Grouped Query Attention), modern models reduce `n_kv_heads` by 4x to 8x compared to legacy Multi-Head Attention, cutting KV cache size by ~75%.

8,192 Context (Standard)

~0.6–1.2 GB VRAM

Fits easily in standard GPU buffer

32,768 Context (Long Documents)

~2.5–5.0 GB VRAM

Requires allocating dedicated GPU headroom

131,072 Context (Full Codebase)

~10–22 GB VRAM (FP16)

Use FP8/Q4 KV cache quantization (`--ctk q4_0`)

5

CPU/GPU Layer Offloading (`-ngl` in llama.cpp)

When a quantized model exceeds your GPU's dedicated VRAM, layer offloading enables you to load a portion of transformer layers directly into GPU VRAM while streaming the remainder from system RAM.

llama.cpp Partial GPU Offload
./llama-server -m models/llama-3.3-70b-q4_k_m.gguf -ngl 35 -c 8192

`-ngl 35` offloads 35 of 80 layers into a 16GB GPU VRAM, keeping remaining 45 layers in System RAM.

6

Hardware Buyer's Guide & Cloud Break-Even Economics

Tier 1: Solo Developer / Entry$499 – $1,299

Hardware: Apple Mac Mini M4 (24GB/32GB) or NVIDIA RTX 4070 Ti Super 16GB.

Capabilities: 8B to 14B models at 40–80 tok/s. Fast coding autocompletion with Qwen 2.5 Coder 14B.

Tier 2: 70B Workhorse (Sweet Spot)$1,599 – $3,200

Hardware: Mac Mini M4 Pro (64GB) or Dual RTX 3090 / 4090 48GB Rig.

Capabilities: Llama 3.3 70B & Qwen 2.5 72B Q4_K_M running at 18–35 tok/s natively with zero offload penalties.

When Does Local Hardware Pay for Itself?

If your team spends >$150/month on cloud API token billing (approx. 50M tokens/month), a dedicated 64GB workstation amortizes its cost within 9 to 14 months while providing complete offline data privacy.

Launch Hardware Sizer

Ready to evaluate your stack?

Calculate your optimal model weights with Stack Engine or compare top models head-to-head.

Explore More AI Intelligence Tools

Back to Homepage Overview