Check whether a model fits your GPU — and how fast it will run.
Test any open-source AI model against NVIDIA DGX nodes, Apple Mac Mini / Studio, AMD Instinct, and consumer GPUs to calculate memory feasibility, quantization tiers, KV cache footprint, generation speed (tok/s), and deployment commands.
New to local hosting? Read the full VRAM & quantization guideSelect an open-source model and target machine to check runnable fit, modeled generation speed (tok/s), memory footprint, and launch-command templates.
Z.ai · Coding Reported rating 1552
30.2 GB Usable · 1792 GB/s Bandwidth
⚠️ Model requires 468.4 GB VRAM, but RTX 5090 (32 GB) provides 30.2 GB usable (438.2 GB deficit). Partially offloads to system RAM at ~2–4 tok/s. Scale to a multi-GPU cluster or apply 4-bit quantization for full GPU throughput.
To run GLM-5.3 (468.4 GB VRAM required), deploy a cluster of 16× NVIDIA GeForce RTX 5090 (32 GB GDDR7) with Tensor Parallelism TP=8 and Pipeline Parallelism PP=2 across 16 GPUs (forming a 512 GB pool with 483.2 GB usable after KV cache & runtime buffers).
Memory values are modeled estimates. Weights use the selected quantization; KV cache uses context, precision, and batch size; runtime overhead is a conservative workspace estimate, not measured CUDA telemetry.
ollama run glm-5.3:q4k_mVerify the model tag, file name, runtime support, and license before running. Availability differs by registry and model family.
Estimated local power cost is ~$0.10/hr. Using the catalog API output price of $4.40/1M tokens, estimated break-even is ~1,200,000 tokens/day.
Dedicated ultra-fast memory on your GPU where AI model weights live during inference.
Compacts model numbers from 16-bit to 4-bit or 5-bit to fit smaller GPUs.
Working memory used to remember conversation history and long documents.
The generation streaming speed of the AI model.
Splits model layers between fast GPU VRAM and slower System Host RAM.
Combines 2, 4, or 8 identical GPUs into a single unified virtual VRAM pool.
Calculates pure model weight footprint across FP16, Q8_0, Q5_K_M, and Q4_K_M plus real-world CUDA context buffers and activation overhead.
Models Grouped Query Attention (GQA) and Multi-Head Attention memory requirements scaling from 2k to 1M token contexts.
Calculates electricity draw, hardware amortization, and cloud rental costs to determine the exact break-even token threshold against commercial APIs.
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.