Check whether a model fits your GPU — and how fast it will run.
Test any open-source AI model against NVIDIA DGX nodes, Apple Mac Mini / Studio, AMD Instinct, and consumer GPUs to calculate memory feasibility, quantization tiers, KV cache footprint, generation speed (tok/s), and deployment commands.
New to local hosting? Read the full VRAM & quantization guideSelect an open-source model and target machine to check runnable fit, modeled generation speed (tok/s), memory footprint, and launch-command templates.
IBM · Coding Estimated rating 1480
30.2 GB Usable · 1792 GB/s Bandwidth
✓ Fits 100% in VRAM on RTX 5090 (32 GB) at ~58–86 tok/s using FP16 (100% accuracy) with 12.7 GB free headroom.
Granite 4.2 8B fits within a single NVIDIA GeForce RTX 5090 (32 GB GDDR7) (6.5 GB / 30.2 GB usable) without requiring multi-unit clustering.
Memory values are modeled estimates. Weights use the selected quantization; KV cache uses context, precision, and batch size; runtime overhead is a conservative workspace estimate, not measured CUDA telemetry.
ollama run granite-4.2-8b:q4k_mVerify the model tag, file name, runtime support, and license before running. Availability differs by registry and model family.
Estimated local power cost is ~$0.10/hr. Using the catalog API output price of $0.15/1M tokens, estimated break-even is ~35,200,000 tokens/day.
Dedicated ultra-fast memory on your GPU where AI model weights live during inference.
Compacts model numbers from 16-bit to 4-bit or 5-bit to fit smaller GPUs.
Working memory used to remember conversation history and long documents.
The generation streaming speed of the AI model.
Splits model layers between fast GPU VRAM and slower System Host RAM.
Combines 2, 4, or 8 identical GPUs into a single unified virtual VRAM pool.
Calculates pure model weight footprint across FP16, Q8_0, Q5_K_M, and Q4_K_M plus real-world CUDA context buffers and activation overhead.
Models Grouped Query Attention (GQA) and Multi-Head Attention memory requirements scaling from 2k to 1M token contexts.
Calculates electricity draw, hardware amortization, and cloud rental costs to determine the exact break-even token threshold against commercial APIs.
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.