Comprehensive benchmark scores, coding performance, latency metrics, and API pricing for all models developed by xAI.
Portfolio at a glanceRanked catalog of 5 xAI models.
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Model fact sheetGrok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Model fact sheetxAI's flagship reasoning model, strongest on knowledge work and STEM and weaker on pure software-engineering suites.
Model fact sheetGrok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...
Model fact sheetGrok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and...
Model fact sheetJump straight to their strongest model, or compare the lab against its peers.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.