xAI's Grok models offer a flagship tier with a large context window and a low output-to-input price ratio, plus a flat cached-input rate, which can suit output-heavy workloads.
| Model | $ input /1M | $ output /1M | $ cached /1M | Batch | ≈ $/mo * |
|---|---|---|---|---|---|
| Grok 4.5FRONTIER | $2 | $6 | $0.30 | — | $342 |
| Grok 4.3FRONTIER | $1.25 | $2.50 | $0.20 | −20% | $178 |
| Grok Build 0.1MID | $1 | $2 | $0.20 | — | $148 |
* Example workload — chatbot, 100k requests/mo, 2,000 input / 300 output tokens per request, 70% of input cached. Computed by the same engine as the calculator. Batch: the −20% is xAI's verified Batch API discount; the ≈ $/mo column is computed without it.
Cache pricing differs per model: Grok 4.3 at $0.20/1M (16% of input); Grok Build 0.1 at $0.20/1M (20% of input); Grok 4.5 at $0.30/1M (15% of input). The calculator models this with your cache share.
The Batch API runs asynchronous jobs at a verified −20% on both input and output across Grok 4.3 — flip the Batch toggle in the calculator to model it.
Grok 4.3 runs a verified 1M-token context window; Grok 4.5 is 500k tokens; Grok Build 0.1 is 256k tokens. Grok 4.5, Grok 4.3 and Grok Build 0.1 bill higher rates on long-context prompts — this table and the calculator use standard rates.
| Use case | Verdict |
|---|---|
| Output-heavy workloads | Grok's output rate is relatively low versus input |
| Large context within a single request | The flagship tier carries a large window |
| You need a confirmed batch discount up front | Grok 4.3 has a verified −20%; the newer tiers don't publish one yet |