On this page
Put Coworker to work on your stack.
Connect Salesforce, Slack, Jira and run your first agent in minutes.
Enterprise AI
Grok API Pricing in 2026: Every Rate, Checked Against xAI's Docs
Grok API pricing starts at $1.00 per million input tokens and runs to $6.00 output on the flagship. Live rates from xAI's docs, plus how they compare.
What Grok API pricing actually costs today
Rates on this page were read directly from xAI's own pricing page on September 18, 2026. Model rates in this category change several times a year, so treat any number older than a few weeks as a starting point and confirm against the vendor before you commit a budget.
Two things about the structure matter more than any single rate. First, xAI bills three separate token categories: fresh input, cached input, and output. Reasoning tokens bill at the output rate, which means a model set to high reasoning effort costs more per visible word than the output column suggests. Second, every current Grok model carries a long-context tier that kicks in at 200k prompt tokens and roughly doubles all three rates.
One naming note worth clearing up before you compare quotes. xAI's live pricing table lists seven text models, and none of them is called plain "Grok 4". The current flagship is grok-4.6. If a comparison article you are reading quotes $3.00 input and $15.00 output for "Grok 4", it is quoting a rate card that xAI no longer publishes.
The full Grok model lineup and what each tier costs
All figures below are USD per one million tokens, from xAI's pricing page as of September 18, 2026. "Short context" applies to prompts under 200k tokens.
| Model | Context | Input | Cached input | Output |
|---|---|---|---|---|
| grok-4.6 | 500k | $2.00 | $0.50 | $6.00 |
| grok-4.5 | 500k | $2.00 | $0.30 | $6.00 |
| grok-4.3 | 1M | $1.25 | $0.20 | $2.50 |
| grok-build-0.1 | 256k | $1.00 | $0.20 | $2.00 |
| grok-4.20-0309-reasoning | 1M | $1.25 | $0.20 | $2.50 |
| grok-4.20-0309-non-reasoning | 1M | $1.25 | $0.20 | $2.50 |
| grok-4.20-multi-agent-0309 | 1M | $1.25 | $0.20 | $2.50 |
Source: docs.x.ai/developers/pricing, checked September 18, 2026.
The spread inside xAI's own lineup is wider than the spread between vendors. Moving a workload from grok-4.6 to grok-4.3 cuts input cost by 38% and output cost by 58%, and buys a 1M token context window instead of 500k. For anything that is not frontier coding work, that is the first optimization to try.
Here is the same lineup with the long-context column, which is where most budget surprises come from. These rates apply to every token in a request once the prompt reaches 200k.
| Model | Long input | Long cached | Long output |
|---|---|---|---|
| grok-4.6 | $4.00 | $1.00 | $12.00 |
| grok-4.5 | $4.00 | $0.60 | $12.00 |
| grok-4.3 | $2.50 | $0.40 | $5.00 |
| grok-build-0.1 | $2.00 | $0.40 | $4.00 |
| grok-4.20 family | $2.50 | $0.40 | $5.00 |
Grok API pricing vs OpenAI, Anthropic, and Google
This is the comparison most buyers are actually running. Every rate below came from the vendor's own live pricing page on September 18, 2026, and all are USD per million tokens at standard (non-batch) processing.
| Provider | Model | Input | Cached input | Output |
|---|---|---|---|---|
| xAI | grok-4.6 | $2.00 | $0.50 | $6.00 |
| xAI | grok-4.3 | $1.25 | $0.20 | $2.50 |
| OpenAI | GPT-5.6 Terra | $2.00 | $0.20 | $12.00 |
| OpenAI | GPT-5.6 Sol | $4.00 | $0.40 | $20.00 |
| OpenAI | GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| Anthropic | Claude Sonnet 5 | $2.00 | $0.20 | $10.00 |
| Anthropic | Claude Opus 5 | $5.00 | $0.50 | $25.00 |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 |
Sources: xAI, OpenAI, Anthropic, Google, all checked September 18, 2026. Google's Gemini 3.8 Flash rates are promotional through December 31, 2026 and double on January 1, 2027. OpenAI's GPT-5.6 Sol rates are promotional at least through November 21, 2026.
The pattern jumps out once the numbers sit next to each other. Four different models from four different vendors all charge exactly $2.00 per million input tokens: grok-4.6, GPT-5.6 Terra, Claude Sonnet 5, and Gemini 3.1 Pro Preview. Input pricing at the mid-frontier tier has converged. Output is where they separate, and the range is large: $6.00 for grok-4.6 against $10.00 for Sonnet 5 and $12.00 for both Terra and Gemini 3.1 Pro. That is a 2x spread on the half of the bill you control least.
So the honest version of "is Grok cheaper" is: yes on output, at parity on input, and the gap grows with how verbose your workload is. A summarization or code-generation job that produces long responses saves real money on Grok. A classification job that reads a lot and writes three words barely notices the difference. Google's Flash tier undercuts everything in the table on both sides, but that is a different capability class and a promotional rate with a published expiry date.
The full five-provider picture, including the budget tiers
Restricting the comparison to frontier models flatters everyone. Once you add each vendor's cheap tier and DeepSeek, the range across the category is roughly 65x on input and 80x on output, which is the number that should actually drive model selection.
| Provider | Model | Input | Cached input | Output | Batch discount |
|---|---|---|---|---|---|
| xAI | grok-4.6 | $2.00 | $0.50 | $6.00 | None |
| xAI | grok-4.3 | $1.25 | $0.20 | $2.50 | 20% |
| xAI | grok-build-0.1 | $1.00 | $0.20 | $2.00 | None |
| OpenAI | GPT-6 Astra | $10.00 | $1.00 | $50.00 | 50% |
| OpenAI | GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | 50% |
| OpenAI | GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | 50% |
| Anthropic | Claude Opus 5 | $5.00 | $0.50 | $25.00 | 50% |
| Anthropic | Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 50% |
| Anthropic | Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 50% |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | 50% | |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | 50% | |
| DeepSeek | deepseek-flash | $0.15 | $0.003 | $0.60 | Off-peak, see below |
| DeepSeek | deepseek-v4-pro | $0.66 | $0.022 | $1.98 | Off-peak, see below |
Sources as above plus DeepSeek's models and pricing page, checked September 18, 2026. DeepSeek rates shown are off-peak; peak rates are exactly double. Anthropic figures from its published model pricing table, where cache hits price at 0.1x base input on the models listed.
DeepSeek's structure is the one that does not fit the category's shape. Instead of a batch API it prices by clock: peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and every other hour bills at exactly half the peak rate. Off-peak deepseek-flash at $0.15 input and $0.60 output undercuts grok-4.3 by 88% on input and 76% on output. It also caps concurrency at 2,500 for flash and 500 for v4-pro, which is a different constraint shape from the requests-per-second limits the US providers publish.
Why the per-token rate is not the per-word rate
One thing no rate-card comparison captures, including the tables above: a token is not a standard unit across vendors. Each provider tokenizes text with its own vocabulary, so the same paragraph produces a different token count depending on who processes it.
Anthropic publishes the clearest example. Its pricing page states that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text than the previous one, with the exact increase depending on content and workload shape. That is a 30% cost difference sitting entirely outside the rate card, in the same direction for both input and output.
The practical consequence: a 15% headline rate advantage can be wiped out by tokenization before a single request is optimized. If a decision is close, the only comparison that settles it is running the same representative sample of your real traffic through both APIs and comparing reported token counts, not comparing published rates. Every provider returns the counts in the response, so this is an afternoon of work, not a project.
Cached input and batch discounts
Cached input on grok-4.6 costs $0.50 per million, a 75% discount off the fresh input rate. That is a weaker cache discount than the competition: OpenAI, Anthropic, and Google all price cache hits at 10% of base input, versus xAI's 25% on the flagship. On grok-4.3 the cache rate is $0.20 against $1.25 base, which is closer to the industry norm.
Batch processing is the bigger structural difference. xAI's batch API applies a 20% discount, and only to four models: grok-4.3 and the three grok-4.20 variants. The flagship grok-4.6 and grok-build-0.1 have no batch discount at all. OpenAI and Google both advertise 50% off for batch work. If a meaningful share of your volume is asynchronous, that difference can erase Grok's output-rate advantage on its own.
Two more line items that do not appear on a headline rate card. Requests routed to xAI's US regional endpoint bill at 1.1x the global rates, a 10% premium applied after cache discounts. And server-side tools bill separately from tokens: web search and code execution run $5 per 1,000 calls, with X Search moving to $5 per 1,000 posts fetched and $10 per 1,000 user profiles from September 21, 2026.
Cache mechanics, compared properly
The cache-hit rate is only one of three numbers that decide whether caching saves money. Two providers also charge to write to the cache, and one charges rent on it. Read across this table before assuming a cache discount is free money.
| Provider | Cost to write | Cost to read | Storage fee | Effective read rate |
|---|---|---|---|---|
| xAI (grok-4.6) | Not separately priced | $0.50/M | None published | 25% of base input |
| xAI (grok-4.3) | Not separately priced | $0.20/M | None published | 16% of base input |
| OpenAI (GPT-5.6 Terra) | $2.50/M | $0.20/M | None published | 10% of base input |
| Anthropic (Claude Sonnet 5) | $2.50/M (5 min), $4.00/M (1 hour) | $0.20/M | None published | 10% of base input |
| Google (Gemini 3.1 Pro Preview) | Priced as cached input | $0.20/M | $4.50 per M tokens per hour | 10% of base input |
| DeepSeek (deepseek-flash) | Not separately priced | $0.003/M | None published | 2% of base input |
All figures from the vendor pricing pages linked above, checked September 18, 2026.
That table reverses the naive reading of the previous section. xAI's 25% cache-read rate is worse than the 10% its competitors publish, but xAI does not charge for the write and does not rent the cache by the hour. Anthropic publishes the break-even arithmetic directly: its five-minute cache write costs 1.25x base input and pays for itself after a single read, while the one-hour write costs 2x base and needs two reads. Google's context cache adds $4.50 per million tokens per hour of storage on Gemini 3.1 Pro Preview, so a 200k-token cached document costs about $0.90 an hour, roughly $22 a day, whether or not anything reads it.
So the right question is not "who has the best cache rate" but "how many times will each cached prefix be read before it expires". For a chat product where every conversation reuses a large system prompt dozens of times, the read rate dominates and the 10% providers win. For a batch job that caches a document, reads it twice and moves on, write costs and storage fees decide it, and xAI's simpler structure comes out ahead.
DeepSeek is the outlier at 2% of base input, and its cache is automatic rather than something you declare. On a workload with a stable prefix, that is the steepest effective discount published by any provider in this comparison.
Multipliers that sit on top of every rate
Four separate multipliers can apply to the same request. They stack in a defined order, and none of them appears in the headline table.
- Priority processing (xAI): 2x. Higher scheduling priority for lower latency, applied to all token types including cached and reasoning tokens. Prompt caching discounts are applied before the multiplier. xAI bills the priority rate only when the response confirms `"service_tier": "priority"`, so a request that falls back to the default tier bills at standard rates.
- US regional endpoint (xAI): 1.1x. Currently available for grok-4.6 only, applied to input, output and cached input including long-context rates, again after caching discounts.
- Fast mode (OpenAI): roughly 2x. OpenAI's fast tier prices GPT-5.6 Terra at $4.00 input and $24.00 output against $2.00 and $12.00 standard. OpenAI also applies a 10% uplift on regional processing endpoints for models released on or after March 5, 2026.
- US-only inference (Anthropic): 1.1x. For Claude 4.6 and later, specifying US-only inference applies a 1.1x multiplier across input, output, cache writes and cache reads.
There is also a small penalty worth knowing about: xAI charges a $0.05 usage guideline violation fee per request for violations caught before generation in the Responses API, and still bills for generation when a violation is caught later.
Coworker
See what your AI stack really costs
Compare model and platform costs, then run it all in one place.
Open the free LLM cost calculatorA worked cost example
The assumptions here are mine, not sourced data. I picked a shape that looks like a lot of real internal deployments: an internal support assistant handling 20,000 requests a month, each sending 8,000 input tokens of ticket history and retrieved context, and producing 700 output tokens of answer. That works out to 160M input tokens and 14M output tokens a month.
| Setup | Input cost | Output cost | Monthly total |
|---|---|---|---|
| grok-4.6, no caching | $320 | $84 | $404 |
| grok-4.6, 75% of input cached | $140 | $84 | $224 |
| grok-4.3, no caching | $200 | $35 | $235 |
| grok-4.3, batch (20% off) | $160 | $28 | $188 |
| Claude Sonnet 5, no caching | $320 | $140 | $460 |
| GPT-5.6 Terra, no caching | $320 | $168 | $488 |
| Gemini 3.1 Pro Preview, no caching | $320 | $168 | $488 |
Every figure is volume multiplied by the published rate, with no discounts beyond the ones named in the row. Three things fall out of it. Grok's flagship runs about 12% under Sonnet 5 and 17% under Terra on identical volume, entirely on the output side. Caching is worth more than switching vendors: the same model with a warm cache beats every uncached option in the table. And dropping to grok-4.3 costs less than a caching project and takes an afternoon.
Note what this example does not include: tool invocations, the 1.1x US endpoint premium, or the long-context tier. Add a retrieval step that pushes prompts past 200k tokens and the grok-4.6 row jumps from $404 to $808.
Three more workload shapes, and how the ranking changes
The support-assistant shape above is balanced enough that the ranking follows the output rate. Most real workloads are not balanced, and the ranking moves when they are not. Assumptions are stated inline for each; all of them are mine, and all arithmetic is volume multiplied by the published rate.
Classification at volume, where output barely exists
Assume 5,000,000 requests a month, 600 input tokens each and 10 output tokens each. That is 3,000M input tokens and 50M output tokens.
| Model | Input cost | Output cost | Monthly total |
|---|---|---|---|
| deepseek-flash (off-peak) | $450 | $30 | $480 |
| GPT-5.6 Luna | $600 | $60 | $660 |
| Gemini 3.8 Flash | $2,250 | $188 | $2,438 |
| Claude Haiku 4.5 | $3,000 | $250 | $3,250 |
| grok-4.3 | $3,750 | $125 | $3,875 |
| grok-4.6 | $6,000 | $300 | $6,300 |
Output is $300 of grok-4.6's $6,300 here, under 5% of the bill, so the output-rate advantage that decides every other scenario on this page is worth almost nothing. The input rate decides everything instead, and xAI does not publish a small model that competes with GPT-5.6 Luna or deepseek-flash there. If your workload is 98% reading, Grok is the wrong tool and the rate card says so plainly.
Summarization, where output dominates and batch flips the answer
Assume 200,000 documents a month, 4,000 input tokens and 1,200 output tokens each. That is 800M input and 240M output.
| Model | Standard total | With batch discount |
|---|---|---|
| grok-4.3 | $1,600 | $1,280 (20% off) |
| grok-4.6 | $3,040 | $3,040 (no batch discount) |
| Claude Sonnet 5 | $4,000 | $2,000 (50% off) |
| GPT-5.6 Terra | $4,480 | $2,240 (50% off) |
| Gemini 3.1 Pro Preview | $4,480 | $2,240 (50% off) |
This is the single clearest illustration of the batch gap. Run it synchronously and grok-4.6 beats Sonnet 5 by 24% and Terra by 32%, exactly as the output rate predicts. Run it as an overnight batch job, which summarization almost always can be, and Sonnet 5 costs 34% less than grok-4.6 while Terra and Gemini both come in cheaper too. The flagship Grok model has no batch discount to apply, so it does not move while everything around it halves.
The fix inside the xAI lineup is to use grok-4.3, which does carry the 20% batch discount and wins the table outright at $1,280. That is a capability tradeoff rather than a free win, and it is the tradeoff worth testing before concluding Grok is expensive for batch work.
Long-context RAG, where the threshold decides the bill
Assume 50,000 requests a month with 250,000 input tokens of retrieved context each and 800 output tokens. That is 12,500M input and 40M output, and every request sits above the 200k line.
| Model | Rate applied | Input cost | Output cost | Monthly total |
|---|---|---|---|---|
| Claude Sonnet 5 | Flat to 1M | $25,000 | $400 | $25,400 |
| grok-4.6 | Long context | $50,000 | $480 | $50,480 |
| Gemini 3.1 Pro Preview | Long context | $50,000 | $720 | $50,720 |
| grok-4.6, retrieval trimmed to 150k input | Short context | $15,000 | $240 | $15,240 |
Two conclusions, both worth acting on. Anthropic's flat pricing across the full context window halves the bill relative to the two providers that tier at 200k. And trimming retrieval so prompts stay under the threshold saves more than any vendor switch: the last row is 70% cheaper than the second row, same model, same output volume.
Agentic loops, where reasoning tokens and tools both bill
Assume 10,000 tasks a month, 12 model calls per task, 15,000 input tokens and 1,500 output tokens per call (reasoning tokens included, since they bill at the output rate), plus 4 web searches per task. That is 1,800M input, 180M output and 40,000 searches.
| Model | Token cost | Tool cost | Monthly total |
|---|---|---|---|
| grok-4.6 | $4,680 | $200 (web search at $5/1k calls) | $4,880 |
| Claude Sonnet 5 | $5,400 | $400 (web search at $10/1k searches) | $5,800 |
Tool invocations are the small number here, about 4% of the bill in the Grok row, which is the opposite of what most teams expect when they first see a per-call tool price. The number that actually moves is reasoning tokens, because they bill at the output rate. Turning reasoning effort up on an agent does not add a rounding error, it scales the most expensive line on the invoice. If an agent's cost is running high, check the reasoning token count in the usage payload before checking anything else.
The long-context cliff nobody budgets for
This is the single most expensive thing to misread on xAI's rate card. Long-context pricing is not a surcharge on the tokens past 200k. Once a prompt reaches the threshold, the higher rates apply to every token in that request, output included. A 199,000-token prompt and a 201,000-token prompt differ in price by roughly 2x, not by 1%.
Competitors handle this differently enough that it is worth checking per vendor rather than assuming. Google publishes the same two-tier structure at the same 200k boundary for Gemini 3.1 Pro Preview, at $2.00 rising to $4.00 input and $12.00 rising to $18.00 output. Anthropic's published table has no equivalent tier on the models listed above. If your workload stuffs long documents into context, model the threshold explicitly before you compare headline rates, or the comparison is measuring the wrong thing.
The practical mitigation is boring and effective: retrieve less. A retrieval layer that returns the 12 relevant chunks instead of the whole document keeps you under the cliff, and costs less on every request rather than just the ones that would have crossed it.
Who tiers, who does not, and where the line sits
| Provider | Long-context tier | Threshold | What the higher rate applies to |
|---|---|---|---|
| xAI | Yes, on every text model | 200k prompt tokens | Every token in the request, including cached and output |
| Yes, on Gemini 3.1 Pro Preview | 200k prompt tokens | Input, cached input and output, priced per band | |
| OpenAI | Yes, published as separate long-context columns | Per model | Input, cached input, cache writes and output |
| Anthropic | No, on Claude 4.6 and later | Not applicable | Full 1M window at standard rates |
| DeepSeek | No | Not applicable | Flat to the 1M context window |
Anthropic states the position explicitly on its pricing page: Claude 4.6 and later models include the full 1M token context window at standard pricing, and a 900k-token request bills at the same per-token rate as a 9k-token request, with caching and batch discounts applying at standard rates across the whole window. OpenAI's long-context columns are steep in the other direction: GPT-6 Astra goes from $10.00 and $50.00 to $20.00 and $75.00, and GPT-5.6 Terra from $2.00 and $12.00 to $4.00 and $18.00.
Three practical habits follow from that table.
Instrument the token count, not the character count. The threshold is measured in prompt tokens, so a retrieval layer that budgets in characters or documents will cross the line without warning. Log the prompt token count per request and alert on the distribution, not the average, because a small tail of oversized prompts can carry a large share of the bill.
Cap retrieval, then check what the cap costs you. Trimming retrieval is the cheapest fix available and the only one that also improves latency. It has a quality cost, and the honest way to size it is an evaluation run at both context lengths rather than an assumption in either direction.
Check whether long context is the requirement or a habit. Many long-context prompts exist because passing the whole document was easier than building retrieval, not because the model needs it. On the RAG scenario above that habit costs $35,240 a month.
Rate limits and other operational limits
Price is not the only thing that decides whether a provider fits. xAI publishes its rate limits per model and per tier, and the spread between models is larger than the spread between tiers.
Every team has two limits per model: requests per second (RPS) and tokens per minute (TPM). xAI notes that the per-second limit is derived from the per-minute request budget divided by 60, so a full minute of requests cannot be spent in a single second. Exceeding either limit returns a 429.
| Model | Tier 0 RPS | Tier 0 TPM | Tier 4 RPS | Tier 4 TPM |
|---|---|---|---|---|
| grok-4.6 | 150 | 50M | 500 | 100M |
| grok-4.5 | 150 | 50M | 500 | 100M |
| grok-4.3 | 37 | 10M | 208 | 85M |
| grok-build-0.1 | 37 | 10M | 208 | 85M |
| grok-4.20-multi-agent-0309 | 9 | 2.5M | 56 | 21M |
From xAI's rate limits page, checked September 18, 2026.
The headline there is that the flagship is the most generously limited model in the lineup, and it starts that way at Tier 0 with zero spend. A team that picks grok-4.3 to save money gets a quarter of the request headroom and a fifth of the token headroom as a side effect. On a bursty workload that tradeoff can cost more in engineering time than the rate difference saves.
Tiers are set by cumulative spend on the xAI API since January 1, 2026: $50 reaches Tier 1, $250 Tier 2, $1,000 Tier 3 and $5,000 Tier 4, with Enterprise available on request. Qualification counts prepaid credit purchases and fulfilled invoices, and xAI states that tiers never downgrade once reached. That is a friendlier structure than it sounds, because a pilot that spends $250 in its first month keeps that headroom permanently.
Two operational details that change architecture decisions rather than budgets. Batch requests do not count toward rate limits at all, which makes the batch API a throughput tool as well as a discount, and one of the better reasons to use it on grok-4.6 even without a price break. And DeepSeek, for contrast, publishes a concurrency cap rather than a rate limit: 2,500 concurrent requests on deepseek-flash and 500 on deepseek-v4-pro.
Finally, storage bills separately if you use xAI's file and collection features. Files cost $0.025 per GiB per day and collections $0.10 per GiB per day, with downloads from either at $0.20 per GiB. Those are small numbers that become visible on a large document corpus, and they are the one part of an xAI bill that accrues whether or not any inference happens.
When Grok is the right choice, and when it is not
Rate cards do not make decisions. Here is the honest version of what the numbers above imply.
Grok is a strong choice when output dominates the bill. Summarization, code generation, long-form drafting and anything conversational where responses run long. This is where grok-4.6's $6.00 output rate against $10.00 and $12.00 does real work, and the advantage compounds with verbosity.
Grok is a strong choice for synchronous, latency-sensitive traffic. The flagship carries the highest rate limits in xAI's lineup from Tier 0, and priority processing exists if you need more, billed only when the request is actually served at that tier.
Grok is a strong choice when a 1M context window at a mid-tier price matters. grok-4.3 and the grok-4.20 family carry 1M windows at $1.25 and $2.50, which is the cheapest 1M-window option among the US providers listed here.
Grok is the wrong choice for read-heavy, low-output work. The classification scenario above shows the flagship costing 13x deepseek-flash and nearly 10x GPT-5.6 Luna. xAI does not publish a small model that competes at that end of the market.
Grok is the wrong choice for batch-first pipelines on the flagship. No batch discount on grok-4.6 means competitors halve and it does not. Either use grok-4.3, or accept that an overnight pipeline will be cheaper somewhere else.
Grok is the wrong choice when prompts routinely exceed 200k tokens. The all-tokens long-context rule is the harshest published structure of the five providers here, and Anthropic's flat pricing to 1M is the direct counter-example.
What token pricing does not cover
Per-token rates buy model inference. They do not buy the part most teams underestimate, which is getting company context to the model in the first place. A support assistant is only worth its token bill if it can see the ticket, the customer record, the runbook, and the last engineering thread about the same bug. That plumbing is a separate build, and it is where the schedule actually goes.
It is also the part that has to survive a model change. Rates in this category moved multiple times in 2026 alone, and the sensible response is to keep the option to switch. Teams that hardwire one vendor's SDK into every integration pay for that decision twice: once when the rate card moves and once when a better model ships. For a wider view of how the pricing tiers stack up across providers, see enterprise AI pricing compared, the Claude API pricing breakdown for Anthropic's rate structure in detail, or DeepSeek API pricing if you are also evaluating the low-cost end of the market. OpenRouter pricing covers the aggregator route if you want one bill across several of these providers, Qwen API pricing covers another option at the low-cost end, and the LLM gateway pattern is the architecture that makes switching cheap. Our enterprise AI price index tracks how these rates move over time.
Coworker MCP is deliberately model-agnostic on this point. It gives whichever model your team picks, Grok included, unified access to company data across 50+ apps through one governed connection, so switching models later is a configuration change rather than a re-integration. It is not an alternative to xAI, and it does not resell tokens. You still pay xAI for inference at the rates above. Coworker sits at the data layer: Pro at $29.99, Max at $149.99, and custom pricing for enterprise deployments.
If cost per answer rather than cost per token is the question you are actually trying to settle, this comparison of the cheapest ways to access Claude, GPT, and Gemini works through the seat-versus-API math, and ChatGPT Enterprise pricing covers what the per-seat route costs at company scale.
Book a demo to see how Coworker MCP connects your stack to the model you have already chosen.
Frequently asked questions
How much does the Grok API cost?
Grok API pricing is per token with no monthly minimum. As of September 18, 2026, the flagship grok-4.6 costs $2.00 per million input tokens and $6.00 per million output tokens for prompts under 200k, with cached input at $0.50. The cheapest current model, grok-build-0.1, runs $1.00 input and $2.00 output.
Is Grok cheaper than OpenAI or Claude?
On output tokens, yes, and by a meaningful margin: grok-4.6 charges $6.00 per million against $10.00 for Claude Sonnet 5 and $12.00 for GPT-5.6 Terra. On input tokens all three charge exactly $2.00. How much you save depends on how verbose your workload is, because the advantage sits entirely on the output side.
What is Grok's context window?
It varies by model. grok-4.3 and the grok-4.20 family carry 1M tokens, grok-4.6 and grok-4.5 carry 500k, and grok-build-0.1 carries 256k. All of them switch to long-context pricing once a prompt reaches 200k tokens.
Does xAI offer a batch discount?
Yes, but narrowly. The batch API takes 20% off all token types, and only on grok-4.3 and the three grok-4.20 variants. The flagship grok-4.6 has no batch discount. OpenAI and Google both offer 50% for equivalent asynchronous processing, so batch-heavy workloads should price this out rather than assume parity. Batch requests also do not count toward rate limits, which is a reason to use the batch API on grok-4.6 even with no price break.
How does prompt caching change the bill?
Cached input on grok-4.6 costs $0.50 per million versus $2.00 fresh, a 75% saving on cached tokens. In the worked example above, a 75% cache hit rate cut the monthly bill from $404 to $224, a bigger saving than switching vendors would have produced. Caching pays off most when input dominates the workload.
Why did the price of my grok-4.6 request double?
Almost certainly the long-context tier. Once a prompt reaches 200k tokens, xAI applies the long-context rates to every token in that request, not just the ones past the threshold, and that takes grok-4.6 from $2.00 and $6.00 to $4.00 and $12.00. Check your prompt token count against the 200k line before looking for anything more complicated.
What are the Grok API rate limits?
They are set per model and per tier, on two dimensions: requests per second and tokens per minute. grok-4.6 starts at 150 RPS and 50M TPM at Tier 0 and reaches 500 RPS and 100M TPM at Tier 4. grok-4.3 starts at 37 RPS and 10M TPM. Tiers are based on cumulative API spend since January 1, 2026, starting at $50 for Tier 1 and reaching $5,000 for Tier 4, and xAI states that tiers never downgrade.
Do Grok's reasoning tokens cost extra?
They bill at the output rate rather than as a separate category, which for grok-4.6 means $6.00 per million at short context and $12.00 at long context. That makes reasoning effort one of the largest cost levers on an agentic workload, because raising it scales the most expensive line on the bill rather than a minor one.
How much do xAI's server-side tools cost?
Web search, X Search and code execution each bill $5 per 1,000 calls, file attachment search bills $10 per 1,000 calls and collections search bills $2.50 per 1,000 calls, all on top of standard token costs. From September 21, 2026, X Search moves to $5 per 1,000 posts fetched and $10 per 1,000 user profiles. For comparison, Anthropic prices web search at $10 per 1,000 searches and Google gives 5,000 free grounding requests a month across Gemini 3.x before charging $14 per 1,000.
Is Grok cheaper than DeepSeek?
No, and not close. Off-peak deepseek-flash costs $0.15 per million input and $0.60 output against grok-4.3's $1.25 and $2.50, with cache hits at $0.003. DeepSeek also prices by time of day rather than through a batch API, with peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays billing at exactly double the off-peak rate. The comparison worth running is on capability at your workload, not on rate.
Are there costs beyond tokens and tools?
Several, and none of them appear on the headline table. Priority processing bills at 2x standard rates across all token types. The US regional endpoint bills at 1.1x, currently for grok-4.6 only. File storage runs $0.025 per GiB per day, collection storage $0.10 per GiB per day, and downloads from either $0.20 per GiB. xAI also charges a $0.05 usage guideline violation fee per request for violations caught before generation in the Responses API.
Related reading
Ready to get started?
Put Coworker to work inside your actual stack
Connect Salesforce, Slack, Jira, whatever you use, and run your first agent in minutes.