xAI launched Grok 4.6 on 12 August. The model reportedly scores 61 on Artificial Analysis' intelligence index, tying with GPT-5.6 Sol Max, while keeping the previous version's pricing, around $2 for input and $6 for output per million tokens, with a 500,000-token context window. The point worth noting isn't the performance: it's how pricing shifts with context length.
Another model release, with a pricing detail that matters well beyond this specific case.
What's been announced
Matching the best model from a competitor while holding pricing steady is a solid commercial feat. The index cited aggregates several evaluations and provides a vendor-independent benchmark, which is preferable to in-house numbers, as we noted in our article on benchmarks.
The 500,000-token window places the model in the long-context category, which appeals to those processing large documents or running agents over extended sessions.
The long-context trap
Here's the mechanism to understand, because it doesn't apply only to this model.
Several providers use tiered pricing: the per-token price changes beyond a certain context volume. A model listed at $2 for input can jump to $4 beyond a threshold, with a comparable increase on output.
The technical reason is real: processing a very long context costs more than proportionally, because the attention mechanism compares elements against each other. Doubling the length doesn't double the cost—it increases it further.
The problem is that this threshold is often crossed without notice. An agent accumulating history gradually exceeds the limit. A document-processing job receiving variable-size files tips over for some of them. You discover on the bill that part of your requests were charged at the higher rate.
Know your threshold. It's in the pricing documentation, rarely in the announcements. Write it down.
Measure your distribution. How many of your requests actually exceed that threshold? Often a minority, but one that weighs heavily.
Trim the context. Many apps send the entire history out of convenience. Summarising older exchanges rather than retransmitting them cuts cost and often improves quality, as the model gets less lost.
What it says about the market
This release confirms two trends we're tracking.
Performance gaps are narrowing. When a model matches a competitor's best on an independent index, the choice stops being about capability and becomes about price, reliability, and ecosystem.
Pricing complexity is rising. Context tiers, differentiated cache rates, variable effort modes, batch processing: comparing two providers on their listed prices no longer makes much sense. That's precisely why we keep saying you need to measure cost per completed task on your own use cases.
What to take away
The real lesson goes beyond this model. We're entering a phase where pricing grids are becoming as complex as telecom operators', with thresholds, options, and tiers.
This complexity isn't necessarily malicious: it reflects real costs that vary with usage. But it has the same effect as in telecoms: those who don't read the terms pay more than those who do. The only defence is to measure your own consumption rather than compare listed prices.