Grok 4.6 Review — xAI's Intelligence-per-Dollar Play
On August 12, 2026, xAI shipped Grok 4.6, the successor to Grok 4.5 and the model that pulls the Grok line back onto the intelligence frontier. The headline is not a new absolute ceiling — Claude Opus 5 still owns that — but a jump from an Artificial Analysis Intelligence Index of 54 to 61 at unchanged $2/$6 pricing, which reframes Grok as the intelligence-per-dollar option rather than the also-ran. This review covers what Grok 4.6 actually scores, where its numbers are strong and where they are thin, the pricing band that the headline hides, and how it lands against the GPT-5.6 Sol tier it now ties.
TL;DR verdict
| Grok 4.6 | |
|---|---|
| Type | Frontier-adjacent reasoning + agentic model |
| Released | August 12, 2026 (xAI) |
| Category peers | GPT-5.6 Sol, Claude Fable 5, Kimi K3 |
| Context window | 500,000 tokens (~32K max output) |
| Modalities | Text, image in; text out |
| Pricing (≤200K) | $2 in / $6 out per 1M ($0.50 cached input) |
| Pricing (>200K) | $4 in / $12 out per 1M (long-context band) |
| Headline numbers | AA Index 61, GPQA Diamond 94.9, Terminal-Bench 2.1 88.0 |
| Best for | Reasoning- and tool-heavy agents on a budget |
| Caveat | Coding/SE components trail; price doubles above 200K tokens |
If you read no further: Grok 4.6 is the best intelligence-per-dollar move of the month at the frontier tier. It ties GPT-5.6 Sol on the composite index while costing a fraction as much per token, and on long-horizon agentic workflows its token efficiency compounds that gap. The rating stops short of the top because the software-engineering numbers trail the knowledge scores, and the flat-price framing quietly ends at 200K tokens.
What xAI actually reported
Grok 4.6 keeps the Grok 4.5 hardware envelope — a 500,000-token context window, text-and-vision input, ~32K max output — and pushes the intelligence numbers up while holding the price line. The figures that anchored the launch:
| Benchmark | Grok 4.6 | Grok 4.5 | GPT-5.6 Sol | Notes |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 54 | 61 | Independent (Artificial Analysis) composite of nine evals |
| GPQA Diamond | 94.9 | 91.0 (est.) | 88.0 (est.) | xAI-reported; ties for the top of the board |
| Terminal-Bench 2.1 | 88.0 | 83.3 | 88.8 | xAI-reported; agentic terminal use |
| SWE-bench Pro | ~66 (est.) | 64.7 | 64.6 | Not published by xAI; leaderboard estimate |
Two readings fall out of that table. First, on the axis xAI cares about most — the Artificial Analysis Intelligence Index, an independent composite of nine benchmarks weighted toward agents (34%), coding (24%) and scientific reasoning (24%) — Grok 4.6 lands a 61, tied with GPT-5.6 Sol and two points behind Claude Opus 5’s 63. That is a genuine return to the frontier group after Grok 4.5 sat at #4. Its knowledge and reasoning components (GPQA Diamond 94.9 ties for the top of that board) do the heavy lifting.
Second, note where it is thin. On the newer, much harder Terminal-Bench 3.0, independent write-ups put Grok 4.6 in the mid-20s — well behind the leaders — and the software-engineering components of the composite trail the knowledge ones. xAI did not publish a SWE-bench Pro number, the memorisation-resistant coding benchmark we lean on most. On our leaderboard we therefore store SWE-bench Pro as an explicit conservative estimate (~66, nudged just above Grok 4.5’s vendor 64.7), greyed out and excluded from every ranking — rather than laundering a harness-specific SWE-bench (Vals) figure into a Pro measurement. The GPQA and Terminal-Bench 2.1 cells are stored vendor-reported, with the caveat that they were run on xAI’s own harness. For how much weight self-reported coding numbers deserve, the LLM Benchmark Comparison 2026 lays out the hierarchy.
The real story is cost efficiency
The benchmark that does not fit in a single cell is money per finished task. xAI’s pitch — and the part Artificial Analysis independently echoed — is that Grok 4.6 does long-horizon agentic work in dramatically fewer turns and tokens than its rivals.
On AA’s agent task set, xAI measured Grok 4.6 at roughly 53 turns and 0.5 billion input tokens, against Claude Opus 5’s ~103 turns and 2.0 billion tokens for the same work. Layer the flat $2/$6 pricing on top of that efficiency and the result is roughly 4x cheaper than Opus 5 on the AA-Briefcase agent workload — not because the per-token price is a little lower, but because it burns far fewer tokens to converge. For a multi-step agent that loops dozens of times per task, that compounding is the whole game, and it is where Grok 4.6 is genuinely differentiated rather than merely competitive.
The pricing — and the 200K catch
This is the part to read slowly, because “unchanged pricing” is true only inside a band.
Standard pricing (prompts ≤ 200K tokens):
- Input: $2 per 1M tokens
- Cached input: $0.50 per 1M tokens
- Output: $6 per 1M tokens
Long-context band (prompts > 200K tokens):
- Input: $4 per 1M tokens
- Output: $12 per 1M tokens
At the standard rate, Grok 4.6 sits 60%+ below GPT-5.6 Sol’s $5/$30 while tying it on the composite index — a rare position on the price-capability frontier. It also undercuts Claude Fable 5 ($10/$50) and Claude Opus 5 ($5/$25) by a wide margin. For a reasoning- or tool-heavy agent that lives under 200K tokens per prompt, that is a compelling default right now.
The catch is the band, not a calendar. The 500K context window is real, but the moment a single prompt crosses 200,000 tokens the rate doubles to $4/$12. That is the same structure Grok 4.5 shipped with, so it is not a regression — but it does mean the “cheaper than everyone” framing applies to the first 200K, and long-document or full-repository prompts land in the pricier tier. If you are sizing a budget for RAG over large corpora or whole-codebase agents, model the $4/$12 numbers. Prompt caching at $0.50 narrows the gap for repeated-context workloads; the leaderboard’s cost calculator will run either tier against your own token mix.
How it stacks up
Against its own predecessor the verdict is simple: Grok 4.6 strictly dominates Grok 4.5 — a seven-point jump on the composite index, higher published GPQA and Terminal-Bench 2.1, same 500K context and same price. If you are on 4.5, this is an upgrade with no obvious downside beyond re-validating your prompts.
Against the frontier tier the picture is a genuine contest with a clear shape. GPT-5.6 Sol ties it on the AA index and still edges it on published Terminal-Bench 2.1 (88.8 vs 88.0) and pure software-engineering work — but costs more than twice as much per input token and far more per output. Claude Opus 5 remains two points ahead on the index and clearly stronger on coding-heavy agentic tasks, so if SWE-bench Pro-style work is your bottleneck, pay for Opus. Where Grok 4.6 wins is any workload dominated by reasoning, tool orchestration and token budget rather than raw code generation. As with the Kimi K3 launch, the honest move is to weight the self-reported coding figure as directional until a neutral harness confirms it, then run your own bake-off.
Who should care
- Teams running reasoning- and tool-heavy agent pipelines: This is the headline audience. The intelligence-per-dollar and token-efficiency story is real — just keep individual prompts under 200K to stay in the cheap band.
- Anyone on Grok 4.5: Upgrade. It is better on every number xAI published at the same price; the only cost is re-testing your prompts.
- Cost-sensitive frontier buyers: If you were paying GPT-5.6 Sol or Opus 5 rates for work that is more reasoning than coding, Grok 4.6 is the obvious A/B test.
- Coding-first teams: Be cautious. The SE and Terminal-Bench 3.0 numbers trail, and there is no published SWE-bench Pro cell — validate on your own repository before switching a coding agent over.
FAQ
When was Grok 4.6 released? xAI shipped it on August 12, 2026, as the successor to Grok 4.5.
How much does it cost? $2 per 1M input tokens ($0.50 cached) and $6 per 1M output for prompts up to 200K tokens; above 200K the rate doubles to $4/$12. Pricing is unchanged from Grok 4.5.
How good is it at coding? Mixed. Terminal-Bench 2.1 is a strong 88.0, but the harder Terminal-Bench 3.0 and the software-engineering components of the AA index trail the leaders, and xAI published no SWE-bench Pro figure — so our leaderboard carries an estimate for that cell.
How does it compare to GPT-5.6 Sol and Claude Opus 5? It ties GPT-5.6 Sol at 61 on the AA Intelligence Index and sits two behind Opus 5’s 63, at a fraction of both models’ per-token price and with markedly better token efficiency on long agent tasks.
Is it open weights? No. Grok 4.6 is a closed model available through the xAI API, OpenRouter, and the Grok apps — not downloadable for self-hosting.
Continue reading
- AI Models Leaderboard — Grok 4.6 versus 78 other models on benchmarks, pricing, context window, and throughput, with every figure sourced.
- Grok 4.5 Review — the predecessor Grok 4.6 strictly improves on, at the same price.
- GPT-5.6 Sol, Terra & Luna Review — the OpenAI tier Grok 4.6 now ties on the composite index and undercuts on price.
- Claude Opus 5 Review — the frontier leader Grok 4.6’s intelligence-per-dollar pitch is measured against.
- LLM Benchmark Comparison 2026 — how to read agentic benchmarks and self-reported numbers without getting fooled.
- All Reviews — index of every head-to-head review on the site.