Grok 4.6 Pushes SpaceXAI Back to the Frontier, and It's Cheap
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and trailing only Anthropic, while undercutting rivals on price and excelling in agentic tasks.

SpaceXAI's Grok 4.6 is out, and it's a serious contender again. The model scores 61 on the Artificial Analysis Intelligence Index, a 5-point jump over Grok 4.5 and a 23-point leap from Grok 4.3. That puts it squarely in the frontier club, tied with GPT-5.6 Sol (max) and just behind Claude Opus 5 (63) and Claude Fable 5 (62). For a model that's priced at $2/$6 per million input/output tokens, that's a statement.
Agentic performance is the headline
Grok 4.6 doesn't just ace static reasoning—it shines where agents actually work. On GDPval-AA v2, its Elo of 1753 trails only Claude Opus 5, and it's statistically tied with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on τ³-Banking (top two alongside Qwen3.8 Max's 51.3%) and 88.4% on Terminal-Bench v2.1, matching the leaders on terminal-based tasks. Few models are this strong across knowledge work, customer service, and terminal use simultaneously.
Cost efficiency that breaks the pattern
Holding pricing flat across a generation is rare at the frontier. Grok 4.6 keeps the $2/$6 pricing from Grok 4.5, while rivals like Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) charge a premium. Measured cost per task is $0.84, same as Kimi K3 but with higher intelligence, placing it on the Pareto frontier for cost vs. performance. For reasoning-heavy workloads, output token price dominates, and Grok 4.6 undercuts the competition by 4-5x on that front.
Efficiency on long-horizon tasks
On AA-Briefcase, a private benchmark for long-horizon agentic knowledge work, Grok 4.6 debuts with an Elo of 1577—Fable 5-tier, behind the Claude Opus 5 family. But the real story is efficiency: it completes tasks in ~53 turns and ~0.5B input tokens, versus ~103 turns and ~2.0B tokens for Claude Opus 5 (max). Long-horizon tasks accumulate context fast, so a model that uses a quarter of the input tokens has a cost advantage that goes beyond per-token pricing.
Context window stays at 500k tokens, and cache hit pricing rose to $0.5 per 1M tokens from $0.3—a minor bump that doesn't dent the overall value proposition.
Grok 4.6 is a reminder that the frontier isn't just about raw intelligence; it's about doing the work without burning through your budget. SpaceXAI is back in the game, and it's bringing the price war with it.
Grok 4.6 delivers a 5-point Intelligence Index gain at unchanged $2/$6 pricing—a rare move at the frontier, where intelligence gains usually come with price hikes.
| Model | Intelligence Index | Price (in/out per 1M) | Cost per Task | GDPval-AA Elo |
|---|---|---|---|---|
| Claude Opus 5 (max) | 63 | $5/$25 | — | Top |
| Claude Fable 5 (max, fallback) | 62 | — | — | — |
| Grok 4.6 | 61 | $2/$6 | $0.84 | 1753 |
| GPT-5.6 Sol (max) | 61 | $5/$30 | — | — |
| Kimi K3 | ~60 | — | $0.84 | — |
Discussion
0 Comments
Be the first to start the discussion.