Skip to content

Author

Published

Reading time

7 min read

Text size

Share

Email

Artificial IntelligenceNews

Grok 4.6 Arrives: xAI Returns to the Intelligence Frontier, With Price as Its Sharpest Weapon

xAI released Grok 4.6 on 12 August 2026, roughly one month after Grok 4.5. The headline of this release is not a bigger parameter count — the company declined to disclose one — but a deliberate focus on long-running agents and ambitious interactive visual work: multi-step research, deep work inside large codebases, and turning rough product ideas into polished prototypes.

Grok 4.6 Arrives: xAI Returns to the Intelligence Frontier, With Price as Its Sharpest Weapon

xAI released Grok 4.6 on 12 August 2026, roughly one month after Grok 4.5. The headline of this release is not a bigger parameter count — the company declined to disclose one — but a deliberate focus on long-running agents and ambitious interactive visual work: multi-step research, deep work inside large codebases, and turning rough product ideas into polished prototypes.

Independent evaluator Artificial Analysis published its own results the same day, scoring Grok 4.6 at 61 on the Intelligence Index — level with GPT-5.6 Sol (max) and enough to return xAI to the frontier tier, trailing only Anthropic.

Training Focus: Not Bigger, Just Better at Checking Itself

xAI framed this release around methodology rather than scale. Grok 4.6 went through a longer supplementary training phase than Grok 4.5, built on three changes:

Curated model-generated data targeting reasoning and advanced technical concepts

Regenerated supervised fine-tuning (SFT) trajectories produced by Grok 4.5 itself, spanning multiple reasoning efforts, agentic harnesses, and domains including STEM and software engineering, with problematic trajectories filtered out by model-based detection

Reinforcement learning across general coding and knowledge work, plus specialised environments covering kernel optimisation, web development, and computer-aided design (CAD)

According to xAI, this produced a behavioural shift worth noting: across longer task trajectories, the model tests and verifies its own work more frequently, checking results before advancing to the next step. For visual and interactive projects, Grok 4.6 tends to establish both the structure and the visual language of an application within a single generation, reducing the iterative back-and-forth that earlier models required before a design settled.

xAI’s Official Evaluation Table

Below is the full comparison table published by xAI, placing Grok 4.6 High alongside its predecessor Grok 4.5 High, GPT-5.6 Sol Max, and Fable 5 Max. Best score per evaluation is in bold; third-party scores represent the best of self-reported or publicly available results.

EvaluationGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%—58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Three signals emerge from the data.

Generational gains are substantial. Grok 4.6 beats Grok 4.5 on all ten evaluations. APEX-Agents jumped from 47.1% to 57.5% — more than ten percentage points — while Terminal-Bench v3.0 climbed from 15.7% to 26%, DeepSWE v1.1 from 54% to 65.9%, and AA-Briefcase from an Elo of 1313 to 1577. The training changes clearly landed on agentic work.

Knowledge work and professional domains lead. Grok 4.6 takes the best score in four evaluations: GDPVal-AA v2 (1753), AA-Briefcase (1577), and Harvey LAB (15.8%). The margin on Harvey LAB — a legal work evaluation — is the most striking of the set: 15.8% against 2.5% for GPT-5.6 Sol Max, a gap of nearly six times. For teams handling contract review, regulatory analysis, or professional document workflows, that is a meaningful signal.

Repository-scale coding remains the weak spot. Terminal-Bench v3.0 at 26% trails GPT-5.6 Sol Max (34.6%) and Fable 5 Max (34.1%) by roughly eight percentage points, and DeepSWE v1.1 at 65.9% sits more than seven points behind GPT-5.6 Sol Max (73%).

One caveat matters here. This table aggregates competitors’ self-reported or publicly available best results, and those figures come from different test configurations, harnesses, and tool-access permissions. It is not a controlled four-way experiment. For genuinely comparable numbers, third-party evaluation is required.

Independent Evaluation: The Artificial Analysis Intelligence Index

Artificial Analysis is among the most closely watched independent model evaluators in the industry. Its Intelligence Index — v4.1.1 in this round — combines nine evaluations spanning reasoning, knowledge, mathematics, coding, and agentic work, with every model run in-house under identical conditions. That makes it considerably more comparable than vendor-published figures.

Frontier Standings

ModelIntelligence IndexInput / Output (per 1M tokens)
Claude Opus 5 (max)63$5 / $25
Claude Fable 5 (max, with fallback)62—
Grok 4.6 (high)61$2 / $6
GPT-5.6 Sol (max)61$5 / $30
Kimi K3Just below 61—
Grok 4.5 (high)56$2 / $6

Grok 4.6 gains five points over Grok 4.5 and 23 points over Grok 4.3. Given that Grok 4.5 shipped barely a month earlier, that iteration pace is notable. Artificial Analysis concluded that the release returns xAI to the intelligence frontier alongside OpenAI, behind only Anthropic.

Agentic Work: Where the Real Strength Sits

Artificial Analysis specifically noted that Grok 4.6’s strongest showing lands on agentic tasks rather than static reasoning:

GDPval-AA v2 (real-world agentic knowledge work): Elo 1753, second only to Claude Opus 5, with confidence intervals overlapping Claude Fable 5 and Qwen3.8 Max — statistically indistinguishable from both

τ³-Banking (multi-turn customer service and tool use): 50.7%, placing it in the top two, marginally behind Qwen3.8 Max at 51.3%

Terminal-Bench v2.1 (terminal-based software tasks): 88.4%, level with leading models

AA-Briefcase (long-horizon agentic knowledge work, private evaluation): Elo 1577, in the Fable 5 tier, behind the Claude Opus 5 family

Few models remain competitive across knowledge work, customer service, and terminal operation simultaneously. Combined with its pricing, Grok 4.6 sits on the cost-performance Pareto frontier for every agentic evaluation within the Intelligence Index.

Turn Efficiency: The Underrated Cost Lever

One finding deserves more attention than the scores themselves. On AA-Briefcase, Grok 4.6 completed tasks in roughly 53 turns and 0.5 billion input tokens. Claude Opus 5 (max) required approximately 103 turns and 2.0 billion input tokens — meaning Grok 4.6 reached comparable output quality with under half the turns and a quarter of the input volume.

For long-horizon agent work, this matters more than the headline price gap suggests. Agentic tasks accumulate context with each turn, so token consumption grows non-linearly, and turn efficiency compounds the advantage well beyond what per-token pricing alone would indicate. Artificial Analysis measured an average cost of $0.84 per task across the Intelligence Index — the same as Kimi K3, at a slightly higher intelligence score.

Speed: Fast Output, Slow First Token

The independent testing also surfaced a clear experiential trade-off. Grok 4.6 produces roughly 85.8 output tokens per second, above the 71.2 median for reasoning models in its price band. But time to first token measures 32.30 seconds, far behind the 2.88-second median. For interactive conversation that latency is conspicuous; for background batch agent work, considerably less so.

Pricing and Availability: Economics Are the Real Differentiator

Grok 4.6 holds Grok 4.5’s pricing at 2 dollars per million input tokens and 6 dollars per million output tokens, with a fast variant at double the price. Against Claude Opus 5 at 5/25 dollars and GPT-5.6 Sol at 5/30 dollars, that puts Grok 4.6’s output cost at one-fifth of GPT-5.6 Sol’s — and total cost more than 60 percent below both rivals — while matching them on the Intelligence Index.

Holding pricing flat across a generational upgrade is uncommon in the frontier market, where capability gains have typically arrived with price increases. Two details temper the picture, however:

Cache discount narrowed: cache-hit pricing rose from 0.30 dollars per million tokens on Grok 4.5 to 0.50 dollars

Long-context pricing doubles: the context window reaches 500K tokens, but once a single prompt hits 200K tokens, input, cached input, and output rates all double to 4/1/12 dollars — and the higher rate applies to every token in the request, not merely the excess

That tiered structure is a critical budgeting variable for RAG and agentic applications working over large document sets or accumulating context over long sessions.

The model is live in Cursor and Grok Build, and available through the xAI API, OpenRouter, Vercel, and Cloudflare. xAI offered double free usage in Grok Build and Cursor during launch week. Grok 4.6 accepts text and image input, supports a 500K-token context window, and carries a knowledge cutoff of 1 February 2026.

Closing Assessment

Grok 4.6 reflects how frontier competition has shifted from raw intelligence scores toward a composite contest of agentic reliability and cost efficiency. Independent and vendor data converge on the same conclusion: the top three models now cluster tightly between 61 and 63 on the Intelligence Index, and the real differentiation lies in strength distribution and economics. Grok 4.6 leads on knowledge work, professional domains such as legal and financial documents, and turn efficiency, while GPT-5.6 Sol and Fable 5 retain their edge on repository-scale software engineering.

For organisations running long-lived agents over research, document analysis, or knowledge work, Grok 4.6 offers one of the more compelling performance-to-cost combinations currently available. That said, controlled evaluations do not guarantee equivalent savings in production. Harness design, prompting, tool calls, caching strategy, and retry logic all reshape actual token and turn consumption — which makes independent verification a prerequisite for any deployment decision.

Vendor evaluation data from xAI’s published Evals table. Independent data from the Artificial Analysis Intelligence Index v4.1.1, tested 12 August 2026. Cross-model comparisons drawn from vendor self-reporting reflect differing test conditions and should be treated as directional.

Leave a comment

Your email address will not be published. Required fields are marked *

FFOO Labs Newsletter

Occasional notes on AI, technology and the space between imagination and practice.

Follow by RSS