Skip to content

Author

Published

Reading time

5 min read

Text size

Share

Email

Artificial IntelligenceNews

Same Price, Stronger Coding: SpaceXAI Ships Grok 4.7 With 46.3% on CursorBench 4.0

Company reports longer task stamina and tighter self-checks; API stays at $2 / $6 per million tokens for short prompts

開發者與 AI 代理多螢幕長時程編碼協作意象
開發者與 AI 代理多螢幕長時程編碼協作意象

Lead

On 21 September 2026, SpaceXAI (xAI) introduced Grok 4.7, positioning it as the company’s most capable model for coding and knowledge work: it works longer on difficult tasks, checks its own work more carefully, and ships with a new safeguard stack. Official materials stress that it is served at the same price and speed as Grok 4.6, while advancing on several company-reported agentic and knowledge-work benchmarks.

The same-day Grok 4.7 Model Card (revision 2026-09-21) documents availability, training cutoffs, multimodal scope, and a high-stakes-use disclaimer. This piece relies only on the official announcement, the model card, and the docs.x.ai models pricing page.

What shipped today

According to the model card, Grok 4.7 is available now through:

  • Cursor — every plan tier
  • Grok Build — default model in SpaceXAI’s terminal coding agent (API and CLI)
  • Grok API — standard chat/completions endpoints from console.x.ai
  • Office add-ins — default model in Grok by SpaceXAI add-ins for Microsoft Word, PowerPoint, and Excel
  • Model gateways — including OpenRouter, Vercel, Cloudflare, Snowflake, Databricks Mosaic, and others

The model card also states that SpaceXAI plans to add Grok 4.7 to consumer surfaces (web, mobile apps, and Grok-in-X on the X platform) at a later date — not as part of today’s launch set.

On modality, the card describes Grok 4.7 as primarily a text model: natural-language text and images in, text out.

Model improvements: larger base, longer RL, native Grok Bot harness

The announcement and model card align on the development story:

  • A new, larger base model versus Grok 4.6
  • A longer reinforcement-learning run on a harder task mix, weighted toward problems that take many hours
  • Better self-verification and longer-context management
  • Training to natively understand the Grok Bot harness, aimed at conversational and general knowledge work
  • Supplemental training on anonymized Cursor workflow data to improve coding and agentic performance (model-card footnote)

Training cutoffs (model card): pretraining data cutoff June 2026; supplemental training uses data generated as late as August 2026.

Note: The docs.x.ai models page separately states that Grok 4.7’s knowledge cut-off is May 2026. This article treats the model card’s pretraining/supplemental wording as primary; any citation of the May date should be labeled as a docs-page statement.

Company-reported benchmark table

Figures below are from the official announcement comparison table. They are SpaceXAI/xAI-reported scores, not independent replications. Grok 4.7’s DeepSWE result is marked * (high effort).

BenchmarkGrok 4.7Grok 4.6GPT-5.6 Sol MaxFable 5.1 Max
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.037.6%20.3%37.3%57.9%
Harvey Legal Agent19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

* high effort (company annotation)

The announcement additionally says Grok 4.7 sits at the price-performance frontier on CursorBench 4.0 (longer-running coding tasks), and that it improves on Grok 4.6 for document/presentation-style knowledge work (e.g. GDPval and AA Briefcase) while performing comparably to other frontier models.

Illustration of an agent self-checking code against tests in a loop

Pricing and speed

docs.x.ai lists grok-4.7 standard API pricing and context as follows (per million tokens; same band as grok-4.6):

ConditionContextInputCached inputOutput
Prompt < 200k tokens500k$2.00$0.50$6.00
Prompt ≥ 200k tokens500k$4.00$1.00$12.00

The announcement also notes a fast variant with roughly 2× output speed at 2× price. Marketing copy on the announcement page frames Grok 4.7 as “twice as fast, at half the price of comparable models” relative to that peer set — a company comparison, not an independent cost study.

Illustration of coding and knowledge-work capabilities balanced on a price-performance scale

Safety and use boundaries

SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack, and calls it the strongest model the company has tested on refusals and jailbreak resistance. In dual-use domains such as cybersecurity and biological work, the company claims strong utility on benign tasks with safe refusal on dangerous ones. Company-reported figures include:

  • LatchBio biosafety-related benchmark: 62.4%
  • HackerBench v0.3: allows only 3.3% of risky dual-use prompts through, while rarely blocking legitimate security work
  • Invite-only red-team access for select cybersecurity partners for defense research

The model card disclaimer states that Grok 4.7 is not intended for autonomous high-stakes decision-making in medicine, law, finance, or safety-critical systems without appropriate human oversight and domain-expert validation. Use remains subject to SpaceXAI’s Acceptable Use Policy, applicable terms, and law.

What changed versus Grok 4.6 (checkable deltas)

Reading the announcement beside the model card, the clearest deltas versus Grok 4.6 are:

  1. 1. Base and training length — larger base plus a longer RL mix weighted to multi-hour tasks.
  2. 2. Behavioural emphasis — longer task stamina, stronger self-checking, better long-context management, and native understanding of the Grok Bot harness.
  3. 3. Product pricing — standard API band unchanged versus 4.6 ($2 / $6 for short prompts; doubles for long prompts), plus a fast variant (~2× speed, 2× price).
  4. 4. Availability scope — developer/enterprise and gateways today; consumer surfaces explicitly later.

On the benchmark table, company figures show Grok 4.7 up versus 4.6 across the listed rows, but not uniformly ahead of GPT-5.6 Sol Max or Fable 5.1 Max — for example Fable leads on CursorBench and Terminal-Bench in the company’s table, while GPT/Fable lead on HealthBench. This article reprints the company table; it does not adjudicate cross-lab methodology.

Two reading notes on the benchmarks

  • CursorBench 4.0 (model card): version 4.0 adds longer-horizon tasks, so scores are not comparable with CursorBench 3.2.
  • DeepSWE: Grok 4.7’s 71.0% carries a * high-effort mark; keep that annotation when reading it beside other models in the same table.

Outlook (official statements only)

What can be verified from today’s primaries: developer and enterprise channels are live, consumer surfaces are planned for later, pricing matches Grok 4.6’s band, and company-reported long-horizon coding and knowledge-work scores moved up versus 4.6. Any later consumer launch dates, third-party replications, or safeguard updates should be taken from subsequent SpaceXAI/xAI notices and model-card revisions.


Primary sources

  • Introducing Grok 4.7|SpaceXAI/xAI https://x.ai/news/grok-4-7
  • Model Card: Grok 4.7 (2026-09-21) https://media.x.ai/v1/website/4p7card-5eccc980.pdf
  • xAI Docs|Models (pricing and knowledge-cutoff note) https://docs.x.ai/docs/models

Source: https://x.ai/news/grok-4-7

Leave a comment

Your email address will not be published. Required fields are marked *

FFOO Labs Newsletter

Occasional notes on AI, technology and the space between imagination and practice.

Follow by RSS