SpaceXAI released Grok 4.7 on September 21, calling it the company's "most capable model for coding and knowledge work" yet. API pricing and speed match Grok 4.6: $2 per million input tokens and $6 per million output tokens, plus a faster tier that costs twice as much and runs twice as fast. The model is already live in Cursor, Grok Build, the API console, and third-party platforms.
"our most capable model for coding and knowledge work."
That's how SpaceXAI frames Grok 4.7's positioning: coding and knowledge work.
According to SpaceXAI, the new model runs on a larger base than Grok 4.6 and went through longer reinforcement-learning training, with a focus on tasks that run for hours at a stretch. The company says it self-checks more reliably, handles longer context, and is now natively tied into Grok Bot.
Official Benchmarks and Third-Party Reviews
SpaceXAI's own numbers: 71.0% on DeepSWE v1.1, 46.3% on CursorBench 4.0, 37.6% on Terminal-Bench 4.0, and 64.0% on EEBench, an electrical-engineering benchmark. On safety, SpaceXAI says its refusal and jailbreak resistance lead the industry, with a 3.3% pass-through rate for risk warnings on HackerBench v0.3; red-team access is invite-only, limited to select cybersecurity partners.
Third-party evaluator Artificial Analysis puts Grok 4.7's overall intelligence index at 46, two points above 4.6; paired with Grok Build, its coding-agent index is 56, nine points above 4.6. On individual benchmarks, DeepSWE v1.1 rose from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. By this measure, Grok 4.7 plus Grok Build ranks fourth among coding agents, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. On AA-Briefcase, a benchmark for long-horizon knowledge work, it scored 1657 Elo, 111 points above 4.6.
SpaceXAI's own release page gives no head-to-head ranking against rivals, and the company has not responded to where third-party evaluations place it.
Doing the Usage Math
The price list hasn't changed, but the token usage Artificial Analysis recorded has: on the xhigh tier, Grok 4.7 outputs roughly 81,000 tokens per task on average, versus about 38,000 for 4.6 at the same tier — a 125% increase. Output speed runs at about 188 tokens per second, with a task taking 7.1 minutes on average.
Doing rough math on output alone, a task on 4.6 costs about $0.23; on 4.7, about $0.49. The sticker price is flat, but the per-task bill has more than doubled — a good chunk of the score gain effectively comes from "thinking longer."
Set against this week's pricing moves, the math stands out even more. The day after Grok 4.7 launched, Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol both went live: Opus 5.5 charges $4 per million input tokens and $20 per million output tokens, while GPT-6 Sol came in at $2 and $10. Measured by sticker price per million tokens, Grok 4.7's output rate is still the lowest of the three; how much that gap narrows once measured by actual per-task spend depends on the other two models' token usage, and that third-party data isn't in yet.
Less than two days after Grok 4.7's 46-point score appeared, Opus 5.5 and GPT-6 had both entered Artificial Analysis's evaluation queue — the top of the leaderboard is about to move again.
Sources: SpaceXAI official release page, CocoLoop, Artificial Analysis evaluation report. API pricing and official benchmark scores are verified against SpaceXAI's release page; the intelligence index, coding-agent index, and token usage figures follow Artificial Analysis's methodology; per-task output cost is a rough estimate based on list price.