Grok 4.7 was launched by xAI on Monday. It beat Anthropicβs Fable 5.1 and OpenAIβs GPT-5.6 Sol on a legal-agent benchmark, while charging a fifth of Fableβs price per input token.
Output runs $6 per million tokens against $50 at Anthropic
xAIβs release said that Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark. Fable 5.1 got 6.7% on the same test and GPT-5.6 Sol got 2.5%.
That puts Grok 4.7 at about three times Anthropicβs score and close to eight times Solβs. Cryptopolitan reported last month that Grok 4.6 already led that duo at 15.8%.
Legal work is one of the few domains where Grok 4.7 outperforms both competitors outright. It also clears Fable 5.1 on EEBench electrical engineering and edges past it on the DeepSWE coding test.
Grok 4.7 pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. That input rate is 1/2 of GPT-5.6 Solβs $4 and 1/5 of Fable 5.1βs $10.
Output runs $6 vs $20 for Sol, and $50 for Fable, about an eighth of Anthropicβs number.
xAI plotted the CursorBench 4.0 scores against the average cost per completed task and claims the model sits on the price-performance frontier. Grok 4.7 gets around 46% on that chart at ~$6 per task, while Claude Opus 5 needs almost double the spend for a similar outcome.
Fable 5.1 performs better at higher budgets, going up to 51.8% at around $17 per task. The release was βa strong combination of intelligence, speed & low cost,β Musk said on X.
Token prices and benchmark scores from SpaceXAIβs Grok 4.7 launch post, published September 21, 2026.Terminal-Bench 4.0 hands Anthropic a 57.9% to 38.0% lead
Grok 4.7 was runner-up to Fable 5.1 on GDPval and the AA Briefcase office-work test. Anthropicβs model keeps a clear lead on longer coding and terminal benchmarks.
The largest margin is on Terminal-Bench 4.0, where Fable 5.1 had 57.9% versus Grok 4.7βs 38.0%, a difference of some 20 points.
Fable also leads on CursorBench and on HealthBench Professional clinical reasoning, where GPT-5.6 Sol also beats Grok.
Grok 4.7 scored 1,695 Elo on GDPval, which scores models on tasks done by lawyers, nurses and financial analysts, up from Grok 4.6βs 1,605. Thatβs behind Fable 5.1βs tally of 1,735, but ahead of the 1,542 OpenAIβs newer GPT-6 Astra got on the same chart.
Grok 4.7 is ahead of GPT-5.6 Sol on five of seven benchmarks. It only loses DeepSWE and the clinical test.
Grok 4.7 runs on 2.1 trillion parameters, 40% more than the 1.5 trillion powering Grok 4.6. xAI has included supplemental SpaceX training data, including Starlink satellite telemetry and manufacturing records.
The company says the model is more likely to spend extra time on arduous problems and double-check its own answers than Grok 4.6.
The model launched on the Grok app, Cursor, Grok Build and the xAI API, with no waitlist.
All numbers here are from xAIβs own testing. CursorBench, the test xAI leads with, is made by Cursor, which SpaceX finished acquiring last month.
xAI benchmarked against GPT-5.6 Sol, and OpenAIβs newer GPT-6 Astra appears only in the GDPval, AA Briefcase and EEBench charts.
Donβt just read crypto news. Understand it. Subscribe to our newsletter. It's free.



















English (US)