Key Notes
- SpaceXAI launched Grok 4.7, describing it as its most capable model yet for coding and knowledge work while keeping pricing and speed the same as Grok 4.6.
- The company's own benchmarks show Grok 4.7 improving on its predecessor and beating OpenAI's GPT-5.6 Sol on several tests, though it trails Anthropic's Fable 5.1 on coding and terminal benchmarks.
- SpaceXAI also highlighted a new safety stack, citing strong results on refusal and jailbreak resistance testing along with a top score on LatchBio's biosafety benchmark.
SpaceXAI has introduced Grok 4.7, the newest version of its Grok family of models, positioning it as the company’s strongest option yet for coding and knowledge work. The model is available immediately through Cursor, SpaceXAI’s own Grok Build tool, the Grok API, and a range of third party coding harnesses and cloud platforms.
Grok 4.7 is served at the same price and speed as its predecessor, Grok 4.6, which SpaceXAI released roughly six weeks earlier. The company says the new model works longer on difficult tasks, checks its own output more carefully before returning an answer, and includes what it calls its best calibrated safety systems to date.
A Larger Base Model With a Harder Training Mix
According to SpaceXAI’s announcement, Grok 4.7 is built on a new, larger base model than Grok 4.6. The company trained it with a longer reinforcement learning run over a harder mix of tasks, weighted toward problems that can take many hours to complete rather than short, single step prompts.
SpaceXAI also says the model is better at verifying its own work and at managing longer context windows during extended sessions. A separate change was training Grok 4.7 to natively understand the Grok Bot harness, which the company says improves its performance on conversational tasks and general knowledge work rather than just coding.
Benchmark Performance and Pricing
SpaceXAI published a series of benchmark comparisons alongside the launch, measuring Grok 4.7 against Grok 4.6, OpenAI’s GPT-5.6 Sol, and Anthropic’s Fable 5.1. On CursorBench 4.0, a test built around longer running coding tasks, SpaceXAI says Grok 4.7 scored 46.3 percent, ahead of Grok 4.6 at 40.4 percent and GPT-5.6 Sol at 41.7 percent, though behind Fable 5.1 at 51.8 percent. No comparison to GPT-6 Astra Ultra has been shared.
On Terminal-Bench 4.0, a measure of multi hour terminal work, SpaceXAI’s own figures put Grok 4.7 at 38 percent, again trailing Fable 5.1’s 57.9 percent but ahead of Grok 4.6.
The company also cited results on EEBench, an electrical engineering benchmark, and the Harvey Legal Agent Benchmark, where it reported Grok 4.7 scoring 64 percent and 19.6 percent respectively, both ahead of the other three models SpaceXAI listed. On HealthBench Professional, a clinical reasoning test, SpaceXAI’s chart put Grok 4.7 behind both GPT-5.6 Sol and Fable 5.1.
For document and presentation work, SpaceXAI pointed to GDPval and its own AA Briefcase benchmark, both of which task AI systems with work typically done by professionals such as lawyers, nurses, and financial analysts. The company’s GDPval chart showed Grok 4.7 improving on Grok 4.6 and landing close to, though still behind, Fable 5.1 and slightly ahead of OpenAI’s GPT-6 Astra.
Pricing starts at 2 dollars per million input tokens and 6 dollars per million output tokens, unchanged from Grok 4.6. SpaceXAI also offers a faster variant at twice the output speed for twice the price. By comparison, the company’s chart shows GPT-5.6 Sol priced at 4 dollars per million input tokens and 20 dollars per million output tokens, while Fable 5.1 is listed at 10 dollars and 50 dollars respectively, making Grok 4.7 the cheapest of the three models SpaceXAI included in its comparison.
On DeepSWE v1.1, a separate software engineering benchmark, SpaceXAI reported Grok 4.7 scoring 71.0 percent when run at its highest effort setting, a figure the company flagged separately in its chart. That put it ahead of Grok 4.6’s 65.2 percent and Fable 5.1’s 70.0 percent, though slightly behind GPT-5.6 Sol’s 72.7 percent. SpaceXAI frames the model’s overall positioning as sitting at the frontier of price to performance on CursorBench 4.0, rather than as an outright leader on every benchmark it published.
Safety Claims and Cybersecurity Testing
SpaceXAI says Grok 4.7 was built around an entirely new safeguard stack and describes it as the strongest model the company has tested on refusals and resistance to jailbreak attempts. In dual use domains such as biology and cybersecurity, the company reported that Grok 4.7 topped LatchBio’s biosafety benchmark with a score of 62.4 percent.
On HackerBench v0.3, a SpaceXAI benchmark for risky and malicious cyber tasks, the company says the model blocked all but 3.3 percent of risky dual use prompts while rarely refusing legitimate security work. SpaceXAI has also begun giving select cybersecurity partners invite only access to the model’s red team capabilities for defense research purposes.
Context And Availability
Grok 4.7 arrives roughly seven months after SpaceX completed its acquisition of xAI in an all stock deal that valued the combined company at 1.25 trillion dollars. The Grok brand and its API continue to operate under the xAI name even as the parent company markets the model under the SpaceXAI label.
The model is live today for developers through the Grok API console, and SpaceXAI says it is also being made available through third party model routers and cloud platforms, following the same distribution pattern the company used for Grok 4.6’s launch in August.
As with prior releases, the benchmark figures in SpaceXAI’s announcement come from the company’s own testing rather than independent verification, and results on tasks such as GDPval and AA Briefcase can vary depending on how effort settings and tool access are configured. SpaceXAI did not publish a technical report alongside the launch, and it is not yet clear how Grok 4.7 performs on third party leaderboards such as the Artificial Analysis Intelligence Index, which has separately tracked comparable scores for models including OpenAI’s GPT-6 Astra.
For enterprise customers, SpaceXAI continues to offer a business affiliate agreement and data processing addendum covering Grok 4.7, the same documents it maintains for its other API models, and the company says government customers can request access through a separate track. Developers can create an API key and start testing the model immediately through the console, with documentation available at docs.x.ai.
The release keeps SpaceXAI on roughly the same six to eight week cadence it has followed for major Grok updates over the past year, a pace that puts pressure on rivals such as OpenAI and Anthropic to match both the frequency and the pricing of new releases.
Whether Grok 4.7’s price to performance claims hold up once developers move it into production, particularly on the long running agentic tasks SpaceXAI is emphasizing with benchmarks like AA Briefcase and Terminal-Bench 4.0, will likely shape how the model is received over the coming weeks.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.