SpaceXAI’s most powerful Grok 4.7 released with crazy price and disappointing performance

📅 2026-09-22

Abstract:

This morning, SpaceXAI’s most powerful model, Grok 4.7, arrived. There is only one very "offensive" positioning at the beginning of the release page: The most powerful model for programming and knowledge work, twice as fast as comparable models and only half the price.


This upgrade does not stop at the benchmarks. Grok 4.7 switches to a larger basic model, strengthens long-term tasks, self-examination and long-context management capabilities, and includes documents, presentations, legal, medical and engineering and other professional work into key tests. The scenario it targets is very clear: allowing the model to avoid detours in tasks that last several hours and deliver results that can be directly used.


Musk tweeted, "Grok 4.7 achieves a very competitive balance between intelligence level, operating speed and cost."


Long missions have become the main battlefield of this generation of models

Grok 4.7 uses a larger base model than Grok 4.6. The training phase extends the reinforcement learning period and the task combinations become more difficult, with more samples taking hours to complete. SpaceXAI said the new model has improved its ability to inspect its own work and manage long contexts.

This type of improvement directly corresponds to the most difficult problems of currently programming intelligent agents. Short code generation often only requires partial judgment, but truly complex software tasks include reading the warehouse, dismantling requirements, modifying multiple files, running tests and repeated error correction. Whether the model can maintain goals and detect errors in a long execution chain is often more important than writing beautiful code in one go.

CursorBench 4.0 specifically examines long-term coding tasks. In the official data, Grok 4.7 scored 46.3% and Grok 4.6 scored 40.4%, an increase of 5.9 percentage points. On DeepSWE v1.1, Grok 4.7's high-inference intensity score was 71.0%; Terminal-Bench 4.0 increased from 20.3% in the previous generation to 38.0%.

Accordingly, Grok 4.7 is called the cutting-edge price and performance model on CursorBench 4.0.


The model is also trained to natively understand the framework within which Grok Bot runs. Officially, this adjustment improves performance on dialogue tasks and general knowledge tasks. In other words, the upgrade focus of Grok 4.7 has extended from a single round of answers to continuous collaboration between models, tools, and execution environments.

From writing code to complete knowledge work

Programming is still the most eye-catching label of Grok 4.7, but the scope of the evaluation given by SpaceXAI is significantly wider. AA Briefcase and GDPval focus on the multi-step work that professionals complete every day, covering common tasks in positions such as lawyers, nurses, and financial analysts, as well as document and presentation production.

On AA Briefcase v1.1, Grok 4.7 scored 1657 points, higher than Grok 4.6's 1546 points; the GDPval Elo score increased from 1605 to 1695. In the official chart, it is close to Fable 5.1’s 1735 points on GDPval and higher than GPT-6 Astra’s 1542 points. SpaceXAI concluded that Grok 4.7 outperformed the previous generation in both tests and overall performed on par with other leading-edge models.


Promotions in the professional field also occur in many directions. The EEBench score increased from 53.0% to 64.0%, the Harvey Legal Agent Benchmark increased from 15.8% to 19.6%, and the HealthBench Professional increased from 48.5% to 56.7%. These results come from an internal comparison table published by SpaceXAI, which can illustrate the changes between the old and new generations of models; cross-model comparisons will still be affected by inference intensity, tool configuration and test environment.

This set of results reveals a clear trend: the competition for cutting-edge models is entering the "complete the whole job" stage.

The model needs to understand the task, call tools, produce files, and review itself. Single question and answer capability is only one part of it. Grok 4.7 extends the training task duration and strengthens self-verification, which is exactly what makes up for this type of workflow.

Safety and price

In addition to improved capabilities, Grok 4.7 has enabled a new security protection system. SpaceXAI said that this is the model that it has tested with the strongest performance in terms of denial of response and resistance to jailbreaking.

In terms of biosecurity testing, Grok 4.7 topped the LatchBio benchmark with a score of 62.4%. Cybersecurity testing HackerBench v0.3 mainly covers high-risk and malicious tasks. According to official data, the model only releases 3.3% of high-risk dual-use tips, and rarely blocks normal security research requests by mistake.

SpaceXAI has also begun opening invitation-only red team capabilities to a limited number of cybersecurity partners for defensive research. Currently, this red team capability is initially available as a controlled collaboration.

The standard version continues at the original price, and the speed of the fast version is doubled

Grok 4.7 has been launched on Cursor and Grok Build, and is provided through Grok API, third-party programming frameworks, model routing services and cloud platforms. The standard version starts at $2 per million input tokens and $6 per million output tokens, consistent with Grok 4.6.

The price strategy makes this upgrade more impactful. Developers do not need to pay additional standard version token costs for upgrading, but can obtain stronger long-term task performance, self-examination capabilities and professional work results.

For programming agents and knowledge work products, the unit price of each model call is of course important. The number of attempts and rework required to complete a task ultimately determines the true cost.

Finally summarize this upgrade:

From training methods to evaluation combinations, Grok 4.7 all reinforces the same thing: staying in tasks for a long time and completing complex tasks. Programming is only the first mature entry point. Documentation, presentation, legal analysis, clinical reasoning and engineering tasks have been put into the same set of capabilities.

But netizens don’t seem to buy it. “This model actually dares to be released. It’s so bad that it shouldn’t be released.” I can only say respectfully that the selling point of Grok 4.7 this time is not to completely crush competing products. It uses an API cost that is far lower than some flagship models in exchange for performance that is quite close and leading in individual professional tasks.


Reference link:

https://x.ai/news/grok-4-7

https://x.com/ArtificialAnlys/status/2102074904560771365

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet