SpaceXAI released Grok 4.7 on Monday as a larger frontier model for coding and knowledge work, available immediately through Grok Build, Cursor, the Grok API and other platforms. The company kept standard pricing at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6. The price tag stayed still; the amount of work behind a completed task did not necessarily do the same.

The launch matters because Grok 4.7 is not just a chatbot refresh. SpaceXAI says it trained the model longer on harder, multi-hour tasks and taught it to work natively with the Grok Bot harness. Its published results show gains over Grok 4.6 on coding, terminal, office, legal and engineering evaluations. Those are company-selected figures, however, and the launch table compares Grok 4.7 at xhigh effort with Grok 4.6 at high effort on most rows. That is useful product evidence, not a neutral finish line.

Same unit price, more units

Artificial Analysis measured Grok 4.7 at 46 on its Intelligence Index, two points above Grok 4.6, and at 56 with Grok Build on its Coding Agent Index, nine points above the previous pairing. The coding system ranked fourth among the native model-and-harness combinations it tested. That is a meaningful improvement, especially for longer work where an agent must inspect, act, check and recover rather than answer once.

The same evaluation also supplies the invoice-shaped footnote. At xhigh effort, Grok 4.7 produced about 81,000 output tokens per Intelligence Index task, more than twice Grok 4.6’s roughly 38,000. It averaged 7.1 minutes per task. Because output tokens carry the higher price, identical per-token rates do not establish identical cost per accepted result. Benchmarks have rediscovered hardware-store economics: the screws can stay cheap while the house uses more of them.

The model is not locked to xhigh. Cursor documents low, medium, high and xhigh effort levels, with high as the default, plus a 256,000-token standard context window and a 500,000-token maximum. That is the strongest counterargument to the cost concern: buyers can choose a lower effort setting, cap budgets and decide that a higher completion rate is worth more tokens. Artificial Analysis’s xhigh run cannot tell every team what its own bill will be.

The wrapper changed the result

Early testing suggests the agent harness matters as much as the model name. Security-testing company XBOW found Grok 4.7 slightly worse than 4.6 in its existing exploit-crafting harness, but substantially better when paired with SpaceXAI’s Build-style orchestration. In one broader test, Build-based systems rose from 42 correct findings with 4.6 to 68 with 4.7. Outside Build, performance slipped slightly.

XBOW observed that Grok 4.7 favored shorter commands and more interactions, a behavior that fit Build’s rapid act-observe-adjust loop. It also saw an old failure—reasoning indefinitely without acting—in about 0.85% of Grok 4.6 runs and none of its 4.7 runs. These were controlled offensive-security tests of an early candidate, not proof of universal superiority. They do show why a model benchmark and a working agent are different products.

TINA’s view: price the finished job

TINA’s view: Grok 4.7 is a credible agent upgrade, particularly inside the system it was trained to use. SpaceXAI’s “same price” framing is technically accurate and economically incomplete. Teams should compare cost per accepted code change, successful investigation or finished document—not celebrate a token rate before counting the tokens, retries, tool calls and human review.

This judgment would soften if default-high evaluations show the same agent gains without the large output-token increase, or if portable tests reproduce the improvement across competing harnesses. It would harden if real workloads require xhigh effort and Build-specific orchestration to reach the launch results.

Watch independent high-versus-xhigh tests, completed-task cost, failure and retry rates, and performance when Grok 4.7 leaves its native wrapper. The most consequential number will not be the benchmark score or the price per million. It will be how much useful work survives at the end of the meter.