On 28 June 2026, Elon Musk announces on X that Grok 4.5, based on a new 1.5-trillion-parameter model, is being tested internally at SpaceX and Tesla with performance "close to, perhaps better than" Claude Opus. No public benchmarks, no technical spec sheet, no external access. Just a tweet. And yet, the story behind it is more interesting than the number itself.
There is a way of launching an AI model that has followed a well-established ritual for years: a detailed blog post, a system card, scores on recognised benchmarks, sometimes early access for independent testers. And then there is Elon Musk's way: a post on X, a parameter count, and a flattering comparison with the industry's most respected competitor. On 28 June 2026, it was the second method that was chosen for Grok 4.5.
What is confirmed
Musk's message amounts to one sentence: "Grok 4.5, based on our V9 foundation model with 1.5 trillion parameters, with Cursor data added in continued training, is now in private beta at SpaceX and Tesla." The underlying model, dubbed V9, finished training on 26 May 2026, with an impressive jump in size: roughly three times more parameters than the v8-small model currently used in production on X and in Teslas.
Notably, the training incorporated data from Cursor, an AI coding assistant very popular among developers, whose parent company was acquired by SpaceX for $60 billion. This acquisition closes a loop: xAI now owns the computing power (the Colossus supercomputer), the model, and the coding tool that generates part of the training data.
What remains unverifiable
This is where you need to slow down. The claim that Grok 4.5 "performs close to, or above, Opus" rests on no independent evaluation. The tests were carried out by engineers at SpaceX and Tesla, two companies that belong to the same group as xAI. There is literally no outsider in the room to challenge or confirm the result.
No recognised independent ranking — neither Artificial Analysis, nor LMArena, nor any of the standard public benchmarks we described in our article on benchmarks — has a score for Grok 4.5. There is not even public API access so an independent developer could test the comparison themselves. The comparison with Opus specifies neither which version of Opus, nor which test set, nor which methodology. It is an unverifiable claim about an inaccessible model.
xAI engineers themselves confirmed on X that the Cursor data was added in continued training, a phase that comes after the main training run, rather than being integrated from the start. One engineer involved in the project admitted that this method is "not quite as good as if the data had been present from the initial training". The next model, with 2 trillion parameters and already in training, is precisely designed to integrate this data from the outset. In other words, even by xAI's own account, the model being tested today has a structural limitation that its own successor is meant to fix.
The real story: vertical integration
What makes this story more interesting than a simple parameter count is the structure taking shape behind it. In a few months, SpaceX has merged its capital with xAI, acquired Cursor for $60 billion, and is now deploying a model trained on that coding tool's data directly within its own subsidiaries. A single shareholder now controls the computing power, the AI model, and the tool that generates part of its training data. No other frontier lab currently has such a vertically integrated loop.
This integration has a concrete consequence for any developer using Cursor today: the tool remains, for now, agnostic about which model is used, but the structural risk of future lock-in — where Cursor quietly favours its own parent company's models — now exists, even if it is not yet visible in the product.
What to take away
Musk has also announced an extremely aggressive cadence: a fully from-scratch trained new model every month until the end of 2026. If that promise holds, it would put unprecedented pace pressure on the entire industry, beyond just raw compute.
For now, the right attitude towards Grok 4.5 is the same as towards any self-announced figure: note it, stay curious, and wait until an independent third party can finally measure it on neutral ground. A model you can neither touch nor test is not a model you can compare. At best, it is a promise.