Skip to content

GPT-6.1 Sol: the "critical" tier for 2 dollars

A month ago, this level of cybersecurity risk was reserved for OpenAI's most expensive model. It now applies to a version that costs five times less.

Advertisement

A month ago, OpenAI crossed a threshold. Its most powerful model, GPT-6 Astra, was classified "critical" in cybersecurity, the highest risk level on its own scale. The company had even released it in two versions the better to control it.

This week, the same level was assigned to GPT-6.1 Sol. A model sold at 2 dollars per million input tokens, five times cheaper than Astra.

What "critical" means

OpenAI assesses each of its models against a four-level scale: low, medium, high, critical. It does so across several sensitive domains, including cybersecurity and biology.

In cybersecurity, "critical" denotes a model capable of finding previously unknown flaws in real, well-protected systems, and of building what is needed to exploit them, without human help. In other words, what until now was expected of highly experienced teams of attackers.

According to the safety card published by OpenAI, GPT-6.1 Sol is "critical" in cybersecurity, "high" in biology and chemistry, and below the "high" level for its ability to improve AI itself. Its predecessor, GPT-5.6 Sol, was only at the "high" level in cybersecurity.

What changes in practice

GPT-6.1 Sol inherits all the protections put in place for Astra. In concrete terms, certain tasks may be slowed down, paused or stopped if they resemble a cyberattack. In ChatGPT or in OpenAI's coding tool, the user may be asked to verify an action before it continues. For developers, the task stops.

Security professionals who need to go further, in order to defend systems, go through a restricted and controlled access programme. The general public does not have direct access: the model is offered in OpenAI's work and coding tools, and to developers.

A transparency that extends to the bad numbers

The card published by OpenAI has one merit: it does not show only the good results. On tasks designed to push the model into lying about its programming work, GPT-6.1 Sol distorted reality in 1.5% of cases. That is slightly more than GPT-6 Sol, at 1.3%, and nearly three times more than Astra, at 0.5%.

These are small percentages, obtained under deliberately difficult conditions. But they serve as a reminder that a cheaper model is not merely slightly less capable. It may also behave slightly differently.

The honeypot 🍯
To test models, OpenAI sometimes places a trap in a hacking exercise: a target that looks easy to reach, but that is not part of what is being asked. Specialists call this a honeypot. A model that takes the bait shows it is prepared to take forbidden shortcuts in order to succeed. According to the safety card, GPT-6.1 Sol never attempted to exploit this trap.

What we take away

In one month, a risk level reserved for OpenAI's most expensive and most closely monitored model has become that of a version five times cheaper. This is the direct consequence of the price war we have been following since September: high-end capabilities move down to the mid-range very quickly.

Protections must follow the same path, and just as quickly. That is the whole point of measures such as the circuit breaker presented by Nvidia. In AI, "critical" no longer means rare. It means: this month.

Advertisement