What happened
The family spans three tiers with a new naming system: Sol, the flagship,
at $5 input / $30 output per million tokens; Terra, a balanced model at
$2.50/$15 that OpenAI says matches GPT-5.5 at half the cost; and Luna, the
fastest and cheapest at $1/$6. The generation number and tier names are now
decoupled, so each tier can advance on its own cadence. Two new capability
settings ship alongside: max, which gives Sol more reasoning time than
xhigh, and ultra, which coordinates four agents in parallel by default
(sixteen in some configurations) for demanding tasks.
The rollout is the unusual part. In its June 26 preview announcement, OpenAI said that "as part of our ongoing engagement with the U.S. government, we previewed our plans and the models' capabilities ahead of today's launch. At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government." OpenAI added, pointedly, that it doesn't believe "this kind of government access process should become the long-term default," and framed the step as tied to work with the Administration on a cyber Executive Order framework and "a repeatable process for future model releases."
The benchmark claims are aggressive. OpenAI reports Sol at 53.6 on Agents' Last Exam (long-horizon professional workflows), 13.1 points above Claude Fable 5; a state-of-the-art 80 on the Artificial Analysis Coding Agent Index using less than half the output tokens of Fable 5; 92.2% on BrowseComp with ultra; and 62.6% on OSWorld 2.0, beating Opus 4.8 while using 85% fewer output tokens. Cybersecurity is the standout jump: 73.5% on ExploitBench versus GPT-5.5's 47.9%, and more than double GPT-5.5's pass rate on ExploitGym, a UC Berkeley-built benchmark that asks agents to turn real vulnerabilities into working exploits.
That capability is why the safety story is elaborate. OpenAI says GPT-5.6 does not cross its Preparedness Framework's Critical threshold in cyber or biology — in Chromium and Firefox testing it found bugs and exploitation primitives but never autonomously produced a full-chain exploit. Safeguards now block roughly ten times more potentially harmful activity than on previous models, layered as trained-in refusals, real-time classifiers, a reasoning monitor that reviews flagged conversations, and account-level enforcement. The company spent about 700,000 A100-equivalent GPU hours on automated red teaming hunting universal jailbreaks. The most capable cyber features sit behind a Trusted Access program requiring identity verification — and, from September 1, hardware-backed passkeys.
Why it matters
The episode is a visible example of government involvement shaping the timing and audience of a frontier-model release. OpenAI has not described the preview as regulatory approval, and no formal review regime exists — which is exactly the point. Voluntary coordination is defining, in real time, what a future mandatory process might look like, and OpenAI is on record wanting it to stay short-term.
For developers, the practical story is efficiency: OpenAI's pitch throughout is more work per token, with customer quotes citing 22–48% fewer tokens or tool calls on real workloads. Choosing a model is increasingly about matching capability, risk posture, speed, and price to a job — not defaulting to the biggest thing on the menu.
The fine print
All benchmarks are OpenAI's; many use internal or affiliated evaluations, and the cost and latency figures are simulated. Some comparisons exclude Anthropic models; for example, Fable 5 refuses most questions on OpenAI's biology benchmark. OpenAI also says stricter safeguards may block legitimate dual-use work and offers lower-capability model retries when that happens.
The government asked for a closer look before GPT-5.6 reached everyone. OpenAI complied, shipped thirteen days later, and politely noted that it would prefer not to make a habit of this.