Tag

Models

5 briefings.

OpenAI ExploitBench chart reports Astra at 100% capability coverage and GPT-5.6 Sol at 78.5%, plotted against output tokens.
Research chart© OpenAIEditorial excerptOriginal storyOpenAI’s ExploitBench results without production safeguards, from Astra’s safety report. Company-reported evaluation; not an independent test or the unrestricted capability of the public release.

openai · models

OpenAI starts GPT-6 Astra rollout with critical cyber capability gated

OpenAI's September 5 Astra launch combines a new computer-use flagship with restricted exploit generation and a staged expansion through Daybreak.

4 sources
Google’s HLE-Verified chart compares Gemini 3.8 Flash at 54.9% with earlier Flash and competing models.
Research chart© Google DeepMindEditorial excerptOriginal storyGoogle’s HLE-Verified comparison for Gemini 3.8 Flash, from its model page. Vendor-reported reasoning scores do not measure reliability on a complete agent workflow.

google · gemini

Gemini 3.8 Flash moves Google's workhorse model toward long-running agents

Google's newest Flash model combines a one-million-token input window, multimodal input, computer use, and a larger emphasis on sustained agent work.

2 sources