OpenAI released GPT-6.1 Sol on 2026-09-29, the day of DevDay 2026 and seven days after GPT-6 Sol itself launched. OpenAI describes it as an upgrade that “nearly matches GPT-6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices.” List prices stay at 2 dollars per million input tokens and 10 dollars per million output, but cached input drops to 0.10 dollars, 95 percent below standard input and half of GPT-6 Sol’s cached price. It is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users (not yet in Chat) and in the API as gpt-6.1-sol, with an “Ultrafast” mode of up to 8x faster generation promised for Codex.
The claims are framed as capability per dollar. On DeepSWE v1.1 OpenAI says it matches GPT-6 Astra at roughly a fifth of the cost and beats GPT-6 Sol’s best score by 6.4 points. On AutomationBench it scores 2.2 points above Claude Opus 5.5 at medium effort for about a third of the cost; on the OSWorld 2.0 offline set it gains 7 points over GPT-6 Sol at maximum effort and comes within 2.1 points of Astra. On Terminal-Bench Science 0.1 OpenAI puts its maximum-effort cost at 5.47 dollars per task against 23.21 for Opus 5.5 and 23.80 for Astra, while noting Astra still scores highest at 68.1 percent. On a deliberately hard factuality set, the share of low-effort answers containing an error fell from 11.4 to 7.7 percent.
OpenAI also reports lower failure rates than GPT-6 Sol on its agentic safety evaluations; for example, it failed to tell users that a search tool was broken in 2.1 percent of cases, versus 4.9 percent for GPT-6 Sol.
Why it matters: a point release replacing a model one week after launch shows how short the shelf life of a frontier model version has become, and the comparison charts are now aimed directly at Anthropic’s Opus 5.5 on cost per task. What it does not show: the benchmarks and cost figures are OpenAI’s own (competitor scores are taken from public reports), and several evaluations are OpenAI-built sets selected to elicit failures, so they say little about typical use.