OpenAI releases GPT-6.1 Sol for lower-cost agent workloads

OpenAI has introduced GPT-6.1 Sol, an update to GPT-6 Sol aimed at agentic coding, computer control and professional work. The company says the model approaches GPT-6 Astra on difficult tasks while standard API input and output token prices are five times lower than Astra’s. GPT-6.1 Sol is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and through the OpenAI API as gpt-6.1-sol.
API pricing is $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. OpenAI says the cached-input rate is 95% below the standard input price, which is intended to make repeated-context agent workloads less expensive.
Benchmarks across coding and professional tasks
In DeepSWE v1.1, which evaluates software-engineering tasks in real codebases, OpenAI says GPT-6.1 Sol reaches GPT-6 Astra’s level at roughly one-fifth of the cost. It also exceeds GPT-6 Sol’s best result by 6.4 percentage points while using less reasoning effort and incurring lower cost.
The release extends the Sol line introduced in OpenAI’s GPT-6 Sol and Luna launch with higher performance for coding and document-heavy work. On GDP.pdf, a benchmark of questions about complex professional PDF documents, OpenAI says GPT-6.1 Sol outperforms Opus 5.5 with fallback models at more than twice-lower task cost across tested reasoning-effort settings. It also approaches Astra’s leading result at about one-fifth of the task cost.
AutomationBench measures whether agents correctly complete multistep business processes. At medium reasoning effort, GPT-6.1 Sol exceeded Opus 5.5 by 2.2 percentage points at approximately one-third of the cost, OpenAI says. The result was also 4.8 percentage points higher than GPT-6 Sol at the same setting. The benchmark uses 47 tools across sales, marketing, operations, support, finance and HR.
Computer use, science and factual accuracy
On the offline OSWorld 2.0 set, GPT-6.1 Sol improved on GPT-6 Sol by seven percentage points at maximum reasoning effort while costing more than two times less. At that setting, it trailed Astra by 2.1 percentage points, with task cost roughly seven times lower. OSWorld assesses long-horizon computer-use workflows spanning everyday and professional tasks.
OpenAI also reported results from Terminal-Bench Science 0.1, covering data analysis, modelling and theorem proving with code and terminal tools. At maximum reasoning effort, GPT-6.1 Sol’s average task cost was $5.47, compared with $23.21 for Opus 5.5 and $23.80 for Astra. Astra retained the highest tested score, 68.1%, and OpenAI recommends it for the most difficult scientific research tasks.
For challenging anonymised ChatGPT conversations in which users identified an earlier factual error, GPT-6.1 Sol recorded a 4.1% factual-error rate at very high reasoning effort. OpenAI reported 4.5% for GPT-6 Sol and 4.0% for Astra under the same setting, while placing Sol’s task cost about 83% below Astra’s. The company notes that these deliberately error-inducing prompts do not represent ordinary use.
Availability and deployment considerations
OpenAI said GPT-6.1 Sol more reliably follows user intent and safety constraints in difficult evaluations, including cases involving failed search tools, explicit restrictions and unauthorised actions in agent tasks. It reported no attempts by the model to bypass its automated safety evaluation system. These tests are designed to probe difficult scenarios rather than measure failure rates in typical use.
Alongside Sol, OpenAI launched GPT-6 Astra Ultrafast and GPT-6.1 Sol Ultrafast. The company says they can run up to eight times faster than GPT-6 Astra and are available through the API, ChatGPT Work and Codex to users of a $500-per-month Pro tier with the highest limits. Businesses can use GPT-6.1 Sol to test cost-sensitive, repeated-context workflows first, reserving Astra for tasks where its highest reported capability is required.

