GPT-6.1 Sol: Near-Astra at One-Fifth the Cost, How Close?

An AI agent that costs less to run can still be expensive to finish with. A cheap refactor becomes a costly afternoon when someone has to untangle its mistakes.

GPT-6.1 Sol makes that tradeoff more interesting. OpenAI’s September 29 release targets complex coding, computer use and sustained work across apps. The company reports performance approaching Astra on several evaluations, sometimes at roughly one-fifth of the cost per task. Standard API rates are $2 per million input tokens and $10 per million output tokens, with cached reads at $0.10.

The headline has substance. It also has boundaries. “Near-Astra” describes particular tasks, reasoning settings and evaluation environments. It doesn’t mean equal judgment everywhere, or an automatic 80% reduction in your bill.

For builders, the question is whether Sol completes their workload without expensive retries, supervision or escalation.

1. GPT-6.1 Sol Benchmarks: The Full Launch Evaluation Set

OpenAI’s launch covers five performance evaluations. The table summarizes every plotted Sol effort setting as score percentage / dollars per task. These approximate chart readings are rounded to whole percentage points and roughly $0.05. Exact reported figures below take precedence.

GPT-6.1 Sol Benchmark Performance by Reasoning Effort and Cost

EvaluationLowMediumHighExtra HighMax
DeepSWE v1.1≈65 / $0.15≈73 / $0.40≈75 / $0.65≈72 / $0.80≈72 / $1.55
GDP.pdf≈27 / $0.35≈30 / $0.35≈32 / $0.35≈32 / $0.35≈31 / $0.40
AutomationBench v1.0.6≈25 / $0.15≈32 / $0.20≈33 / $0.20≈36 / $0.25≈36 / $0.30
OSWorld 2.0, Offline Partial Reward≈59 / $0.40≈67 / $0.75≈70 / $0.95≈69 / $1.05≈71 / $1.25
Terminal-Bench Science v0.1≈44 / $1.75≈48 / $2.30≈51 / $2.75≈54 / $2.85≈57 / $5.45

On DeepSWE, which tests demanding repository work, High reaches an explicitly reported 75.2%. The previous Sol’s best result was 68.8% at Max. That’s a 6.4 percentage-point improvement at approximately 76% lower task cost. Notice that additional effort doesn’t improve the new model’s result.

On AutomationBench, Medium scores 31.7%, versus 26.9% for its predecessor. That’s progress, but plenty of these difficult workflows remain unsolved.

On OSWorld, Max reaches 71.4%, compared with Astra’s 73.5% and the previous Sol’s 64.4%. Its task cost is roughly one-seventh of Astra’s. This is a partial-reward evaluation on an offline set, not a claim that 71.4% of arbitrary desktop jobs finish successfully.

For GDP.pdf, OpenAI reports results approaching Astra at roughly one-fifth of the cost, and exceeding Opus 5.5 with fallbacks across tested efforts. On Terminal-Bench Science, Sol more than doubles its predecessor’s Max score. Reported task costs are $5.47 for Sol, $23.21 for Opus and $23.80 for Astra. Astra’s 68.1% score still leaves a meaningful science capability gap. Introducing GPT-6.1 Sol _ OpenA…

These percentages measure different things. Averaging them into one “overall success rate” would produce a tidy number with little meaning.

2. GPT-6.1 Sol Pricing: The Complete API Cost Picture

The rates below are in US dollars per million tokens unless another unit is specified. Cache writes and reads are separate billing categories. Output pricing includes reasoning tokens.

GPT-6.1 Sol Pricing for Standard, Fast, Batch, Flex and Tool Usage

Processing or ChargeInputCache ReadCache WriteOutput or Other Charge
Standard, Up to 272K Input$2$0.10$2.50$10
Standard, Above 272K Input$4$0.20$5$15
Fast, Up to 272K Input$4$0.20$5$20
Fast, Above 272K Input$8$0.40$10$30
Batch or Flex, Up to 272K Input$1$0.05$1.25$5
Batch or Flex, Above 272K Input$2$0.10$2.50$7.50
Eligible Regional Processing+10%+10%+10%+10%
Web SearchModel Rates for ContentN/AN/A$10 per 1,000 Calls
File Search Calls, Responses APIModel Token RatesN/AN/A$2.50 per 1,000 Calls
File Search StorageN/AN/AN/A$0.10 per GB/Day, First GB Free
Hosted Shell or Code Interpreter ContainersN/AN/AN/APer 20 Minutes: 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92

Fast doubles Standard token rates. Batch and Flex halve them. Regional premiums apply where available, and Fast is unavailable with EU data residency. The long-context multiplier affects the full request, including output. OpenAI API

Eligible container sessions use per-minute billing with a five-minute minimum. Search content also consumes billed tokens, so the tool-call fee isn’t the entire cost. Image generation and other separately priced tools can add charges beyond this table’s text, search and execution costs. OpenAI API

For a simple illustration, 10,000 ordinary input tokens and 2,000 output tokens cost $0.04, excluding cache writes and tools. A chatty agent can quickly spend more on output than on reading your instructions.

3. What Changed From GPT-6 Sol?

The new release followed GPT-6 Sol by just seven days. That pace invites questions about whether this is merely a relabelled model.

The documented answer is narrower: the evaluations show stronger repository coding, automation and computer use, alongside cheaper cached input. A cache read costs half the previous Sol rate. The system card describes a training approach related to Astra’s, but doesn’t establish a public parameter count or prove a particular distillation story.

For GPT-6.1 Sol vs GPT-6 Sol, the clearest upgrade evidence is completing harder agent tasks at lower measured cost. For users still on GPT-5.6 Sol, the practical comparison should include their existing prompts and tools rather than assuming every generation responds identically.

This is a work-oriented reasoning model. It accepts text and images and produces text. Access to an image-generation tool doesn’t make image generation a native output modality. OpenAI API

4. Access Through Work, Codex And The API

Launch access includes Plus, Pro, Business, Enterprise and Edu through ChatGPT Work and Codex. Free and Go aren’t included. Enterprise and Edu administrators must enable the model because it starts disabled for those plans.

It isn’t available in ordinary Chat. If you’re checking that picker, its absence doesn’t necessarily indicate an account problem. Availability also depends on client rollout and workspace controls.

For GPT-6.1 Sol Codex access, select the model in the available picker or run codex -m gpt-6.1-sol. Check saved defaults as well as the current session. Choosing a model cannot override your workspace’s permissions. ChatGPT Learn

API developers use gpt-6.1-sol. OpenAI directs tool-calling workflows to the Responses API. Chat Completions is supported without tool calling. Keep API token charges separate from subscription credits and allowances, which don’t translate into a fixed number of completed tasks.

5. Why Independent Benchmarks Tell A Different Story

Artificial Analysis reports an Intelligence Index score of 52 at Max, one point behind Astra and four ahead of the previous Sol. Its evaluation costs $0.72 per task for the new model, compared with $3.26 for Astra.

Its coding results are especially instructive. Max improves on the previous Sol but remains behind Astra. Extra High performs better than Max on the Coding Agent Index, showing that the largest reasoning budget isn’t automatically the best setting. Output consumption also increased relative to the previous Sol. Artificial Analysis

None of this contradicts OpenAI’s narrower claims. Different suites use different tasks, tool environments, scoring rules and model configurations. A benchmark name alone isn’t enough to establish comparability.

BridgeBench adds another distinction. Its published methodology includes editorial assessments across several dimensions. Those judgments can help readers inspect practical output, particularly design, but should not be treated as controlled measurements equivalent to repository tests. BridgeMind

Community screenshots are useful leads. Without prompts, tool traces and repeated trials, they remain examples rather than rankings.

6. Sol Vs Astra, Sonnet 5.5 And Opus 5.5

The GPT-6.1 Sol vs Astra decision depends on how expensive a mistake is. Repository changes and desktop work show a compelling cost-quality tradeoff. Difficult science tasks preserve a larger gap. A business-critical analysis may justify Astra even when Sol is close on an aggregate index.

For GPT-6.1 Sol vs Claude Sonnet 5.5, Artificial Analysis’s broader Intelligence Index favours Sonnet: 56 versus Sol’s 52. Opus scores 58. Those rankings don’t settle every coding or design comparison. Tools, instructions and repository knowledge influence agent results, so test with matching acceptance criteria and budgets. Artificial Analysis

The GPT-6.1 Sol vs Claude Opus 5.5 comparison has more launch evidence. OpenAI reports favourable PDF results and lower science task costs, but its charts don’t establish that Sol wins every workload. Preserve fallback labels when discussing Opus results. A model with fallback behaviour is a different evaluated configuration.

Code review and frontend design deserve separate trials. Finding a subtle security bug, repairing a repository and producing an attractive interface are different skills. There isn’t a complete, controlled comparison for every combination here. Mark those gaps rather than filling them with confident guesses.

7. What The 95% Cache Discount Actually Saves

Bar comparison showing GPT-6.1 Sol prompt caching cutting a $5.00 workload to $1.44, about 71% saved
Bar comparison showing GPT-6.1 Sol prompt caching cutting a $5.00 workload to $1.44, about 71% saved

The discount applies to eligible cached input reads, not your whole bill. Writes cost $2.50 per million tokens, while reads cost $0.10. Output and tool charges remain separate.

Caching requires a matching prefix and at least 1,024 input tokens. Stable instructions and shared reference material belong early in the prompt, with changing task details later. Inspect reported cache usage instead of assuming repeated material always produces a hit. OpenAI API

Consider a simplified agent workload: a 100,000-token prefix reused across 20 requests, each generating 5,000 output tokens. One cache write costs $0.25, nineteen reads cost $0.19, and output costs $1. That totals $1.44, excluding new task inputs and tools.

Without caching, the same prefix reads and output would cost $5. The saving is approximately 71% overall, despite the 95% read discount. The result depends on actual cache hits and reuse, but it explains why sustained agents benefit more than unrelated short requests.

8. Context Capacity And The 272K Pricing Threshold

Step chart showing GPT-6.1 Sol cost jumping from $0.744 to $1.392 when input crosses 272K tokens
Step chart showing GPT-6.1 Sol cost jumping from $0.744 to $1.392 when input crosses 272K tokens

The GPT-6.1 Sol context window is 1,050,000 tokens, with up to 128,000 output tokens. Those limits describe API capacity. A client can impose smaller working limits or compact a conversation sooner.

Capacity and usable recall are separate questions. A model may accept a huge repository yet overlook the file that determines the right answer. Retrieval, focused file selection and concise instructions still matter.

The pricing boundary creates a particularly sharp budgeting issue. With 20,000 output tokens, an uncached 272,000-token input costs $0.744 at Standard rates. Increase input to 273,000 tokens and the full request costs $1.392. These examples exclude cache writes and tools.

That extra thousand tokens nearly doubles this illustrative bill because every token crosses into the higher tier. Compaction can reduce cost and latency, but a summary may lose requirements or evidence. Preserve decisions, constraints and source references explicitly rather than treating compaction as lossless storage.

9. Reasoning Effort, Speed And Completion Time

The API supports Low, Medium, High, Extra High and Max. Medium is the default, while None and Minimal aren’t supported. Work and Codex controls may use different labels depending on the client.

Start at the default, then increase effort when a task needs deeper planning or verification. The launch curves and independent coding results both warn against choosing Max reflexively. More reasoning can mean more tokens without a better accepted result.

Measure elapsed time to usable output, including tests, tool calls and corrections. Tokens per second tells you how fast text arrives. It doesn’t tell you how quickly a refactor becomes safe to merge.

Standard and Fast are available at launch. Ultrafast support is forthcoming, so don’t plan today’s workflow around it. Fast can be worthwhile for interactive work, but its doubled token rate needs to buy meaningful time savings.

10. Factuality And Agent Reliability Still Need Scrutiny

On deliberately difficult factuality prompts, Low-effort responses containing errors fell from 11.4% to 7.7% compared with the previous Sol. That’s roughly a 32% relative reduction, not proof of 92.3% accuracy on everyday questions.

The system card gives a mixed reliability picture. Broken-search nondisclosure improves, but coding misrepresentation is 1.50%, compared with 1.30% for the previous Sol and 0.51% for Astra in the elicitation evaluation. A separate warning-respect test reports unwanted persistence of 23.5%, versus Astra’s 17.4%, without system-level control measures.

These are targeted evaluation results, not production incident probabilities. They still support checking whether an agent actually ran tests, disclosed tool failures and respected restrictions. OpenAI observed no reviewer-bypass attempts in its automated safety review evaluation. A finite test cannot establish that bypasses are impossible. gpt-6-1-sol system card

Claims that a model has been “nerfed” need comparable evidence. Preserve model identifiers, effort, prompts and tool conditions before attributing a disappointing session to a hidden capability change.

11. Should You Switch? Measure Cost Per Accepted Result

Start with a representative set of real tasks: a refactor, a difficult bug, a review, a PDF deliverable and an app workflow. Define success before running the models.

Record four things:

  • Whether the result meets your acceptance criteria.
  • Total token and tool spending, including retries.
  • Human review and correction time.
  • Time until the result is usable.

For coding, passing tests is necessary but may miss regressions, maintainability problems or unsupported claims. Review the diff and verify important behaviour independently.

Calculate cost per accepted result by dividing total evaluation spending by accepted completions. Add labour cost when it matters to your decision. Include migration effort and subscription allowances separately.

A sensible starting policy is Sol for repeated complex work, with Astra available for unresolved or unusually demanding tasks. Set a clear escalation point. Letting a cheaper model retry indefinitely is an excellent way to spend more slowly.

12. A Better Default, If Your Work Confirms It

GPT-6.1 Sol makes a credible case for running demanding agents more often. Its strongest evidence combines better repository performance, competitive computer use and substantially cheaper repeated context.

The gaps matter. Astra retains advantages on harder work, cache savings depend on usage, and long-context pricing can erase an efficiency gain. Your accepted results should decide the switch.

Choose one recurring workflow, test both models under matching conditions, and compare the complete bill. Follow Binary Verse AI for evidence-led model comparisons that connect benchmark claims to the work you actually need finished.

What is GPT-6.1 Sol?

GPT-6.1 Sol is OpenAI’s reasoning model for complex coding, computer use and professional work at a lower cost than GPT-6 Astra. It upgrades GPT-6 Sol and supports text and image inputs, a 1.05-million-token API context window and up to 128,000 output tokens.

Is GPT-6.1 Sol as good as GPT-6 Astra?

It approaches or matches Astra on selected evaluations, but does not establish equal performance across all work. OpenAI reports comparable DeepSWE performance at substantially lower cost, while Astra remains ahead on difficult scientific workflows. Artificial Analysis places Sol one point below Astra on its Intelligence Index.

Is GPT-6.1 Sol better than Claude Sonnet 5.5 or Opus 5.5?

There is no universal winner. Artificial Analysis reports Intelligence Index scores of 52 for Sol, 56 for Sonnet 5.5 and 58 for Opus 5.5 with default fallback at the reported configurations. Sol offers strong cost efficiency, while selected OpenAI evaluations favor it on PDF and business tasks. Compare the specific workload and configuration you will use.

How much does GPT-6.1 Sol cost, and does caching make it 95% cheaper?

Standard API rates per million tokens are $2 for input, $0.10 for cached input, $2.50 for cache writes and $10 for output. The 95% discount applies to cache reads compared with ordinary input—not the whole bill. Above 272,000 input tokens, input/cache rates double and output rises to $15 for the full request.

Can I use GPT-6.1 Sol with ChatGPT Plus? Why is it missing from Chat?

Yes. Launch access includes Plus, Pro, Business, Enterprise and Edu through ChatGPT Work and Codex, with availability depending on client rollout and workspace settings. GPT-6.1 Sol is not yet available in ordinary Chat. Enterprise and Edu administrators must enable it, and Free and Go are excluded at launch. Developers can also access it through the API.

Leave a Comment