All posts

GPT-6.1 Sol review with four hands on tests

GPT-6.1 Sol handled the core requirements in our four tests, from Python edge cases to an impossible schedule. This hands-on review examines the results, pricing and vendor benchmarks.

10 min read
On this page

GPT-6.1 Sol from OpenAI

GPT-6.1 Sol makes a strong case for becoming your everyday work model. We ran four practical tests at Medium effort: a misleading revenue claim, a Python function with awkward edge cases, a refund reply and an impossible schedule. It handled the central requirement in every run. The standard input and output rates also match GPT-6 Sol, making this an upgrade worth testing if you already use Sol.

The useful question is whether that translates into work you can trust enough to review quickly. Here are the prompts, actual responses and our observations.

GPT-6.1 Sol model card

What changed and what it costs

OpenAI released GPT-6.1 Sol on September 29, 2026, positioning it close to Astra on several evaluations at a much lower price. The launch covers Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Official announcement.

Standard API rate per million tokensGPT-6 SolGPT-6.1 SolGPT-6 Astra
Input$2.00$2.00$10.00
Cached input$0.20$0.10$1.00
Output$10.00$10.00$50.00

The ordinary token rates are one fifth of Astra’s. Repeated cached input is also cheaper than on GPT-6 Sol. That does not guarantee a particular saving per completed task: reasoning length, retries and tool use still affect the bill.

For developers, the model ID is gpt-6.1-sol. It supports a 1,050,000-token context window and up to 128,000 output tokens. Medium is the default reasoning effort. Tool calling uses the Responses API. Model documentation.

For the earlier model choices, see our GPT-6 Sol vs Astra comparison and GPT-6 Sol vs Luna comparison.

How we ran the four tests

We selected GPT-6.1 Sol in Work for the tests. Each test used a fresh conversation, Medium effort and the first completed response. We requested no browsing or tools. We then checked the arithmetic, executed the generated Python separately and verified the policy and scheduling constraints.

TestWhat we checkedObserved result
Funnel analysisArithmetic and unsupported causal claimsCorrect calculations and cautious recommendation
Python deduplicationTies, missing dates, ordering and input mutationSupplied asserts and 106 additional cases passed
Customer supportPolicy boundaries and missing informationAsked for purchase date; promised no action
SchedulingInfeasibility and the smallest permitted repairCorrect proof and 10-minute repair

Test 1: Revenue growth can hide a weaker funnel

We gave Sol a SaaS funnel where revenue increased while paid-search conversion deteriorated. The company wanted to double its ad budget. The challenge was to calculate the numbers without accepting the company’s conclusion.

Prompt

This is a hands-on editorial test of GPT-6.1 Sol. Do not browse or read local files. Answer using only these fictional figures. Keep the response under 250 words.

A SaaS company says: “Revenue rose, so our acquisition strategy is working. Should we double paid-search spending next month?”

Channel | August leads | August customers | August revenue | September leads | September customers | September revenue  
Paid search | 1000 | 100 | $10000 | 2400 | 144 | $14400  
Referrals | 500 | 150 | $15000 | 400 | 140 | $14000

Calculate overall and channel-level conversion rates for both months, total revenue growth, and the percentage-point change in overall conversion. Give a recommendation in three sentences, separating what the data proves from what is missing. Do not invent spend, profit, retention, or causal explanations.

Response

GPT-6.1 Sol doing social media marketing

Observation: Sol calculated paid-search conversion at 10% then 6%, referrals at 30% then 35%, and overall conversion at 16.67% then 10.14%. Revenue grew 13.6%, while overall conversion dropped 6.52 percentage points. Its three-sentence recommendation separated the observed growth from missing acquisition economics and causal evidence.

Verdict: A useful first-pass analysis. The model challenged the premise instead of producing a confident budget recommendation. The figures alone cannot tell us whether the business should spend more.

Test 2: A small Python function with real edge cases

Deduplication looks easy until equal timestamps, missing dates and output order interact. We asked for a function that follows all three rules and leaves the input untouched.

Prompt

This is a hands-on editorial test. Do not browse, read files, or execute tools. Write a Python 3 function latest_by_id(records). Each record is a dictionary with an id and an optional updated_at in canonical UTC ISO format YYYY-MM-DDTHH:MM:SSZ. Keep the newest record for each id; missing or None timestamps are older than any real timestamp. Equal timestamps, including two missing timestamps, must keep the LAST input record. Return winning records in the order their ids FIRST appeared. Do not mutate the input or its dictionaries. Empty input returns []. No external packages.

Include three assert tests covering timestamp ties, missing timestamps and first-seen ordering. Give the function and asserts in one Python code block, followed by an explanation under 80 words. Do not introduce extra validation requirements.

Response

GPT-6.1 Sol doing code review

Observation: Sol used an insertion-ordered dictionary and >= to make the last equal-timestamp record win. It treated missing timestamps as older than real timestamps. The three generated asserts passed. We also ran 106 independent cases covering empty input, replacement order, ties, missing values and mixed records; the input stayed unchanged throughout.

Verdict: Correct for the specified contract, with a short explanation of why it works. The timestamp comparison relies on the canonical UTC format in the prompt. Accepting arbitrary offsets or malformed dates would require a different specification and fresh validation.

Test 3: A refund reply that stays inside policy

This test mixed an angry customer, an invoice date that could be mistaken for a purchase date, nonrefundable credits and an account deletion request. We wanted a usable reply without invented eligibility or promises.

Prompt

This is a hands-on editorial test using a fictional support case. Do not browse or use tools. Write a customer reply under 120 words, then a separate internal note under 60 words.

Policy: monthly plans can qualify for a refund within 14 days of the original purchase; annual plans within 30 days of the original purchase. Used API credits are nonrefundable. Refunds require review; agents cannot promise approval or issue money. Account deletion is a separate process and requires identity verification.

Today: September 30, 2026. Customer: annual plan, $240 charge, latest invoice September 1, last login September 29, $18 of used API credits. The original purchase date is NOT provided; an invoice date is not proof of it. Customer says: “I barely used it. Give me every dollar back and delete my account now.”

Do not infer refund eligibility from login or invoice dates. Ask for the missing information, explain the used-credit exclusion, and avoid claiming that a refund or deletion has happened.

Response

GPT-6.1 Sol doing data analysis

Observation: The customer reply stayed at 82 words and the internal note at 46. Sol asked for the original purchase date, stated the annual plan’s review window, excluded used credits and kept deletion subject to identity verification. It did not claim that either action had happened.

Verdict: The policy handling passed. The tone needs editing: “I can’t promise approval or issue money” sounds more like an internal rule than polished customer support. We would soften that sentence before sending. The case also does not establish a specific refundable dollar amount.

Test 4: Catching an impossible schedule before fixing it

We made both original meeting days infeasible, then allowed exactly one repair: extending Faro’s Tuesday availability. A convincing-looking calendar was not enough; Sol had to prove the conflict and find the minimum change.

Prompt

Do not browse or use tools. All times are local. One room, one meeting at a time, exactly 10 minutes of buffer between consecutive meetings. Dana must be before Eli, and Faro must be after Dana.

Dana: 40 minutes; Tuesday 10:00–12:00 or Thursday 14:00–16:00.  
Eli: 30 minutes; Tuesday 11:00–13:00 or Thursday 15:00–17:00.  
Faro: 50 minutes; Tuesday 10:00–11:30 or Thursday 14:00–15:30.  
All three meetings must happen on the same day and finish by Tuesday 12:30 or Thursday 16:00.

Find the earliest-finishing valid schedule, or prove that none exists. If none exists, change ONLY the end of Faro’s Tuesday availability by the smallest number of minutes necessary, then give the repaired schedule. Keep the answer under 250 words. Do not silently relax any other constraint.

Response

GPT-6.1 Sol fixing race condition code issues

Observation: Sol noticed that Dana must go first. With the required buffer, Faro’s earliest finish is 11:40 on Tuesday or 15:40 on Thursday, ten minutes past the respective availability windows. It extended only Tuesday’s cutoff to 11:40 and scheduled Dana 10:00–10:40, Faro 10:50–11:40 and Eli 11:50–12:20. It also explained why 12:20 is the earliest possible finish.

Verdict: This was the strongest reasoning result of the four. It found the original impossibility, made the permitted repair and preserved the meeting durations, order, buffers and final deadline.

Sol 6.1 Benchmarks

OpenAI’s evaluations report improvements in coding, PDF work and factuality. The figures below use matching reasoning effort within each chart. They are vendor-reported results, separate from our four Work tests. Release page and methodology notes.

Coding at Medium effort

GPT-6.1 Sol coding benchmarks

DeepSWE v1.1, Medium effort. Sol 6.1 reaches 73.0% at $0.42 per task; Astra reaches 72.8% at $3.08.

Factuality at Low effort

GPT-6.1 Sol factuality error rate

OpenAI reports error falling from 11.4% to 7.7% at Low effort on its adversarial factuality evaluation. This is not an everyday error rate.

Should you switch to GPT-6.1 Sol

Absolutely! Regardless of which model of OpenAI you were using previously, moving over to GPT-6.1 Sol has clear benefits.

It offers a performance comparable to GPT-6 Astra at a fraction of the cost. If GPT-6 Sol already handles your everyday tasks, GPT-6.1 Sol is an obvious upgrade. Our four runs produced correct core results across analysis, code, support and scheduling. The support test also showed why a technically compliant response can still need a human edit.

Switching to GPT-6.1 Sol

Try it on a small set of work you can independently judge: a calculation with a known answer, code with awkward inputs, a policy reply and a task with conflicting constraints. Keep the prompt and effort fixed, then compare correctness, cleanup time and total cost. Use the older Sol vs Astra and Sol vs Luna articles as background for where those models fit.

Our first impression is straightforward: Sol 6.1 handled these four jobs well, and the unchanged standard Sol pricing makes it easy to justify.

Frequently Asked Questions

What is GPT-6.1 Sol?

GPT-6.1 Sol is an upgraded OpenAI work model offering near-Astra intelligence across coding, analysis, and reasoning at a significantly lower cost.

How much does GPT-6.1 Sol cost?

Standard API rates match GPT-6 Sol at $2.00 per million input tokens and $10.00 per million output tokens, with cheaper cached input rates.

What is the context window for GPT-6.1 Sol?

It features a 1,050,000-token context window and supports up to 128,000 output tokens per request.

Who can access GPT-6.1 Sol?

It is available in Work and Codex for Plus, Pro, Business, Enterprise, and Edu tier users.

Last updated: Sep 30, 2026

Build your agent team in 30 seconds.

Build agent teams that work along with your team. Free to start, no card required.