🔍 Read the full analysis: The September 2026 AI Stack Behind My Daily Workflow on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
GPT-6.1 Sol launched on 29 September 2026, scoring 51 on the Artificial Analysis Intelligence Index v4.3.x at roughly $0.39 per task — far below rivals at similar scores. Analyst Thorsten Meyer now runs Claude Opus 5.5 for building and Sol for review, arguing the frontier has become a price curve rather than a capability race.
OpenAI’s GPT-6.1 Sol, released on 29 September 2026, has turned the frontier AI market into a price curve, according to analyst Thorsten Meyer: six leading models now sit within roughly 20 index points of each other on the Artificial Analysis Intelligence Index v4.3.x, while their cost per task differs by about 100 times. Meyer, writing on ThorstenMeyerAI.com, says the practical question has shifted from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?”
Meyer’s current stack pairs Claude Opus 5.5 (released 22 September, index score 58 at max, $5.98 per task at that setting) as the main builder with GPT-6.1 Sol as a low-cost reviewer and detail model. Sol scores 51 at its xhigh setting for $0.39 per task — about one-eighth of GPT-6 Astra’s cost and one-twentieth of Claude Fable 5.1’s, for a score only 1 to 2 points lower, according to Meyer’s figures, all drawn from the Artificial Analysis Intelligence Index v4.3.x.
Meyer highlights three findings from the index data. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task. Second, Sonnet 5.5 at maximum effort costs more per task than Opus at maximum for 2 fewer points. Third, Sol’s pricing undercuts every model in its score tier.
The effort setting, not model choice, is the largest cost lever, Meyer reports. On Opus 5.5, moving from xhigh to max adds 2 index points but 73% more cost per task; from medium to max, cost rises 4.46 times for 7 points. Meyer runs Opus at high (54 points, $1.82 per task) or xhigh (56 points, $3.46) and reserves max for rare cases. Sonnet 5.5 illustrates the ceiling: at max it produces about 193k output tokens per task — the most Artificial Analysis has measured — with cost jumping from $2.74 to $7.60 for 4 points.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Cheap Review Passes Change Development Economics
The core shift Meyer identifies is that routine, independent review has become affordable. A cross-family review pass at $0.32 to $0.39 per task can run on every meaningful change, whereas Astra or Fable in the same role cost roughly 8 to 20 times more. Meyer argues a different model family reviewing Opus output is a stronger check than Opus reviewing itself.
Meyer also cautions against over-reading the rankings: the Artificial Analysis index is a map of general capability, not a verdict on any specific workload, and he recommends shadow-testing before switching models. He notes that one index point is within measurement noise, and that halving model price saves only a fraction of real project cost — a single extra minute of human review can erase the saving, in an example he describes as illustrative rather than measured.
A Month of Front-End Releases in September 2026
The pieces of Meyer’s stack all shipped within four weeks: Claude Fable 5.1 and GPT-6 Astra on 1 and 3 September, GPT-6 Luna and Claude Opus 5.5 on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — the same day as his report. Sol launched at the same list price as its roughly week-old predecessor: $2 per million input tokens and $10 per million output tokens, compared with $4/$20 for Opus 5.5, $10/$50 for Fable and Astra, and $0.10/$0.50 for Luna.
Sol’s efficiency is tied to its brevity: at the high effort setting it used 25 million output tokens on the index benchmark against a median of 82 million for comparable models. Artificial Analysis lists three Sol effort settings so far and has not yet published low or max.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer, ThorstenMeyerAI.com
Benchmark Gaps and Latency Trade-Offs
Several limits remain. Artificial Analysis has not yet published Sol’s low or max effort settings, and Meyer notes that one index point falls inside measurement noise — so the 1-to-2-point gap between Sol and Astra or Fable may not be meaningful. Sol’s high and xhigh settings take 57 to 69 seconds to produce a first token, which Meyer says rules it out as an interactive model at those levels. Opus 5.5 still leads Sol by 5 points at xhigh.
Meyer’s workflow conclusions are also one practitioner’s interpretation of index data, not a measured result: he explicitly labels his cost-saving example as illustrative, and his “shadow-test before switching” advice reflects that index scores do not predict performance on any individual workload.
Pending Benchmarks and Shifting Model Roles
Watch for Artificial Analysis to publish Sol’s remaining effort settings, which could shift the value comparison further. Meyer says he will keep Sol in the review seat as long as its cross-family disagreement with Opus catches real defects, and will re-evaluate whether Opus at max is ever justified as pricing evolves. He also flags that GPT-6 Luna, at $0.07 per task, handles high-volume classification and routing via a decision model called Jev — a pattern that could expand as specialized, non-generative models take over high-volume yes/no judgements. Readers should treat all index figures as of 29 September 2026; this market has repriced monthly.
Key Questions
What is GPT-6.1 Sol, and when was it released?
GPT-6.1 Sol is an OpenAI model released on 29 September 2026, priced at $2 per million input tokens and $10 per million output tokens. It scores 51 on the Artificial Analysis Intelligence Index v4.3.x at its xhigh setting, costing about $0.39 per task.
Why does Thorsten Meyer prefer Opus 5.5 over higher-effort settings?
On Opus 5.5, moving from xhigh to max adds only 2 index points but 73% more cost per task, per Artificial Analysis data. Meyer runs high (54 points, $1.82) for everyday building and xhigh (56 points, $3.46) for architecture, migrations and trust boundaries.
What are GPT-6.1 Sol’s main drawbacks?
Its high and xhigh settings take 57 to 69 seconds to first token, making it unsuitable for interactive use, and it trails Opus 5.5 by 5 index points at xhigh. Its low and max settings have not yet been benchmarked by Artificial Analysis.
Are these benchmark scores a reliable guide for choosing a model?
Meyer cautions that the Artificial Analysis index measures general capability, not performance on your workload, and that single-point differences fall within noise. He recommends shadow-testing candidates against real tasks before switching.
Does a cheaper model automatically reduce project costs?
Not necessarily. Meyer notes that halving model price saves only about 12.5% of real cost in his illustrative example, and one extra minute of human review can erase the saving. He presents this as an illustration, not a measured result.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
