🔍 Read the full analysis: How Opus, Sol, And Jev Divide The Work In My AI Stack on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A September 29 article by Thorsten Meyer describes using Opus 5.5 for building, GPT-6.1 Sol for detailed review and Jev for high-volume routing decisions. The proposed split reflects the article’s reported capability index scores and per-task costs; those benchmarks may not predict results on other workloads.
Meyer bases comparisons among six models on the Artificial Analysis Intelligence Index v4.3.x, which he describes as a measure of general capability rather than a verdict on an individual workload. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task; GPT-6.1 Sol scores 51 at xhigh and costs $0.39. He says Luna is cheaper still, at $0.07 per task, with an index score of 37.
The recommended split varies by task and effort setting. Meyer uses Opus at high effort for ordinary development, citing a score of 54 and a $1.82 task cost, and at xhigh for more difficult work such as architecture and migrations. He assigns Sol high or xhigh settings to detailed file or code-diff reviews, and says Jev handles routing and other high-volume binary judgments. The source does not provide Jev benchmark scores or per-decision costs.
Meyer says GPT-6.1 Sol was released on September 29 at the same token prices as its predecessor, GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. He reports a $0.32 task cost at high and $0.39 at xhigh, compared with $3.26 for Astra and $7.63 for Fable 5.1 in his table. These are figures he attributes to Artificial Analysis; the article advises readers to test models against their own work before switching.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Cost Shapes the Model Assignment
The account illustrates a change in how Meyer chooses models: he weighs the reported cost per completed task against capability scores, instead of treating a leaderboard position as the sole criterion. If a lower-cost model clears a team’s quality bar for routine review or routing, that could make it practical to run those checks more often. Meyer says Sol’s reported price makes it affordable for a review pass on every meaningful change.
That approach also makes the review step part of the division of labor. Meyer argues that a model from a different family can provide a useful second perspective on Opus’s output. He cautions, however, that a second model cannot fix an incomplete specification on its own, and that passing tests do not by themselves authorize a release. These are his operating rules, rather than independently evaluated findings.
The reported costs also depend on the workload and effort setting. Meyer’s figures show Opus at max costing more than three times its high setting for a higher index score, while he describes Sol’s high and xhigh settings as slow to begin generating output. Teams considering the same arrangement would need to weigh quality, latency and human review time on their own tasks.
Benchmarks Behind Meyer’s Stack
The article compares six models: Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. Meyer reports that their top-setting index scores fall between 37 and 58, while listed costs range from $0.07 to $7.63 per task. He presents this spread as a reason to consider the cost of the task alongside the score.
Effort settings account for part of the difference. In the reported table, Opus scores 42 at low effort and 58 at max; its per-task cost rises from $0.55 to $5.98. Sonnet reaches 56 at max for $7.60 per task, while its high setting scores 47 for $1.08. Meyer says he favors Sonnet at high for scoped subtasks and documents, rather than using its most expensive setting by default.
Meyer also reports that Sol’s high and xhigh index runs took 57 and 69 seconds, respectively, to produce a first token. He describes those settings as unsuitable for interactive use. His article cautions that a one-point index difference may fall within measurement noise and recommends shadow-testing a model before adopting it.
“The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.”
— Thorsten Meyer, in the September 29 article
Workload Results Still Need Testing
The article reports benchmark and cost figures but does not provide the underlying task set, full measurement method or independent replication. It is not clear how closely the listed per-task costs would match a particular team’s prompts, output lengths or usage patterns. Meyer says one index point is within the noise and advises shadow-testing before a switch.
The source gives no comparable benchmark scores or per-decision prices for Jev, the decision model in Meyer’s stack. It also says Artificial Analysis had not published GPT-6.1 Sol’s low- or max-effort results at the time of publication. The article’s claims about review quality and the value of using different model families are Meyer’s assessment; no controlled comparison is included.
Shadow Tests Before Adoption
Meyer recommends running a candidate model alongside the existing setup on real tasks before replacing a default. That testing can show whether a model meets the required quality bar, how long it takes to respond and what the task costs in practice. He says failures in review should be returned to Opus with the failing case and evidence, rather than a general instruction to try harder.
The next comparison may change as more effort settings and measurements become available. Meyer’s account does not specify a date for another update, a formal rollout schedule or results from a completed shadow test. For now, the stack describes his reported practice on September 29, 2026.
Key Questions
How does Thorsten Meyer divide work across the models?
He uses Opus 5.5 for building, GPT-6.1 Sol for detailed review and Jev for high-volume yes-or-no and routing decisions.
Why does Meyer use GPT-6.1 Sol for review?
He says Sol’s reported cost of $0.32 to $0.39 per task makes it affordable for routine review. Those costs come from the benchmark figures he cites and may differ on other workloads.
Does the article say Sol is better than Opus 5.5?
No. Meyer reports a lower top-setting index score for Sol: 51 at xhigh, compared with 56 for Opus at xhigh and 58 for Opus at max. He assigns them different jobs based on reported scores and costs.
What remains unknown about Jev?
The source describes Jev as a decision model for routing and binary judgments, but gives no benchmark score, per-decision cost or evaluation results.
Should teams adopt Meyer’s model assignments?
The article recommends shadow-testing before switching. Benchmark rankings and listed costs do not establish which model will perform best on a team’s particular work.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
