Skip to content
ProductAI ArchitectureEngineering

GPT-5.6 Sol: Frontier Intelligence at 67% Off and Half the Tokens

Gargeya SharmaFounder & Architect
August 26, 202618 min read
Essays
GPT-5.6 Sol: Frontier Intelligence at 67% Off and Half the Tokens
On this page

If you have been waiting for a clean reason to put a frontier model into production without flinching at the invoice, this is that reason. GPT-5.6 Sol is already one of the smartest models you can call. It is also, unusually, one of the cheapest ways to buy that intelligence — not because it is a budget model, but because it finishes the job with fewer tokens, in less time, and is now sitting on a stacked discount that most comparison charts have not fully priced in.

OpenAI dropped official Sol pricing from the July launch rate of $5.00 input / $30.00 output per million tokens to a promotional $4.00 / $20.00. That cut is live at least through 21 November 2026. On top of that, OpenRouter is applying an automatic 50% off on the OpenAI Standard route. The number you actually pay today through OpenRouter is $2.00 input / $10.00 output per million tokens — with cached reads at $0.20. Check Current Price of GPT 5.6 Sol at OpenRouter

That is 60% off launch input and 67% off launch output. For a model that already sits on the Pareto line for intelligence versus tokens, time, and cost, the discounted rate is not a rounding error. It is the whole story.


The stacked discount, in one table

The $2 / $10 figure is not a blended guess. It is the live OpenRouter list for openai/gpt-5.6-sol. Official OpenAI still posts $4 / $20. Azure and Bedrock are still closer to the old world at $5.00–$5.50 input and $30.00–$33.00 output. If you are calling Sol anywhere except the discounted OpenRouter OpenAI route, you are leaving money on the table.

Price basisInput / 1MOutput / 1MCache readvs launch invs launch out
Launch OpenAI (9 Jul 2026)$5.00$30.00$0.50
OpenAI promo (from 21 Aug)$4.00$20.00$0.40−20%−33%
OpenRouter → OpenAI (50% off)$2.00$10.00$0.20−60%−67%
OpenRouter batch / flex floor$1.00$5.00$0.10−80%−83%

USD per million tokens. OpenRouter 50% applies automatically on `openai/gpt-5.6-sol`. Flex/batch floors reported by OpenRouter in mid-August.

Live OpenRouter model page for GPT-5.6 Sol showing $2 / $10 with a 50% off badge. Azure and Bedrock still sit at $5.00–$5.50 / $30–$33.
Live OpenRouter model page for GPT-5.6 Sol showing $2 / $10 with a 50% off badge. Azure and Bedrock still sit at $5.00–$5.50 / $30–$33.

Figure 1. Live OpenRouter model page for GPT-5.6 Sol. Headline rate $2 / $10 with the 50% badge. Azure and Bedrock still sit at $5.00–$5.50 / $30–$33.


What the charts actually say

Artificial Analysis published a clean 26 August snapshot across intelligence, agentic work, output tokens, time, and cost. Two things jump off the page at the same time.

First: Sol is not a mid-pack model hiding behind a sale. On the Intelligence Index, GPT-5.6 Sol (max) scores 61. That is one point behind Claude Fable 5 with fallback (62) and two behind Claude Opus 5 max (63). Grok 4.6 high is tied with Sol at 61. Sol xhigh is 59. Sol high is 57. Even Sol medium is 56 — the same band as Gemini 3.7 Flash high.

Second: Sol gets there while spending far fewer tokens and far less wall-clock time than almost every model in that intelligence band. That is the part the list-price charts understate, and the part the OpenRouter discount multiplies.

Intelligence: frontier, not "good enough"

Artificial Analysis Intelligence Index bar chart, 26 August 2026. Claude Opus 5 max leads at 63, Fable 5 at 62, GPT-5.6 Sol max and Grok 4.6 high at 61.
Artificial Analysis Intelligence Index bar chart, 26 August 2026. Claude Opus 5 max leads at 63, Fable 5 at 62, GPT-5.6 Sol max and Grok 4.6 high at 61.

Figure 2. Artificial Analysis Intelligence Index v4.1.1 (26 Aug 2026). Nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.

Read that chart left to right and you will notice something builders care about more than a one-point gap at the top. The top of the board is crowded. Opus, Fable, Sol, Grok, Kimi K3, and GLM-5.3 are all in a six-point band. The decision is no longer "who is smartest on a poster." It is "who stays that smart after you pay for tokens, wait for the decode, and run the agent loop a hundred times."

Agentic work: Sol belongs in the first group

Artificial Analysis Agentic Index bar chart. Opus, GLM-5.3, and Grok sit at 59; GPT-5.6 Sol max is 58.
Artificial Analysis Agentic Index bar chart. Opus, GLM-5.3, and Grok sit at 59; GPT-5.6 Sol max is 58.

Figure 3. Artificial Analysis Agentic Index — weighted average of agentic benchmarks inside the Intelligence Index (GDPval-AA v2, τ³-Banking).

On the Agentic Index, Claude Opus 5 max, GLM-5.3 max, and Grok 4.6 high sit at 59. GPT-5.6 Sol max is 58. That is not a participation trophy. For long-horizon coding, command-line work, and multi-step tools — the jobs Sol was explicitly trained for — you are buying a top-shelf agent at a mid-shelf invoice.

Token efficiency: this is where Sol wins the argument

Bar chart of output tokens per Intelligence Index task. Sol variants occupy the entire left side, from 3k (low) to 17k (max). Claude and Qwen sit at 36k–47k.
Bar chart of output tokens per Intelligence Index task. Sol variants occupy the entire left side, from 3k (low) to 17k (max). Claude and Qwen sit at 36k–47k.

Figure 4. Weighted average output tokens per Intelligence Index task. Dark = answer tokens, light = reasoning tokens. Sol variants occupy the entire left side of the chart.

This is the chart I keep coming back to.

  • GPT-5.6 Sol (low): ~3,000 output tokens
  • Sol medium: 5,000
  • Sol high: 8,000
  • Sol xhigh: 11,000
  • Sol max: 17,000

Now look at the models that score in the same intelligence neighborhood. Grok 4.6 high uses 22,000. Kimi K3 max uses 25,000. Claude Fable 5 uses 36,000. Claude Opus 5 max uses 40,000. GLM-5.3 max uses 41,000. Qwen3.8 27B xhigh uses 47,000.

Sol max is producing frontier-grade answers with less than half the output of Fable or Opus. Sol high is doing 57-index work on 8,000 tokens — the kind of budget other labs spend before the model has even started reasoning. Token efficiency is not a footnote on Sol. It is the product.

Scatter of Intelligence Index versus output tokens per task on a log scale. Sol variants sit on the Pareto line inside the high-intelligence, low-token quadrant.
Scatter of Intelligence Index versus output tokens per task on a log scale. Sol variants sit on the Pareto line inside the high-intelligence, low-token quadrant.

Figure 5. Intelligence Index versus output tokens per task (log scale). The green quadrant is high intelligence, low token spend. Sol owns that quadrant. The dotted line is the Pareto frontier — and Sol is the line.

The Pareto line on this scatter is almost a Sol family portrait: low, medium, high, xhigh, max, climbing in intelligence without exploding token spend. Everyone else is to the right — same or lower intelligence, more tokens. That is why a price cut on Sol is more valuable than the same percentage cut on a chatty model. You were already buying fewer tokens. Now each of those tokens is cheaper.

Speed: minutes, not a coffee break

Bar chart of decode time per Intelligence Index task. Sol variants range from 0.6 to 3.8 minutes; Claude, Grok, Kimi, and Qwen run 5.2 to 18.3 minutes.
Bar chart of decode time per Intelligence Index task. Sol variants range from 0.6 to 3.8 minutes; Claude, Grok, Kimi, and Qwen run 5.2 to 18.3 minutes.

Figure 6. Weighted average decode time per task, excluding TTFT and overhead. Lower is better.

  • Sol low: 0.6 min
  • Sol medium: 1.1
  • Sol high: 1.7
  • Sol xhigh: 2.6
  • Sol max: 3.8

Compare that with Claude Fable 5 at 5.2, Grok 4.6 high at 6.5, Claude Opus 5 max at 7.2, Kimi K3 max at 9.7, and Qwen3.8 2.4T A95B at 18.3.

Scatter of Intelligence Index versus time per task. Sol variants stack on the left edge of the attractive quadrant.
Scatter of Intelligence Index versus time per task. Sol variants stack on the left edge of the attractive quadrant.

Figure 7. Intelligence versus time per task. Sol variants stack on the left edge of the attractive quadrant — high index, low minutes.

If you are looping an agent, those minutes compound. A 3.8-minute Sol max task versus a 7.2-minute Opus max task is not a benchmark curiosity. It is twice the iteration speed on a workday. Pair that with Sol high at 1.7 minutes and 57 intelligence, and you have a default setting that feels instant relative to the rest of the frontier.

Cost per task — and the asterisk the charts cannot show

Stacked bar chart of official list-price cost per Intelligence Index task. Luna max is $0.05; Sol max is $0.96; Claude Opus 5 max is $2.34; Claude Fable 5 is $3.14.
Stacked bar chart of official list-price cost per Intelligence Index task. Luna max is $0.05; Sol max is $0.96; Claude Opus 5 max is $2.34; Claude Fable 5 is $3.14.

Figure 8. Official list-price cost per Intelligence Index task, segmented by answer, reasoning, cache, and input. These bars do not include the OpenRouter 50% Sol discount.

Scatter of Intelligence Index versus official cost per task. Sol already sits on the Pareto line at list price.
Scatter of Intelligence Index versus official cost per task. Sol already sits on the Pareto line at list price.

Figure 9. Intelligence versus official cost per task. Sol already sits on the Pareto line at list price. The live OpenRouter rate moves every Sol point left.

This is the caveat that changes the reading of every cost chart in this pack. Artificial Analysis prices tasks off official provider rates. For Sol, that is the OpenAI promotional list of $4 / $20 — already cheaper than launch, but not the $2 / $10 you pay on OpenRouter today. The bars above are the pre-discount world.

Halve the Sol numbers and the picture gets almost unfair.

  • Sol max: $0.96 → ~$0.48
  • Sol high: $0.43 → ~$0.22
  • Sol medium: $0.30 → ~$0.15
  • Sol low: $0.19 → ~$0.10

Claude Fable 5 stays near $3.14. Claude Opus 5 max stays near $2.34. You are looking at a six-to-one or worse gap for a one- or two-point difference on the index.

Model (reasoning)AA IndexList cost / taskOpenRouter-adj.Index per $1
Claude Opus 5 (max)63$2.34$2.3427
Claude Fable 5 (fallback)62$3.14$3.1420
GPT-5.6 Sol (max)61$0.96~$0.48~127
Grok 4.6 (high)61$0.84$0.8473
Kimi K3 (max)60$0.84$0.8471
GPT-5.6 Sol (high)57$0.43~$0.22~259
GPT-5.6 Terra (max)57$0.51$0.51112
GPT-5.6 Luna (max)52$0.05$0.051,040

Approximate OpenRouter-adjusted costs assume a clean 50% cut on Sol only. Terra and Luna currently price at official OpenAI rates on OpenRouter. Index-per-dollar is Index ÷ adjusted task cost — a rough value density, not a benchmark score.

Two readings fall out of that table. If you want the top of the board, discounted Sol max is the only way to buy a 61-index model for under fifty cents a task. If you want the everyday default, discounted Sol high is a 57-index model at about twenty-two cents — cheaper than Terra max at list, faster than Terra max, and more token-efficient than Terra max.


Sol, Terra, and Luna — pick on purpose

The 5.6 family is a ladder, not a single SKU. Live OpenRouter rates as of 26 August 2026:

SolTerraLuna
RoleFlagship reasoning, coding, agentsBalanced everyday driverHigh-volume, low-latency
OpenRouter in / out$2.00 / $10.00 (50% off)$2.00 / $12.00$0.20 / $1.20
OpenAI official in / out$4.00 / $20.00 promo$2.00 / $12.00$0.20 / $1.20
Cache read / write$0.20 / $2.50$0.20 / $2.50$0.02 / $0.25
AA Index (max)615752
AA Agentic (max)585047
Output tokens / AA task (max)17k21k20k
Time / AA task (max)3.8 min3.0 min2.5 min
List cost / AA task (max)$0.96 → ~$0.48 on OR$0.51$0.05

OpenRouter API snapshot, 26 Aug 2026. Terra and Luna no longer carry the extra mid-summer 50% overlay; Sol still does. Context window is 1M across the family.

A year ago the instinct was "use the cheap model unless the task is hard." The discounted Sol rate breaks that habit. Sol high at ~$0.22 per AA-style task is cheaper than Terra max at $0.51, scores the same 57 on the index, uses fewer tokens (8k vs 21k), and returns faster (1.7 min vs 3.0). Terra still has a job — especially if you want a stable mid-tier without hunting effort knobs — but the default for serious work should be Sol until the 50% goes away.

Luna is the opposite lesson. At $0.20 / $1.20 it is a different product: 52 intelligence, $0.05 a task, 2.5 minutes. Classification, routing, cheap first-pass agents, and anything you will run ten thousand times belongs on Luna. Do not spend Sol tokens on work Luna will finish before Sol has warmed up.


How Sol compares with the rest of the field

Live OpenRouter sticker prices next to the 26 August Artificial Analysis snapshot. This is the competitive frame I want in your head when someone says "just use the other lab's flagship."

ModelOR inOR outIndexAgenticTok/taskMin/task
GPT-5.6 Sol (max)$2.00$10.00615817k3.8
Grok 4.6 (high)$2.00$6.00615922k6.5
Kimi K3 (max)$3.00$15.00605425k9.7
GLM-5.3 (max)$1.40$4.40605941k6.9
Claude Opus 5 (max)$5.00$25.00635940k7.2
Claude Fable 5$10.00$50.00625736k5.2
Gemini 3.7 Flash (high)$0.38$1.88564537k1.5
DeepSeek V4 Pro 0813lowlow535039k6.6

OpenRouter in/out is USD per million tokens, 26 Aug 2026. Index / agentic / tokens / minutes are Artificial Analysis 26 Aug 2026 at the reasoning level in the label. Gemini 3.7 Flash rates rounded from $0.375 / $1.875.

What this comparison is really saying

Grok 4.6 high is the honest rival. Same 61 intelligence, cheaper output tokens at $6 / 1M, slightly higher agentic score. It also spends 22k tokens and 6.5 minutes where Sol max spends 17k and 3.8. On a long agent loop, Sol's time and token discipline often erase Grok's output-price advantage. If your workload is short answers, Grok is in the conversation. If your workload is tools, terminals, and multi-step coding, Sol is the cleaner buy at $2 / $10.

Claude is the quality ceiling people still reach for on taste, design, and certain writing jobs. Fair. Opus is 63. Fable is 62. They also cost 2.5× to 5× more on input and 2.5× to 5× more on output, and they emit two times the tokens. I like Claude. I do not like paying Claude prices for work Sol will finish in half the tokens and half the time. Use Claude when you can taste the difference. Use Sol when you can measure it.

GLM-5.3 and Kimi K3 are the pressure from the other side of the map. Both score 60. GLM is cheaper per token than discounted Sol. It also burns 41k output tokens and 6.9 minutes. Cheap rates on a verbose model are how invoices quietly grow. Measure cost per finished task, not cost per million.

Gemini 3.7 Flash is fast and inexpensive, and it is not in Sol's intelligence or agentic band. That is a different shelf. Flash belongs next to Luna, not next to Sol max.


Why the design and writing point matters

Benchmarks will not tell you this as cleanly as an afternoon in a harness will. Sol has a design sense. Slide layouts, docs, UI copy, and structured artifacts come out looking like someone cared. Artificial Analysis has separately flagged Sol's presentation quality on AA-Briefcase — highest Presentation Elo in that suite at launch. Combined with the coding and command-line strength OpenAI trained into the 5.6 flagship, you get a model that can plan the work, do the work, and package the work.

That is the product you want under a discount. Not a toy you burn through on lorem ipsum. A collaborator you can leave on for a week of building.


How to use the discount without thinking

You do not need a new stack.

  1. 1Create an OpenRouter key.
  2. 2Point your existing harness — Codex, Claude Code-compatible wrappers, Aider, Cursor custom endpoint, OpenClaw, your own agent loop — at https://openrouter.ai/api/v1.
  3. 3Set the model to openai/gpt-5.6-sol. The 50% off applies automatically. No coupon code.
  4. 4Start on reasoning effort high for daily coding and writing. Move to xhigh or max when the task is long-horizon. Use Luna for the firehose.
  5. 5Confirm you are landing on the OpenAI Standard route, not Azure or Bedrock. Those providers are still at the expensive sticker.

OpenRouter also discounted batch and flex. If the work can wait a bit, flex has been quoted as low as $1.25 input / $7.50 output. Batch on OpenRouter is $1 / $5. That is launch Sol at a fifth of the original output price. Use it for eval sweeps, overnight refactors, and anything that does not need to answer in the same breath as the user.


The honest caveats

Discounts end. OpenAI's official Sol promo is dated through at least 21 November 2026. OpenRouter's 50% is a route promotion, not a law of physics. Build as if the window is open now, not as if $2 / $10 is the permanent list.

Long context is priced differently. Prompts over roughly 272k input tokens on official OpenAI bill at 2× input and 1.5× output for the full request. Cache writes on GPT-5.6 bill at 1.25× uncached input. If your agents rewrite the same system prompt every turn instead of hitting cache, you will not see the invoice you expected. Put stable instructions first. Reuse prefixes.

Artificial Analysis cost charts will lag promotions. If you quote those charts in a budget meeting, say the sentence out loud: "This is list. Live Sol on OpenRouter is about half of that bar."


Go build

I do not think anyone should blink twice before giving this a go. You are not compromising on intelligence to save money. You are taking one of the two or three smartest models available, the most token-efficient model in that band, one of the fastest, and you are buying it at 60 to 67 percent off the price it launched at ten weeks ago.

Use the discount. Build more than you planned to build this month. Try the idea you were going to postpone until the model got cheaper — it just did. Connect the OpenRouter key to the harness you already like. Leave Sol on high. Turn it up when the problem deserves it.

Happy coding. Happy OpenAI. Happy Sol.

Next: I will be writing a follow-up on OpenAI's latest chip launch, Jalapeno. Stay tuned for that.


Sources and notes

Pricing: OpenAI API pricing page and business pricing (GPT-5.6 Sol promotional $4 / $20 at least through 21 Nov 2026; Terra $2 / $12; Luna $0.20 / $1.20). OpenRouter API and model pages, 26 Aug 2026 (openai/gpt-5.6-sol $2 / $10 with 50% off; terra $2 / $12; luna $0.20 / $1.20). OpenRouter public notes on Sol 50% off (17 Aug 2026) and earlier Terra/Luna overlay (30 Jul 2026). Launch Sol list was $5 / $30.

Charts: Artificial Analysis, 26 August 2026 snapshots — Intelligence Index v4.1.1, Agentic Index, output tokens per task, time per task, cost per task, and the three Pareto scatters. Reasoning models marked with a lightbulb on the originals. Cost charts use official list rates, not the live OpenRouter Sol promotion.

Adjusted task costs for Sol apply a straight 50% to Artificial Analysis list-cost bars and should be treated as estimates. Mix of input, output, reasoning, and cache will move the exact figure. Always run a billing canary on your own traffic before you rewrite a forecast.

Continue