The title asks "who is strongest." The September 2026 answer is: no single leaderboard covers reasoning, coding agents, high-volume JSON, long context, and self-hosting at once. This piece is dated 2026-09-04 and only covers flagships that are announced or already GA: GPT-5.6, GPT-6 Astra just starting to roll out, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.6, Llama 4. Below: first the facts, then a split by task, then fold "which model to pick" into validatable JSON.
"Strongest model" assumes one name that sweeps every task. 2026 product lines have already split: some sell intelligence, some sell coding agents, some sell price per million tokens, some sell whether you can download the weights. Collapse five vendors into one champion, and the docs are left with a model field that expires next week.
What you can confirm as of 2026-09-04
The last seven days were dense. Only record facts you can put on a schedule:
- 2026-09-01, Anthropic released Claude Fable 5.1 (API:
claude-fable-5-1). Official guidance still starts most work on Claude Opus 5, and only escalates hard problems to Fable. - 2026-09-02, Google released Gemini 3.8 Flash. Gemini 4 has no release date; 2026-07-21 only confirmed that pre-training has started. See the previous piece When will Gemini 4 be released.
- 2026-09-03, OpenAI announced GPT-6 Astra, first for Daybreak, then expanding to paid ChatGPT and the API. Until Astra is stably callable in your account, the production default is still GPT-5.6, GA on 2026-07-09.
- 2026-08-12, xAI released Grok 4.6 (API:
grok-4.6), 500K context. - Llama 4 Scout / Maverick are still Meta's publicly downloadable weight line (2025-04-05). Don't put open weights and closed-source APIs in the same "strongest" column.
Lock the boundary. Independent leaderboards (Artificial Analysis Intelligence Index) can be a reference, not a contract. OpenAI's own AGI wording, unpublished price lists, and model IDs that are not in your account yet do not go into production docs. DevDay 2026 is still 2026-09-29; after the event, the official record wins.
Five flagships at a glance
Prices are vendor list prices, per million tokens. Independent scores are this week's Artificial Analysis readings; effort tiers differ, so you cannot compare them to a decimal point.
| Family | ID you can write into docs now | Context | List price (in / out) | Better for |
|---|---|---|---|---|
| GPT | gpt-5.6 (โ Sol); Astra rolling out |
~1M | Sol about $5 / $28โ30 | Coding agents, Responses / PTC |
| Claude | claude-opus-5 / claude-fable-5-1 |
1M / 128K out | Opus $5 / $25; Fable $10 / $50 | Hard reasoning, long-running agents, knowledge work |
| Gemini | gemini-3.8-flash |
1M / 64K out | Promo $0.75 / $3.75 (through 2026-12-31) | High-volume JSON, multimodal, cheap workhorse |
| Grok | grok-4.6 |
500K | $2 / $6 (prompts >200K double) | Long-running agents, visual interaction, mid price band |
| Llama | Scout / Maverick weights | Scout 10M; Maverick 1M | Self-host: you pay power, not an API price | Data stays in-perimeter, fine-tunable, marginal cost |
The same week's independent readings can be a direction, not a schedule: Fable 5.1 (max) Intelligence Index 66; GPT-5.6 Sol and GPT-6 Astra both at 61; Grok 4.6 (high) 61; Gemini 3.8 Flash (high) 59. On the coding-agent index, Astra reaches 67; Fable 5.1 reaches 70 inside Claude Code. A few points are not enough to change architecture; they are enough to pick a default route.
Where each of the five product lines is strong
1. GPT: production default is still 5.6; Astra is rolling out
2026-07-09, GPT-5.6 (Sol / Terra / Luna) GA. The flagship alias gpt-5.6 points to gpt-5.6-sol. The previous piece DevDay 2026 Agent already covered this: the Responses API, Programmatic Tool Calling, and MCP are already on this line.
GPT-6 Astra on 2026-09-03 is a next-generation name, not another product line. On announcement day it went to Daybreak first; paid ChatGPT and the API roll out "in the next few days." In independent evals, Astra's Intelligence Index ties Sol (61); coding agents are stronger, and the price is higher. Treat Greg Brockman's AGI wording as news, not a schedule.
The safe way to write it into docs: default gpt-5.6; after Astra is stably callable on your side, only change model_id. The tool contract stays three layers โ parameters / output_schema / Structured Output. Don't rewrite fields for a new name.
2. Claude: top of the independent intelligence index
Fable 5.1 is the 2026-09-01 flagship. Anthropic itself locked the split: most work uses Opus 5; Fable is for reasoning and long-running agents where a higher-effort Opus is still not enough. Price follows that split: Opus $5 / $25, Fable $10 / $50, cache reads down to $0.25 / MTok.
Mythos 5.1 and Fable are the same weights, only a different safety tier, on an invite path. Don't write it into a routing table open to every user.
On the independent Intelligence Index, Fable 5.1 (max) is 66, currently the highest. That is the closest answer to the "strongest" headline, and it is limited to this intelligence axis. It is not cheap, and it should not be the default for high-volume JSON.
3. Gemini: Flash is pushing; 4 is not here yet
3.8 Flash is the 2026-09-02 workhorse, promo-priced at $0.75 / $3.75 through 2026-12-31. Context is 1M, output cap 64K โ narrower than Claude / GPT's 128K out. Independent score 59, but price per unit of intelligence sits near the Pareto frontier.
Gemini 4 still only has "pre-training has started." 3.x is dual-track: Flash ticks every few weeks; Pro has not followed to 3.8. Write the production path as 3.8 Flash, reserve public_model_id for a later 4, and don't lock "wait for 4 before we ship" into the docs.
4. Grok: long-running agents and the mid price band
Grok 4.6 shipped 2026-08-12, grok-4.6, 500K context, reasoning tiers low / medium / high / xhigh. List price $2 / $6; prompts over 200K tokens double. Official emphasis is long-running agents, coding, and interactive vision โ not another "first across the board" leaderboard.
Independent intelligence readings (high) sit in the same band as GPT-5.6 Sol. It is on the comparison table because the price sits between Gemini Flash and the Claude / GPT flagships โ a fit for routes that need long context and do not want to pay Fable prices.
5. Llama: open weights, not the same leaderboard
Llama 4 Scout (17B active / 109B total, 10M context) and Maverick (17B active / 400B total, 1M context) are the 2025-04-05 public weights. This piece does not write an unreleased next-generation open-source name as fact.
Comparing "who is strongest" against closed-source flagships compares the wrong thing. Llama wins on: downloadable weights, fine-tuning, data that can stay in-perimeter, and marginal cost that is power, not an API bill. It loses on: no first-class vendor-hosted PTC / MCP tools; Structured Output has to be held by your own constrained decoding or validation layer.
The self-host path's contract is still JSON Schema. If the model changes, the sample and required do not.
Pick by task, not by "strongest"
Split the headline into five executable rules. Write the rules into routing JSON, not into slogans.
- Hard reasoning, long-running knowledge work, evals still not enough on Opus: Claude Fable 5.1. Default to Opus 5 first.
- Coding agents, Responses API, Programmatic Tool Calling: GPT-5.6; switch the ID after Astra is stable. Tool contracts: see the DevDay Agent piece.
- High-volume Structured Output, multimodal, need to crush unit cost: Gemini 3.8 Flash. Output cap 64K; chunk long documents first.
- Long-running agents, mid price, 500K context is enough: Grok 4.6.
- Data stays in-perimeter, you need fine-tuning, volume makes the API uneconomic: Llama 4 Scout or Maverick. Output must pass your own Schema.
The same product can hang two defaults at once: the cheap path on Gemini / Grok, then escalate to Claude / GPT on failure or low confidence. Write escalation conditions as fields, not as "it doesn't feel smart enough."
A routing sample that is enough to generate a Schema
Below is JSON for the act of picking a model itself. Enums as strings, scores as number, callable-or-not as boolean. Don't let the model say out loud "just use Claude, more or less."
{
"task": "structured_output",
"priority": "cost",
"constraints": {
"data_leaves_vpc": false,
"max_output_tokens": 8000,
"needs_tool_calling": true
},
"candidates": [
{
"family": "gemini",
"model_id": "gemini-3.8-flash",
"role": "workhorse",
"available": true,
"input_usd_per_mtok": 0.75,
"output_usd_per_mtok": 3.75
},
{
"family": "claude",
"model_id": "claude-opus-5",
"role": "escalation",
"available": true,
"input_usd_per_mtok": 5,
"output_usd_per_mtok": 25
}
],
"selected": {
"family": "gemini",
"model_id": "gemini-3.8-flash",
"reason": "high_volume_structured_output"
}
}
With this site's inference rules, that sample yields:
| Field | Infer types | What to do after generate |
|---|---|---|
task / priority / reason |
string | Collapse to an enum; don't leave free text |
data_leaves_vpc / available / needs_tool_calling |
boolean | Samples must be true / false; writing "yes" drifts to string |
max_output_tokens / price fields |
number | Writing "0.75" becomes a string |
candidates |
array of object | Every item needs family and model_id |
selected.model_id |
string | This is the swappable production field; a generation swap only changes this |
Once the sample is ready, open Structured Output to generate each platform's wrapper, then review required and additionalProperties by hand. Routing results can also go through Function Calling: make "pick a model" itself a tool whose return must match this Schema.
How to accept it in JSON Toolbox
- Paste the routing sample above into JSON Format and confirm it parses.
- Generate a first draft with Structured Output or JSON โ Schema.
- Take a second set of real routing results to Schema validation. The first sample is always optimistic.
- When you switch from Gemini to Claude, or from 5.6 to Astra, use JSON Diff to see whether a field dropped or a type drifted.
- Convert to Zod when the frontend needs checks, and to OpenAPI 3.1 when you need API docs.
Everything stays in the browser and is never uploaded.
FAQ
Which model is strongest in 2026?
There is no single champion. The independent intelligence index currently peaks at Claude Fable 5.1. On coding agents, GPT-6 Astra and Claude are close. High-volume Structured Output more often picks Gemini 3.8 Flash. For self-hosting, the answer is Llama 4.
Can GPT-6 Astra already be the production default?
It was announced 2026-09-03, first via Daybreak. This piece is dated 2026-09-04; the production default is still GPT-5.6. Change model_id after Astra is in your account.
Does Gemini 4 count as already released?
No. There is no announced date. The production path is still 3.x; the latest workhorse is Gemini 3.8 Flash.
Can Llama 4 still compete with closed-source flagships?
They are not the same leaderboard. Llama compares self-hosting and the data perimeter; closed-source flagships compare API scores and hosted tools.
Do you need to rewrite Structured Output when you swap models?
No. Lock the field contract; swap model_id and each platform's wrapper. Generate the Schema from the same sample, then accept it with a second set of real output.
Summary
September 2026's "strongest" is five tracks, not one name. Claude carries the hardest reasoning, GPT carries coding agents and tool loops, Gemini carries high-volume cheap JSON, Grok carries mid-price long-running agents, Llama carries weights and the data-perimeter boundary. What survives next week's generation swap is the routing rules and the Schema, not a champion poster. Grow the contract from one sample; validation, Zod, and OpenAPI all follow it.