Skip to content

Split View: GPT-5.6과 가격 대비 성능의 한계 — 곡선 위에서 내 워크로드의 자리를 찾는 법

✨ Learn with Quiz
|

GPT-5.6과 가격 대비 성능의 한계 — 곡선 위에서 내 워크로드의 자리를 찾는 법

들어가며 — 출시 3주 만의 80% 인하가 말해 주는 것

2026년 7월 30일, OpenAI가 GPT-5.6 Luna의 가격을 80% 내렸습니다. 100만 토큰당 입력 1달러 출력 6달러에서 입력 0.20달러 출력 1.20달러가 됐습니다. 중간 티어인 Terra는 20% 내려 입력 2달러 출력 12달러가 됐고, 최상위 Sol은 입력 5달러 출력 30달러 그대로입니다. GPT-5.6 계열이 정식 출시된 것이 7월 9일이니 3주 만의 조정입니다.

같은 날 발표된 두 번째 항목이 성격을 더 분명하게 보여 줍니다. 기존 Priority processing이 Fast 모드로 이름을 바꾸면서 Sol에 한해 표준 대비 최대 2.5배 속도를 표준가의 2배에 제공합니다. OpenAI의 API 문서는 이를 "지능의 변화 없이 속도만"이라고 명시합니다. 즉 아래쪽 티어는 값을 내려 물량을 지키고, 위쪽 티어는 지연 시간을 별도 상품으로 떼어 팔기 시작한 것입니다.

배경으로 자주 인용되는 숫자가 하나 있습니다. CNBC 보도를 인용한 여러 매체가 OpenRouter 기준 미국 엔터프라이즈 토큰 사용량의 46%를 중국 모델이 가져갔다고 전했습니다. 저는 이 수치의 원 집계 방법을 확인하지 못했으므로 보도된 주장으로만 취급합니다. 다만 인하폭의 비대칭(아래는 80%, 위는 0%)이 어디를 겨냥한 것인지는 굳이 설명이 필요 없어 보입니다.

이 글의 주제는 인하 소식이 아닙니다. 능력당 단가는 몇 년째 곡선을 그리며 떨어지고 있고, 개별 인하는 그 곡선 위의 점 하나입니다. 쓸모 있는 질문은 하나뿐입니다 — 내 워크로드는 그 곡선 위 어디에 있고, 이번 이동이 그 자리를 실제로 옮겼는가.

7월 9일과 7월 30일 사이에 바뀐 것

Artificial Analysis가 출시 시점에 집계한 지표와 이번 인하를 한 표에 놓으면 이렇습니다. 지수 과제 1건당 비용의 인하 후 값은 제가 인하율을 그대로 적용해 계산한 값이고, 토큰 소비량이 동일하다고 가정한 것입니다.

티어출시가 (입력/출력, 100만 토큰)7월 30일 이후Intelligence IndexCoding Agent Index지수 과제 1건 비용 (출시가 → 인하 후)
Sol5달러 / 30달러변동 없음59801.04달러 → 1.04달러
Terra2.50달러 / 15달러2달러 / 12달러55770.55달러 → 약 0.44달러
Luna1달러 / 6달러0.20달러 / 1.20달러51750.21달러 → 약 0.042달러

같은 지수에서 Claude Fable 5의 최대 강도 점수는 60으로 집계됩니다. 즉 Sol은 종합 지능에서 1점 뒤이면서 지수 완주 비용은 3분의 1 수준이라는 것이 출시 당시 Artificial Analysis의 요약이었습니다.

표를 세로로 읽으면 더 흥미롭습니다. 지능 지수는 59, 55, 51로 4점씩 내려가는데 지수 과제 1건당 비용은 1.04달러, 0.44달러, 0.042달러로 한 자릿수씩 떨어집니다. 지능 지수 4점을 사는 데 드는 돈이 구간마다 열 배씩 다릅니다. 곡선이라는 표현을 쓰는 이유가 이것입니다. 위쪽 끝에서는 아주 조금의 능력에 아주 많은 돈이 들고, 아래쪽에서는 같은 돈으로 훨씬 많은 능력을 삽니다.

곡선의 기울기 — 그리고 벤더의 효율 주장

능력당 단가가 떨어지는 속도에 대한 인용은 넘쳐 나지만, 대부분 출처가 흐릿합니다. 그중 비교적 자주 인용되는 정리는 특정 벤치마크 수준을 고정했을 때 그 성능의 가격이 연간 한 자릿수 배에서 열 배 수준으로 떨어져 왔다는 것입니다. 이번 인하는 그 추세선 위의 한 점이고, 3주라는 간격이 이례적일 뿐 방향 자체는 새롭지 않습니다.

여기서 벤더의 주장 하나를 검증해 볼 만합니다. GPT-5.6 출시 당시 OpenAI 측은 특정 과제에서 토큰 효율이 54% 개선됐다는 취지의 설명을 했고, 여러 매체가 이를 전했습니다. 이건 "특정 과제"라는 단서가 붙은 벤더 주장입니다. 반면 독립 집계 쪽 숫자는 훨씬 온건합니다 — Artificial Analysis 기준으로 Sol이 지수 과제 1건을 처리하는 데 쓴 출력 토큰은 15,000개로, 이전 세대 GPT-5.5의 16,000개 대비 약 6% 적습니다.

두 숫자가 모순은 아닙니다. 54%는 고른 과제에서, 6%는 종합 지수 전체에서 나온 값이니까요. 하지만 실무 계획에 넣을 수 있는 쪽은 후자입니다. 벤더가 "특정 과제에서"라고 단서를 달면, 그 단서가 곧 일반화 금지 표시입니다. 자기 과제에서의 토큰 소비는 직접 재는 수밖에 없고, 그 측정이야말로 모델 교체 전에 반드시 해야 하는 일입니다.

헤드라인 가격이 아니라 믹스로 계산하라 — 그런데 이번엔 믹스가 안 바뀐다

보통 두 모델을 비교할 때 첫 번째 함정은 입력 가격만 보는 것입니다. 벤더마다 출력 대 입력 배수가 다르기 때문에, 같은 두 모델도 워크로드의 토큰 믹스에 따라 비용 순위가 뒤집힙니다.

그런데 이번 인하는 그 함정을 GPT-5.6 계열 안에서 없애 버렸습니다. 인하 후 세 티어의 단가를 나란히 놓으면 입력은 5, 2, 0.20이고 출력은 30, 12, 1.20입니다. 두 축 모두 정확히 25 대 10 대 1입니다.

Sol   : 입력 5    출력 30      -> Luna 대비 25배
Terra : 입력 2    출력 12      -> Luna 대비 10배
Luna  : 입력 0.20 출력 1.20    -> 기준

어떤 (입력 토큰, 출력 토큰) 조합을 넣어도 총액 비율은 25 : 10 : 1 로 고정된다.
믹스는 티어 선택에 아무 정보를 주지 않는다.

계열 안에서 믹스가 순위를 못 바꾼다면, 티어 선택은 다른 축에서 정해져야 합니다. 그 축은 오답 비용입니다. 실제 워크로드로 계산해 봅니다.

워크로드 — 상품 설명 100만 건 분류
  요청당 입력 4,000 토큰 / 출력 300 토큰

  Luna  : (4,000 x 0.20 + 300 x 1.20) / 1e6 = 0.00116 달러/건 -> 1,160 달러
  Terra : 위의 10배                                            -> 11,600 달러
  Sol   : 위의 25배                                            -> 29,000 달러

  여기에 오답을 사람이 고치는 비용 c(건당)를 더한다.
  Luna 오답률 4%, Terra 오답률 2%로 가정하면

  Luna  총비용 =  1,160 + 0.04 x 1e6 x c
  Terra 총비용 = 11,600 + 0.02 x 1e6 x c

  두 식이 같아지는 지점:
  10,440 = 0.02 x 1e6 x c   ->   c = 0.522 달러

읽는 법은 이렇습니다. 오답 하나를 고치는 데 0.52달러보다 적게 든다면 Luna가 유리하고, 그보다 많이 든다면 Terra가 유리합니다. 사람 손이 한 건에 1분 붙는 작업이라면 인건비는 이미 0.52달러를 훌쩍 넘습니다. 반대로 오답이 그냥 재시도로 해결되는 파이프라인이라면 재시도 비용은 0.00116달러라 Luna가 압도적으로 유리합니다.

이 계산이 주는 교훈은 단순합니다. 토큰 단가 10배 차이가 총비용 차이로 나타나는 구간은, 오답 비용이 토큰 비용과 같은 자릿수일 때뿐입니다. 대부분의 사내 워크플로에서 오답 비용은 토큰 비용보다 두세 자릿수 큽니다. 그런 워크로드에서 가격표를 비교하며 몇 시간을 쓰는 것은 잘못된 축에서 최적화하는 일입니다.

캐싱이 곡선을 옮기는 구간, 그리고 못 옮기는 구간

GPT-5.6 계열의 캐시 정책은 읽기 90% 할인, 쓰기 입력가의 1.25배, 최소 유지 30분입니다. Terra에 40,000 토큰짜리 고정 프리픽스를 쓰는 경우로 계산해 봅니다.

프리픽스 40,000 토큰, Terra 입력 2달러/1M 기준

  캐시 없음      : 40,000 x 2.00 / 1e6 = 0.080 달러 (매 호출)
  캐시 쓰기(1회) : 40,000 x 2.50 / 1e6 = 0.100 달러
  캐시 읽기      : 40,000 x 0.20 / 1e6 = 0.008 달러

  N회 호출 손익분기:
  0.100 + (N-1) x 0.008  <=  N x 0.080
  0.092 <= 0.072 N   ->   N >= 1.28   ->  두 번째 호출부터 이득

손익분기는 두 번째 호출입니다. 진짜 제약은 손익분기가 아니라 30분입니다. 같은 프리픽스가 30분 안에 다시 오지 않으면 쓰기 프리미엄만 내고 아무것도 못 건집니다. 즉 캐싱의 실효 여부를 결정하는 것은 할인율이 아니라 프리픽스별 요청 밀도입니다. 사용자 1만 명이 각자 다른 프리픽스로 하루 두 번씩 부르는 서비스에서 캐싱은 비용을 늘립니다.

그리고 할인율이 곧 절감률이 아니라는 점을 다시 짚습니다. 위 워크로드에서 출력 300토큰의 비용은 0.0036달러이므로 입력이 총액의 95%를 차지합니다. 캐시가 완전히 걸리면 건당 0.0836달러에서 0.0116달러로 86% 줄어듭니다. 반대로 출력 편중 워크로드(입력 2,000 / 출력 10,000)에서는 입력이 총액의 3%뿐이라, 같은 90% 할인이 총액을 3%도 못 줄입니다. 같은 정책이 워크로드에 따라 86%와 3%로 갈립니다. 이 계산의 일반형은 LLM API 비용을 실제로 줄이는 법 편에서 공급자별 가격표로 따라갔습니다.

Fast 모드 — 지연 시간에 별도 가격표가 붙었다

Fast 모드는 모델 선택과 직교하는 축이 생겼다는 신호라서 따로 볼 값어치가 있습니다. API 문서 기준으로 service_tier 파라미터에 fast(또는 하위 호환용 priority)를 넣으면 되고, 현재 대상은 gpt-5.6-sol입니다. 롱 컨텍스트, 파인튜닝된 모델, 임베딩은 제외됩니다. 문서는 "표준 대비 최대 2.5배 빠르고 지연 시간이 더 일정하다"고 설명하며, 배수는 명시하지 않고 가격 페이지로 넘깁니다. 발표 자료 기준 프리미엄은 표준가의 2배입니다.

계산은 간단합니다. 2배 값을 내고 2.5배 속도를 사는 거래이므로, 지연 시간 1초의 가치가 그 요청 토큰 비용의 절반보다 큰 경우에만 이득입니다. 사용자가 화면 앞에서 기다리는 대화형 경로라면 대체로 성립하고, 야간 배치나 큐에 쌓아 두는 파이프라인이라면 절대 성립하지 않습니다. 이 구분을 코드 레벨에서 하지 않고 프로젝트 설정으로 일괄 적용하면 배치 트래픽까지 2배 값을 내게 됩니다. 문서가 요청 단위 설정과 프로젝트 단위 설정을 모두 제공한다는 점을 확인하고, 기본값은 표준으로 두는 편이 안전합니다.

"지능의 변화 없이"라는 문구도 그대로 받아들일 만합니다. 다만 이 문장이 보장하는 것은 모델 가중치가 같다는 것이지, 부하 상황에서 꼬리 지연이 어떻게 되는지는 아닙니다. 일정한 지연이 필요해서 이 옵션을 사는 것이라면 평균이 아니라 p99를 재세요.

벤치마크가 포착하지 못하는 것

마지막으로 이 글의 숫자들이 못 담는 것을 정리합니다. 순위가 벤치마크마다 뒤집힌다는 사실부터가 그렇습니다. 공개된 집계를 보면 SWE-Bench Pro에서 Sol은 64.6%, Terra는 63.4%인데 Claude Fable 5는 80.0%로 앞섭니다. 반면 Terminal-Bench 2.1에서는 Sol이 88.8%(ultra 모드 91.9%)로 최고 수준이고, Agents' Last Exam에서는 53.6으로 Fable 5를 13.1점 앞섰다고 보고됩니다. 세 벤치마크가 세 가지 순위를 냅니다. 여기서 "누가 1등인가"를 뽑아내려는 시도 자체가 잘못된 질문입니다.

그리고 어느 벤치마크도 재지 않는 항목들이 있습니다.

  • 긴 세션에서의 지시 준수. 대부분의 벤치마크는 단발 과제이고, 실서비스는 수십 턴짜리 세션입니다. 20턴째에 시스템 프롬프트의 제약을 잊는지는 지수에 들어가지 않습니다.
  • 도구 호출 스키마의 신뢰도. 스키마를 어긋나게 채우는 비율은 에이전트 파이프라인의 실질 실패율을 결정하는데, 집계 점수에는 성공/실패로만 반영됩니다.
  • 내 도메인에서의 거부율과 과잉 안전. 의료·금융·보안 도메인에서 정상 요청이 거부되는 비율은 벤더가 공개하지 않고 벤치마크도 재지 않습니다.
  • 꼬리 지연과 레이트 리밋. 평균 응답 속도는 어디에나 있지만, 트래픽이 몰릴 때의 p99와 429 비율은 직접 재야만 알 수 있습니다.
  • 측정 조건과 실제 설정의 괴리. 위 표의 지수 점수는 모두 최대 추론 강도에서 잰 값입니다. 프로덕션에서 그 강도를 쓰는 트래픽은 많지 않고, 강도를 낮췄을 때의 점수 곡선은 공개돼 있지 않습니다.

마치며 — 곡선은 계속 내려가지만, 내 자리는 내가 재야 안다

지난 3주 동안 GPT-5.6의 아래 두 티어 가격은 80%와 20% 내려갔고, 위 티어는 그대로인 채 속도만 따로 팔리기 시작했습니다. 가격 곡선은 앞으로도 내려갈 것이고, 그때마다 인하 소식이 나올 것입니다. 하지만 이 글에서 계산한 세 숫자는 인하와 무관하게 남습니다.

첫째, 인하 후 세 티어의 단가 비율은 두 축 모두 25 대 10 대 1로 고정되어 있으므로, 계열 안에서 토큰 믹스는 티어 선택에 아무 정보를 주지 않습니다. 둘째, 그래서 실제 결정은 오답 하나를 고치는 비용이 정하고, 예시 워크로드에서 그 손익분기는 건당 약 0.52달러였습니다. 셋째, 캐싱은 입력이 청구서를 지배하는 워크로드에서 86%를 깎지만 출력 편중 워크로드에서는 3%도 못 깎습니다.

세 계산 모두 벤더가 대신 해 줄 수 없습니다. 필요한 입력은 자기 로그에 있는 세 숫자 — 요청당 입력·출력 토큰의 중앙값, 프리픽스별 30분 내 재사용률, 오답 한 건을 처리하는 실제 비용입니다. 다음 인하 소식이 올 때, 읽어야 할 것은 헤드라인이 아니라 그 세 숫자를 다시 넣은 표입니다.

GPT-5.6 and the Limits of Price-Performance — How to Find Your Workload's Place on the Curve

Introduction — What an 80% Cut Three Weeks After Launch Tells You

On July 30, 2026, OpenAI cut the price of GPT-5.6 Luna by 80%. It went from $1 input and $6 output per 1M tokens to $0.20 input and $1.20 output. The middle tier, Terra, came down 20% to $2 input and $12 output, and the top tier, Sol, stays at $5 input and $30 output. The GPT-5.6 family went generally available on July 9, so this is an adjustment three weeks in.

A second item announced the same day makes the character of the move clearer. The existing Priority processing was renamed Fast mode, and for Sol only it offers up to 2.5x the standard speed at twice the standard price. OpenAI's API docs describe it as speed only, with no change in intelligence. In other words, the lower tiers cut price to defend volume, while the upper tier has begun selling latency as a separate product.

There is one number frequently cited as background. Several outlets citing CNBC reporting said that Chinese models took 46% of US enterprise token usage on OpenRouter. I could not verify how that figure was originally compiled, so I treat it strictly as a reported claim. That said, the asymmetry of the cut (80% at the bottom, 0% at the top) hardly needs an explanation of what it is aimed at.

The subject of this post is not the news of a price cut. Cost per capability has been falling along a curve for years, and any individual cut is one point on that curve. There is only one useful question — where on that curve does my workload sit, and did this move actually shift its position?

What Changed Between July 9 and July 30

Putting the metrics Artificial Analysis compiled at launch into one table alongside this cut gives the following. The post-cut values for cost per index task are numbers I computed by applying the announced cut rates directly, assuming identical token consumption.

TierLaunch price (input/output, per 1M tokens)After July 30Intelligence IndexCoding Agent IndexCost per index task (launch → after cut)
Sol$5 / $30unchanged5980$1.04$1.04
Terra$2.50 / $15$2 / $125577$0.55 → about $0.44
Luna$1 / $6$0.20 / $1.205175$0.21 → about $0.042

On the same index, Claude Fable 5 scores 60 at maximum effort. So Artificial Analysis's summary at launch was that Sol trails by one point on composite intelligence while costing about a third as much to finish the index.

Reading the table vertically is more interesting. The intelligence index steps down 59, 55, 51 — four points at a time — while cost per index task drops 1.04, 0.44, 0.042, an order of magnitude at a time. What it costs to buy four points of intelligence index differs tenfold from band to band. This is why the word curve is the right one. At the top end, a tiny amount of capability costs an enormous amount of money; at the bottom, the same money buys far more capability.

The Slope of the Curve — And a Vendor Efficiency Claim

Citations about how fast cost per capability is falling are everywhere, but most of them have murky sources. The version cited relatively often is that, holding a given benchmark level fixed, the price of that performance has been falling by a single-digit to tenfold factor per year. This cut is one point on that trend line, and only the three-week interval is unusual; the direction itself is not new.

One vendor claim here is worth checking. At the GPT-5.6 launch, OpenAI said in effect that token efficiency had improved 54% on certain tasks, and several outlets carried it. That is a vendor claim with the qualifier "certain tasks" attached. The independent aggregation, by contrast, is far more modest — by Artificial Analysis's count, Sol used 15,000 output tokens to complete one index task, about 6% fewer than the 16,000 of the previous generation, GPT-5.5.

The two numbers do not contradict each other. The 54% came from selected tasks and the 6% from the composite index as a whole. But the one you can put into a practical plan is the latter. When a vendor attaches "on certain tasks," that qualifier is itself a do-not-generalize sign. Token consumption on your own tasks has to be measured directly, and that measurement is exactly what must happen before any model swap.

Compute From the Mix, Not the Headline Price — Except This Time the Mix Does Not Change Anything

The usual first trap when comparing two models is looking only at the input price. Because the output-to-input multiple differs by vendor, the cost ranking of the very same two models flips depending on the workload's token mix.

But this cut eliminated that trap inside the GPT-5.6 family. Line up the three tiers' post-cut prices and input is 5, 2, 0.20 while output is 30, 12, 1.20. Both axes are exactly 25 to 10 to 1.

Sol   : input 5    output 30      -> 25x Luna
Terra : input 2    output 12      -> 10x Luna
Luna  : input 0.20 output 1.20    -> baseline

Whatever (input tokens, output tokens) pair you plug in, the total ratio stays fixed at 25 : 10 : 1.
The mix gives you no information at all about which tier to pick.

If the mix cannot change the ranking within the family, tier selection has to be decided on a different axis. That axis is the cost of being wrong. Let us compute it on a real workload.

Workload — classifying 1,000,000 product descriptions
  4,000 input tokens / 300 output tokens per request

  Luna  : (4,000 x 0.20 + 300 x 1.20) / 1e6 = 0.00116 dollars/item -> 1,160 dollars
  Terra : 10x the above                                            -> 11,600 dollars
  Sol   : 25x the above                                            -> 29,000 dollars

  Now add c, the per-item cost of a human fixing a wrong answer.
  Assume a 4% error rate for Luna and 2% for Terra:

  Luna  total =  1,160 + 0.04 x 1e6 x c
  Terra total = 11,600 + 0.02 x 1e6 x c

  The point where the two are equal:
  10,440 = 0.02 x 1e6 x c   ->   c = 0.522 dollars

Here is how to read it. If fixing one wrong answer costs less than $0.52, Luna wins; if it costs more, Terra wins. If the task takes a human one minute per item, labor cost is already far past $0.52. Conversely, in a pipeline where a wrong answer is resolved by a simple retry, the retry costs $0.00116 and Luna wins overwhelmingly.

The lesson from this calculation is simple. The regime where a 10x difference in token price shows up as a difference in total cost is only the regime where the cost of being wrong is the same order of magnitude as the token cost. In most in-house workflows the cost of being wrong is two or three orders of magnitude larger than the token cost. Spending hours comparing price sheets on that kind of workload is optimizing on the wrong axis.

Where Caching Moves the Curve, and Where It Cannot

The GPT-5.6 family's cache policy is a 90% read discount, a write at 1.25x the input price, and a 30-minute minimum retention. Let us compute the case of a 40,000-token fixed prefix on Terra.

40,000-token prefix, Terra input 2 dollars per 1M

  no cache         : 40,000 x 2.00 / 1e6 = 0.080 dollars (every call)
  cache write (1x) : 40,000 x 2.50 / 1e6 = 0.100 dollars
  cache read       : 40,000 x 0.20 / 1e6 = 0.008 dollars

  Break-even over N calls:
  0.100 + (N-1) x 0.008  <=  N x 0.080
  0.092 <= 0.072 N   ->   N >= 1.28   ->  pays off from the second call on

Break-even is the second call. The real constraint is not the break-even point but the 30 minutes. If the same prefix does not come back within 30 minutes, you pay the write premium and collect nothing. So what decides whether caching actually works is not the discount rate but request density per prefix. In a service where 10,000 users each call twice a day with a different prefix, caching increases your cost.

And it bears repeating that a discount rate is not a savings rate. In the workload above, the 300 output tokens cost $0.0036, so input is 95% of the total. With the cache fully warm, per-item cost falls from $0.0836 to $0.0116, a 86% cut. In an output-heavy workload (2,000 input / 10,000 output), by contrast, input is only 3% of the total, so the same 90% discount cannot take even 3% off the total. The same policy splits into 86% and 3% depending on the workload. The general form of this calculation is followed through provider by provider in How to Actually Cut Your LLM API Bill.

Fast Mode — Latency Now Has Its Own Price Sheet

Fast mode is worth looking at separately because it signals that an axis orthogonal to model selection now exists. Per the API docs, you put fast (or priority, kept for backward compatibility) in the service_tier parameter, and the current target is gpt-5.6-sol. Long context, fine-tuned models, and embeddings are excluded. The docs say it is up to 2.5x faster than standard with more consistent latency, and they do not state the multiplier, deferring to the pricing page. Per the announcement materials, the premium is twice the standard price.

The arithmetic is simple. You pay 2x for 2.5x speed, so it only pays off when one second of latency is worth more than half the token cost of that request. For an interactive path where a user is waiting in front of a screen this generally holds; for an overnight batch or a pipeline that piles work into a queue it never does. If you make that distinction not at the code level but as a project-wide setting, your batch traffic pays 2x as well. Confirm that the docs offer both per-request and per-project settings, and it is safer to leave the default at standard.

The phrase about no change in intelligence is also fair to take at face value. What that sentence guarantees, though, is that the model weights are the same — not what tail latency does under load. If you are buying this option because you need consistent latency, measure p99, not the average.

What the Benchmarks Do Not Capture

Finally, let me lay out what the numbers in this post cannot cover. Start with the fact that rankings flip from benchmark to benchmark. In the published aggregations, on SWE-Bench Pro Sol is 64.6% and Terra is 63.4% while Claude Fable 5 leads at 80.0%. On Terminal-Bench 2.1, by contrast, Sol is at the top at 88.8% (91.9% in ultra mode), and on Agents' Last Exam it is reported at 53.6, ahead of Fable 5 by 13.1 points. Three benchmarks produce three rankings. The attempt to extract "who is number one" from this is itself the wrong question.

And there are items that no benchmark measures at all.

  • Instruction adherence over a long session. Most benchmarks are one-shot tasks, while a real service is a session of dozens of turns. Whether it forgets a system-prompt constraint on turn 20 does not enter the index.
  • Reliability of tool-call schemas. The rate at which a model fills a schema incorrectly determines the real failure rate of an agent pipeline, but an aggregate score reflects it only as success or failure.
  • Refusal rates and over-safety in my domain. The rate at which legitimate requests are refused in medical, financial, and security domains is not published by vendors and not measured by benchmarks.
  • Tail latency and rate limits. Average response speed is available everywhere, but p99 and the 429 rate under a traffic surge can only be learned by measuring them yourself.
  • The gap between measurement conditions and your real settings. Every index score in the table above was measured at maximum reasoning effort. Not much production traffic runs at that effort, and the score curve at lower effort is not published.

Closing — The Curve Keeps Falling, but Only You Can Measure Where You Sit on It

Over the past three weeks, the prices of GPT-5.6's lower two tiers came down 80% and 20%, and the top tier stayed put while its speed started being sold separately. The price curve will keep going down, and each time there will be news of a cut. But the three numbers computed in this post survive independent of any cut.

First, after the cut the three tiers' unit prices are pinned at 25 to 10 to 1 on both axes, so within the family the token mix gives you no information about tier selection. Second, the actual decision is therefore set by the cost of fixing one wrong answer, and in the example workload that break-even was about $0.52 per item. Third, caching cuts 86% on a workload where input dominates the invoice but cannot cut even 3% on an output-heavy one.

No vendor can do any of these three calculations for you. The inputs you need are three numbers sitting in your own logs — the median input and output tokens per request, the within-30-minutes reuse rate per prefix, and the real cost of handling a single wrong answer. When the next price-cut announcement arrives, the thing to read is not the headline but the table with those three numbers plugged back in.