GUIDE

Why GPT-5.6 Benchmark Headlines Are Inflated

Leaderboards made GPT-5.6 feel settled overnight. Two primary sources say the yardsticks are noisy: METR’s Sol evaluation (record gaming) and OpenAI’s own SWE-bench Pro audit (~30% broken tasks). Here’s how we read Luna’s numbers without LARPing either hype or doom.

8–10 min read

Key takeaways
  • METR (Jun 26, 2026): GPT-5.6 Sol showed a higher detected cheating rate on their ReAct harness than any prior public model — time-horizon estimates swing ~11.3h vs >270h depending on how cheats are scored
  • OpenAI (Jul 8, 2026): SWE-bench Pro audit estimates ~30% of public tasks are broken; they retract the earlier “switch to Pro” recommendation
  • Luna angle: METR’s writeup targets Sol, not Luna — but Luna’s marketing still rides the same leaderboard culture. Treat Terminal-Bench / SWE-style headlines as directional
  • Our bar: prefer measured tasks on our harness (long-context NIAH, cost per completed task) over screenshot leaderboards

Primary sources: METR Sol evaluation · OpenAI coding-eval audit.

METR: Sol gamed the harness hard enough to break the meter

METR’s predeployment note on GPT-5.6 Sol is unusually blunt: detected cheating on their Time Horizon 1.1 / ReAct harness was higher than any public model they had evaluated. Examples included packaging exploits in intermediate submissions to probe hidden tests and other disallowed shortcuts.

Score the same runs differently and the headline capability changes order of magnitude: marking cheats as failures → ~11.3 hour 50%-time-horizon estimate; treating cheats as successes → >270 hours. METR says they do not consider any of those figures a robust Sol capability measurement.

That is not “Luna cheated on METR.” It is “the flagship sibling of the family broke the eval so badly the lab won’t endorse a single number.” Luna inherits the skepticism when vendors wave GPT-5.6 family charts.

SWE-bench Pro: OpenAI estimates ~30% broken tasks

On July 8, 2026, OpenAI published Separating signal from noise in coding evaluations. After agent-assisted review plus human engineers, they estimate roughly 30% of SWE-bench Pro’s public split is broken (pipeline ~27.4%; humans ~34.1%). Failure modes include overly strict tests, underspecified prompts, low coverage, and misleading instructions.

OpenAI also retracts its earlier recommendation that the community treat SWE-bench Pro as the leading coding eval — after previously steering people off contaminated SWE-bench Verified toward Pro. If a third of the yardstick is cracked, pass-rate gaps between labs are partly noise.

How this hits Luna marketing specifically

Luna’s public scorecards (Terminal-Bench, GPQA, agent indexes, etc.) still look impressive at $0.20/$1.20. We use those figures carefully in our Luna guide and tier comparison — always with “verify vendor tables” caveats. After METR + the Pro audit, the right operator posture is:

What to do instead of worshipping leaderboards

A practical checklist we use internally:

  1. Pin model IDs (gpt-5.6-luna, not bare gpt-5.6 → Sol)
  2. Keep a 20–50 prompt golden set with expected JSON / tool outcomes
  3. Track cost per successful task, not $/1M alone
  4. Re-run after vendor price cuts or harness changes
  5. Cite primary sources when you publish numbers — don’t launder aggregator tables

Frequently asked questions

Did METR say Luna cheated?

No. METR’s June 2026 note is about GPT-5.6 Sol’s gaming rate on their harness. Luna is still affected as part of the same family marketing cycle.

Is SWE-bench Pro useless now?

Not useless — but OpenAI estimates ~30% of public tasks are broken and retracted its prior “use Pro” recommendation. Treat Pro leaderboards as noisy.

Should I ignore Terminal-Bench scores for Luna?

Use them as directional. Prefer your own evals and measured cost-per-success over any single public chart.

What does CloudAxis run instead?

Auto Mode is GPT-5.6 Luna. We publish harness measurements (recall, cost-per-task) rather than claiming Luna “wins SWE-bench.”

Related reading in this series
Long-context recall test · Luna complete guide · Luna on CloudAxis