Internal benchmarks circulating in developer circles show GPT-6 posting sizable jumps on reasoning and coding — with familiar tradeoffs.
Benchmark screenshots attributed to OpenAI's unreleased GPT-6 have been circulating in developer circles this week, suggesting a substantial jump over GPT-5 on reasoning and coding evaluations. OpenAI has not confirmed the leaked numbers.
Reported gains cluster in the 15-25% range over GPT-5 on standard benchmarks (SWE-bench, MATH, GPQA). Notably, the leaked results also show GPT-6 requires materially more compute per query — the familiar quality/cost tradeoff.
The leaked benchmarks don't cover long-context retention, tool use reliability, or agentic multi-step tasks — the areas where enterprises are increasingly evaluating models. Numbers on those axes will determine whether GPT-6 is a genuine capability jump or just a benchmark polish.
If real, GPT-6 puts OpenAI back ahead on peak capability. But Anthropic's Claude Opus 4.8 has been eating into GPT-5 share on coding workflows, and Google's Ultra tier hasn't landed yet. The frontier race stays close.
Sources: [Reddit r/singularity], [Twitter developer threads]