OpenAI’s GPT-5.4 Pro has reached 58.7% on Humanity’s Last Exam as of early September 2026, trailing Anthropic’s leading Claude Fable 5.1 at 65% and Opus 5 at 64.7% on the expert-curated 2,500-question benchmark. Recent gains for OpenAI stem from iterative GPT-5 releases incorporating stronger chain-of-thought scaffolding and tool integration, yet Anthropic maintains an edge in multidisciplinary reasoning and adaptive performance across mathematics, sciences, and humanities. Trader sentiment centers on whether further 2026 GPT variants or new architectures can close this gap before year-end, with key catalysts including additional model launches, benchmark protocol updates distinguishing tool-assisted versus closed-book scores, and any competitive shifts in frontier capability demonstrations. The market reflects aggregated views on OpenAI’s release cadence and technical progress relative to rivals.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$76,741 Vol.
50%+
90%
55%+
78%
60%+
38%
65%+
19%
70%+
11%
$76,741 Vol.
50%+
90%
55%+
78%
60%+
38%
65%+
19%
70%+
11%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:53 PM ET
Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolver
0x65070BE91...OpenAI’s GPT-5.4 Pro has reached 58.7% on Humanity’s Last Exam as of early September 2026, trailing Anthropic’s leading Claude Fable 5.1 at 65% and Opus 5 at 64.7% on the expert-curated 2,500-question benchmark. Recent gains for OpenAI stem from iterative GPT-5 releases incorporating stronger chain-of-thought scaffolding and tool integration, yet Anthropic maintains an edge in multidisciplinary reasoning and adaptive performance across mathematics, sciences, and humanities. Trader sentiment centers on whether further 2026 GPT variants or new architectures can close this gap before year-end, with key catalysts including additional model launches, benchmark protocol updates distinguishing tool-assisted versus closed-book scores, and any competitive shifts in frontier capability demonstrations. The market reflects aggregated views on OpenAI’s release cadence and technical progress relative to rivals.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions