Anthropic's Claude 5-series models currently lead Humanity's Last Exam (HLE) leaderboards, with Claude Fable 5.1 at 65% and closely trailing variants like Opus 5 and Mythos 5 around 64.5-64.7% as of early September 2026, according to BenchLM.ai snapshots. This edge stems from iterative advances in reasoning architectures, extended context handling, and tool-augmented evaluation on the 2,500-question expert-curated benchmark spanning mathematics, physics, biology, and humanities. Meta's Muse Spark 1.1 sits at 62.1%, while OpenAI's GPT-5.4 Pro reaches 58.7%, reflecting a tight frontier cluster that has climbed rapidly from low-40s percent earlier in the year. Remaining headroom before saturation, plus potential late-2026 releases or HLE-Rolling updates, keeps competitive positioning fluid for end-of-year outcomes.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$54,171 Vol.
60%+
59%
65%+
31%
70%+
17%
75%+
12%
80%+
7%
90%+
4%
$54,171 Vol.
60%+
59%
65%+
31%
70%+
17%
75%+
12%
80%+
7%
90%+
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:41 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Anthropic's Claude 5-series models currently lead Humanity's Last Exam (HLE) leaderboards, with Claude Fable 5.1 at 65% and closely trailing variants like Opus 5 and Mythos 5 around 64.5-64.7% as of early September 2026, according to BenchLM.ai snapshots. This edge stems from iterative advances in reasoning architectures, extended context handling, and tool-augmented evaluation on the 2,500-question expert-curated benchmark spanning mathematics, physics, biology, and humanities. Meta's Muse Spark 1.1 sits at 62.1%, while OpenAI's GPT-5.4 Pro reaches 58.7%, reflecting a tight frontier cluster that has climbed rapidly from low-40s percent earlier in the year. Remaining headroom before saturation, plus potential late-2026 releases or HLE-Rolling updates, keeps competitive positioning fluid for end-of-year outcomes.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions