Recent releases of frontier models have accelerated progress on math benchmarks tracked by leaderboards like MathArena and the Chatbot Arena Math category. OpenAI’s GPT-6 Astra (max), released September 4, 2026, leads with expected performance near 91% across uncontaminated competitions, while Anthropic’s Claude-Opus-5 (max) and Claude-Fable-5.1 (max) sit close behind at roughly 72%. Alibaba’s Qwen3.8-Max and Google’s latest Gemini Flash variants also posted gains in early September evaluations. These updates reflect ongoing scaling, improved reasoning architectures, and targeted post-training on competition-style problems. With multiple labs planning further releases before year-end, traders are weighing the pace of incremental Elo or accuracy gains against historical saturation trends on final-answer math tasks.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$119,150 Vol.
1575
67%
1600
29%
$119,150 Vol.
1575
67%
1600
29%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Market Opened: Apr 2, 2026, 6:07 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent releases of frontier models have accelerated progress on math benchmarks tracked by leaderboards like MathArena and the Chatbot Arena Math category. OpenAI’s GPT-6 Astra (max), released September 4, 2026, leads with expected performance near 91% across uncontaminated competitions, while Anthropic’s Claude-Opus-5 (max) and Claude-Fable-5.1 (max) sit close behind at roughly 72%. Alibaba’s Qwen3.8-Max and Google’s latest Gemini Flash variants also posted gains in early September evaluations. These updates reflect ongoing scaling, improved reasoning architectures, and targeted post-training on competition-style problems. With multiple labs planning further releases before year-end, traders are weighing the pace of incremental Elo or accuracy gains against historical saturation trends on final-answer math tasks.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions