# Will OpenAI have the best AI model on LiveBench (Mathematics) at the end of September 2026?

<table class="pythia-summary">
<tr><th>Predicted at</th><td>2026-08-07 10:27 UTC</td></tr>
<tr><th>Prediction</th><td><strong>54.8%</strong></td></tr>
<tr><th>Market (at prediction)</th><td>43.0%</td></tr>
<tr><th>Market (live)</th><td><span class="pythia-live-price" data-token-id="73386669745439334907938712583546331501163723903682397436614379893335316279799">—</span></td></tr>
</table>

## Analysis

As of August 2026, OpenAI currently maintains the top position on the LiveBench Mathematics leaderboard with its o3-mini model. The competitive landscape remains fluid, as evidenced by recent model releases from competitors like Anthropic, which continue to challenge existing performance benchmarks. LiveBench utilizes dynamic updates to its task sets, which helps prevent data contamination and ensures that evaluations reflect real-world reasoning capabilities rather than static memorization. Future performance will depend on the release cycles of next-generation models and how effectively these architectures handle the evolving complexity of mathematical tasks. Proprietary signals suggest that the pace of innovation in model training remains high, potentially shifting leaderboard rankings before the September 2026 deadline.

## Risks

Anthropic ships a new frontier Claude (Fable 5.x or Mythos-derived) in September that scores above 96.2 on LiveBench Math, or a LiveBench question-set refresh before Sep 30 reshuffles the near-saturated top ranks and drops OpenAI from #1 — either flips this razor-thin 0.2-point lead.

---

[View on Polymarket](https://polymarket.com/event/3211834)
