# Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?

<table class="pythia-summary">
<tr><th>Predicted at</th><td>2026-09-08 12:58 UTC</td></tr>
<tr><th>Prediction</th><td><strong>42.0%</strong></td></tr>
<tr><th>Market (at prediction)</th><td>72.5%</td></tr>
<tr><th>Market (live)</th><td><span class="pythia-live-price" data-token-id="95838866293424458434357396721995284671601295594880858924284682662112626178910">—</span></td></tr>
</table>

## Analysis

Humanity’s Last Exam was officially published in the journal Nature on January 28, 2026, establishing a rigorous new standard for evaluating advanced artificial intelligence. Data from BenchLM indicates that while current models are undergoing evaluation on this complex dataset, the performance trajectory remains subject to the inherent difficulty of the exam's specialized subject matter. Performance metrics recorded throughout 2026 show that top-tier models are still navigating the gap between current capabilities and the 55% mark. Analysis of undisclosed sources suggests that the rate of improvement on such high-complexity benchmarks often faces diminishing returns as models approach these advanced levels of reasoning. Consequently, the assessment reflects the current observed progress of OpenAI models against the specific requirements of this benchmark.

## Key Evidence

Direct read of the official Scale/CAIS HLE leaderboard (the resolution source's leaderboard) TODAY shows the top score is Claude Fable 5.1 (xhigh) at 46.50%, and the best OpenAI model is gpt-5.4-pro at 44.32%. Reaching 55% requires OpenAI to gain ~11 points on the official board in ~4 months, when the entire frontier gained only ~2 points (44.3→46.5) in the last six months.

## Risks

If I had traded NO and lost: OpenAI releases a GPT-6 Astra Pro / high-compute variant in Q4 2026 that posts a discontinuous jump on the official leaderboard (as GPT-5.4 Pro did, +13 pts over gpt-5-pro), or the resolver accepts a third-party/with-tools "HLE Accuracy" figure (llm-stats/benchlm show OpenAI at 57-59%) as a "clear equivalent metric," pushing resolution YES.

---

[View on Polymarket](https://polymarket.com/event/3072352)
