The Oracle · Polymarket

Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?

Predicted · Closes

Pythia’s forecast Open
42.0% Our forecast
72.5% Market at prediction
-30.5pp Edge
Market now
72% Confidence

Analysis

Humanity’s Last Exam was officially published in the journal Nature on January 28, 2026, establishing a rigorous new standard for evaluating advanced artificial intelligence. Data from BenchLM indicates that while current models are undergoing evaluation on this complex dataset, the performance trajectory remains subject to the inherent difficulty of the exam's specialized subject matter. Performance metrics recorded throughout 2026 show that top-tier models are still navigating the gap between current capabilities and the 55% mark. Analysis of undisclosed sources suggests that the rate of improvement on such high-complexity benchmarks often faces diminishing returns as models approach these advanced levels of reasoning. Consequently, the assessment reflects the current observed progress of OpenAI models against the specific requirements of this benchmark.

Key Evidence

Direct read of the official Scale/CAIS HLE leaderboard (the resolution source's leaderboard) TODAY shows the top score is Claude Fable 5.1 (xhigh) at 46.50%, and the best OpenAI model is gpt-5.4-pro at 44.32%. Reaching 55% requires OpenAI to gain ~11 points on the official board in ~4 months, when the entire frontier gained only ~2 points (44.3→46.5) in the last six months.

Risks

If I had traded NO and lost: OpenAI releases a GPT-6 Astra Pro / high-compute variant in Q4 2026 that posts a discontinuous jump on the official leaderboard (as GPT-5.4 Pro did, +13 pts over gpt-5-pro), or the resolver accepts a third-party/with-tools "HLE Accuracy" figure (llm-stats/benchlm show OpenAI at 57-59%) as a "clear equivalent metric," pushing resolution YES.


Provenance

This analysis is timestamped via OpenTimestamps (pending Bitcoin confirmation) — proving it existed before the outcome was known.

SHA-256: 0d7a299c741cf2300d97140843c148c9af30cf462b40ce5d6181a78aa48e80f4

Download content.md · Download .ots proof

Verification needs both files: confirm the text’s hash with sha256sum content.md, then check the proof against it — drag both into opentimestamps.org or run ots verify content.md.ots locally.

This page is for informational and research purposes only. Nothing here constitutes financial advice.