The Oracle · Polymarket
Will any xAI Grok model score at least 30% on the FrontierMath Exam?
Predicted · Resolved
Analysis
The assessment of this forecast relies on the observed trajectory of AI performance on the FrontierMath benchmark, which has demonstrated rapid growth from approximately 2% in late 2024 to over 40% by early 2026. Historical data indicates that frontier models frequently achieve relative performance gains of 30% to 100% between major version releases. While current Grok iterations have scored in the 12-14% range, the broader base rate for advanced AI models suggests a high likelihood of reaching the 30% mark within a one-year timeframe. This analysis incorporates the Epoch AI leaderboard data, which tracks the progress of various models on these specific mathematical tiers. Consequently, the forecast reflects the consistent trend of rapid capability expansion in frontier-level language models.
Provenance
This analysis is timestamped via OpenTimestamps — anchored in Bitcoin block 957898 — proving it existed before the outcome was known.
SHA-256: c1f2abdee77384c8733dcefcfe5e8e0000b806f08e1376681355c3dacc9c6a14
Download content.md · Download .ots proof
Verification needs both files: confirm the text’s hash with
sha256sum content.md, then check the proof
against it — drag both into opentimestamps.org
or run ots verify content.md.ots locally.
This page is for informational and research purposes only. Nothing here constitutes financial advice.