#1 OpenAI, o3 announcement (as reported by TechCrunch) said on 2024-12-20 · source
o3 could answer just over a fourth of the questions on FrontierMath.
The evidence supports that OpenAI did announce a score of over 25% for o3 on FrontierMath. However, Epoch AI's independent test of the publicly released model showed a significantly lower score of around 10%. This discrepancy suggests that the initial claim may have been based on an internal version with different settings. Additionally, Epoch AI's funding by OpenAI and its access to the problems raise questions about the independence of the benchmark. These points lead to the conclusion that the claim is partially contradicted.
Citations
- techcrunch.com: “Epoch AI's independent evaluation of the publicly released o3 scored around 10% on FrontierMath; OpenAI's December figure came from an internal version with a more aggressive test-time compute setting and a different problem set.”
- techcrunch.com: “Epoch AI disclosed only on the day of the o3 announcement that OpenAI had funded FrontierMath and had access to most of the problems.”
What would change this
- Direct comparison between the internal and public versions of o3 on the same problem set and settings.
- More information on the nature and extent of OpenAI's involvement with Epoch AI and the development of the FrontierMath benchmark.