OpenAI

OpenAI Launches FrontierScience to Benchmark Expert-Level AI Scientific Reasoning


Executive Summary:

OpenAI has introduced FrontierScience, a new, advanced benchmark designed to evaluate the expert-level scientific reasoning of AI models across physics, chemistry, and biology. Created and verified by PhD scientists and international science Olympiad medalists, the benchmark aims to address the limitations of existing tests that have become saturated or lack focus. By providing difficult, original problems, FrontierScience is intended to serve as a "north star" for measuring progress and guiding the development of AI to accelerate real-world scientific discovery.

Key Takeaways:

* Product Name: FrontierScience.

* Primary Function: A benchmark consisting of over 700 difficult, expert-written questions to measure AI's scientific reasoning capabilities.

* Two-Track Structure:

* FrontierScience-Olympiad: 100 questions from Olympiad medalists testing constrained, short-answer scientific reasoning.

* FrontierScience-Research: 60 multi-step subtasks from PhD scientists mimicking real-world research problems, graded on a 10-point rubric.

* Initial Performance: The company's latest model, GPT-5.2, is the top performer, scoring 77% on the Olympiad track and 25% on the Research track, highlighting significant headroom for improvement, especially in open-ended tasks.

* Availability: A "gold set" of 160 questions (100 Olympiad, 60 Research) is being open-sourced to the research community.

* Scalable Grading: The benchmark uses a model-based grader (GPT-5) to evaluate responses against the rubric, enabling large-scale and efficient testing.

Strategic Importance:

This launch establishes the company as a leader in both developing and evaluating frontier AI for science. It pushes the industry beyond simpler benchmarks toward measuring the complex, open-ended reasoning required for AI to contribute to genuine scientific breakthroughs.

Original article