Meta Says Its AI Models Achieved Gold-Medal-Level Performance Across Top Global STEM Olympiads
Meta said the evaluations were designed to measure pure reasoning ability rather than a model's capacity to retrieve information or use software tools.
Meta has announced that its latest AI reasoning models achieved top-tier results across five of the world's most prestigious STEM Olympiad competitions.
According to the company, its models earned perfect scores on the theory exams of both the Asian Physics Olympiad (APhO) and the International Physics Olympiad (IPhO).
They also secured a gold medal at the International Mathematical Olympiad (IMO) and delivered gold-medal-level performances at the International Chemistry Olympiad (IChO) and the Romanian Master of Mathematics (RMM).
Meta said the evaluations were designed to measure pure reasoning ability rather than a model's capacity to retrieve information or use software tools.
"The types of problems in the Olympiad competitions are exceptionally hard, demanding deep chains of reasoning, creative insight, and flawless argumentation," Meta said.
To ensure the results reflected reasoning alone, Meta prohibited all external assistance during testing. The company also acknowledged the role of the competition organisers and participants in enabling the evaluation.
The announcement comes as leading AI companies increasingly use elite academic competitions to benchmark reasoning capabilities beyond conventional AI tests.
Solving Olympiad-level problems requires multi-step logical reasoning, abstract thinking, and mathematically rigorous proofs—areas long considered among the toughest challenges for AI.
Over the past year, Google DeepMind's Gemini models and OpenAI have reported gold-medal-level performances on the International Mathematical Olympiad, while Anthropic has highlighted Claude's strong results on advanced mathematics and coding benchmarks.
These achievements have become key indicators of progress toward more capable reasoning models, even as researchers caution that success in academic competitions does not necessarily translate into reliable performance across real-world enterprise or scientific applications.
Recently, Meta launched Muse Code in beta, a competitor to apps like OpenAI Codex and Anthropic's Claude. Available for macOS and Linux, Muse Code is built to handle end-to-end engineering workflows, including planning code changes, writing code, debugging and validating results.
Comments ()