OpenAI Unveils O3 and O3-Mini Simulated Reasoning Models

OpenAI Unveils Latest AI “Reasoning” Models, o3 and o3-mini

On Friday, during Day 12 of its “12 days of OpenAI,” OpenAI CEO Sam Altman announced its latest AI “reasoning” models, o3 and o3-mini, which build upon the o1 models launched earlier this year. The company is not releasing them yet but will make these models available for public safety testing and research access today.

Private Chain of Thought

The models use what OpenAI calls “private chain of thought,” where the model pauses to examine its internal dialog and plan ahead before responding, which you might call “simulated reasoning” (SR)—a form of AI that goes beyond basic large language models (LLMs).

The o3 Model: A Record-Breaking Achievement

According to OpenAI, the o3 model earned a record-breaking score on the ARC-AGI benchmark, a visual reasoning benchmark that has gone unbeaten since its creation in 2019. In low-compute scenarios, o3 scored 75.7 percent, while in high-compute testing, it reached 87.5 percent—comparable to human performance at an 85 percent threshold.

OpenAI also reported that o3 scored 96.7 percent on the 2024 American Invitational Mathematics Exam, missing just one question. The model also reached 87.7 percent on GPQA Diamond, which contains graduate-level biology, physics, and chemistry questions. On the Frontier Math benchmark by EpochAI, o3 solved 25.2 percent of problems, while no other model has exceeded 2 percent.

Naming the Model: A Cautionary Tale

The company named the model family “o3” instead of “o2” to avoid potential trademark conflicts with British telecom provider O2, according to The Information. During Friday’s livestream, Altman acknowledged his company’s naming foibles, saying, “In the grand tradition of OpenAI being really, truly bad at names, it’ll be called o3.”

Conclusion

The o3 model is a significant achievement in AI research, demonstrating impressive performance on various benchmarks. Its ability to reason and solve complex problems makes it an exciting development in the field of artificial intelligence.

Frequently Asked Questions

Q: What is the private chain of thought concept used in the o3 model?

A: The private chain of thought concept involves the model pausing to examine its internal dialog and plan ahead before responding, a form of simulated reasoning.

Q: What are the benchmark scores achieved by the o3 model?

A: The o3 model scored 75.7 percent in low-compute scenarios and 87.5 percent in high-compute testing on the ARC-AGI benchmark, 96.7 percent on the 2024 American Invitational Mathematics Exam, and 87.7 percent on GPQA Diamond.

Q: Why did OpenAI choose not to name the model “o2”?

A: OpenAI chose not to name the model “o2” to avoid potential trademark conflicts with British telecom provider O2.

Post Views: 87

OpenAI Unveils O3 and O3-Mini Simulated Reasoning Models

OpenAI Unveils Latest AI “Reasoning” Models, o3 and o3-mini

Private Chain of Thought

The o3 Model: A Record-Breaking Achievement

Naming the Model: A Cautionary Tale

Conclusion

Frequently Asked Questions

Q: What is the private chain of thought concept used in the o3 model?

Q: What are the benchmark scores achieved by the o3 model?

Q: Why did OpenAI choose not to name the model “o2”?

Generate single title from this title AgentWatch: Proactive AWS monitoring with ambient agents in 100 -150 characters. And it must return only title i...

How is Technology Changing Recreational Boating?

When robots start to feel: HBK and Siléane bring tactile intelligence to high-speed cosmetics packaging

Generate single title from this title I tested a 4TB quantum-resistant USB drive – but you don’t have to spend $3000 for this much...

Generate single title from this title Data Science • AI • Advanced Analytics in 100 -150 characters. And it must return only title i...

Generate single title from this title AgentWatch: Proactive AWS monitoring with ambient agents in 100 -150 characters. And it must return only title i...

How is Technology Changing Recreational Boating?

When robots start to feel: HBK and Siléane bring tactile intelligence to high-speed cosmetics packaging

Generate single title from this title I tested a 4TB quantum-resistant USB drive – but you don’t have to spend $3000 for this much...

Generate single title from this title Data Science • AI • Advanced Analytics in 100 -150 characters. And it must return only title i...

Strider Robotics demonstrates 40 kg payload quadruped robot as commercial pilots begin

mimic Robotics unveils full-stack platform for dexterous robot manipulation

Aetina expands Nvidia Jetson Thor portfolio with T3000 and T2000 support

LEAVE A REPLY Cancel reply

Latest

Generate single title from this title AgentWatch: Proactive AWS monitoring with ambient agents in 100 -150 characters. And it must return only title i...

How is Technology Changing Recreational Boating?

When robots start to feel: HBK and Siléane bring tactile intelligence to high-speed cosmetics packaging

Categories

Useful Links

Our Newsletter