OpenAI o3-mini System Card
This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external red teaming, and Preparedness Framework evaluations.
Auto Added by WPeMatico
This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external red teaming, and Preparedness Framework evaluations.
An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you. Available to Pro users today, Plus and Team next.
Trading Inference-Time Compute for Adversarial Robustness
Trading inference-time compute for adversarial robustness Read More »
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations according to our Preparedness Framework.
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
We’ve simplified, stabilized, and scaled continuous-time consistency models, achieving comparable sample quality to leading diffusion models, while using only two sampling steps.
Simplifying, stabilizing, and scaling continuous-time consistency models Read More »
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Read More »
We’ve analyzed how ChatGPT responds to users based on their name, using AI research assistants to protect privacy.