Research
224 posts
Ten advances in mathematics and theoretical computer science
OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and c…
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
Scientific computing in the age of agentic AI
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics a…
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
Introducing GeneBench-Pro
Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
OpenAI and Molecule.one show how a near-autonomous AI chemist using GPT-5.4 improved a key drug-making reaction, advancing medicinal chemistry research.
Introducing LifeSciBench
Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decis…
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluatio…
Dreaming: Better memory for a more helpful ChatGPT
ChatGPT introduces a new memory system to better remember preferences, keeping context fresh and relevant across conversations.
An OpenAI model has disproved a central conjecture in discrete geometry
An OpenAI model solved the 80-year-old unit distance problem, disproving a major conjecture in discrete geometry and marking a milestone in AI-driven mathema…
What Parameter Golf taught us about AI-assisted research
Parameter Golf brought together 1,000+ participants and 2,000+ submissions to explore AI-assisted machine learning research, coding agents, quantization, and…
Where the goblins came from
How goblin outputs spread in AI models: timeline, root cause, and fixes behind personality-driven quirks in GPT-5 behavior.
Introducing OpenAI Privacy Filter
OpenAI Privacy Filter is an open-weight model for detecting and redacting personally identifiable information (PII) in text with state-of-the-art accuracy
Introducing GPT-Rosalind for life sciences research
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protein reasoning, and scientific research …
Inside our approach to the Model Spec
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
Improving instruction hierarchy in frontier LLMs
IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.
GPT-5.4 Thinking System Card
[Read on openai.com](https://openai.com/index/gpt-5-4-thinking-system-card)
Reasoning models struggle to control their chains of thought, and that’s good
OpenAI introduces CoT-Control and finds reasoning models struggle to control their chains of thought, reinforcing monitorability as an AI safety safeguard.
Extending single-minus amplitudes to gravitons
A new preprint extends single-minus amplitudes to gravitons, with GPT-5.2 Pro helping derive and verify nonzero graviton tree amplitudes in quantum gravity.
GPT-5.3 Instant System Card
[Read on openai.com](https://openai.com/index/gpt-5-3-instant-system-card)
Why we no longer evaluate SWE-bench Verified
SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend …
Our First Proof submissions
We share our AI model’s proof attempts for the First Proof math challenge, testing research-grade reasoning on expert-level problems.
Introducing EVMbench
OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-severity smart contract vulnerabilities.
GPT-5.2 derives a new result in theoretical physics
A new preprint shows GPT-5.2 proposing a new formula for a gluon amplitude, later formally proved and verified by OpenAI and academic collaborators.
GPT-5 lowers the cost of cell-free protein synthesis
An autonomous lab combining OpenAI’s GPT-5 with Ginkgo Bioworks’ cloud automation cut cell-free protein synthesis costs by 40% through closed-loop experiment…
GPT-5.3-Codex System Card
GPT‑5.3-Codex is the most capable agentic coding model to date, combining the frontier coding performance of GPT‑5.2-Codex with the reasoning and professiona…
Evaluating chain-of-thought monitorability
OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations across 24 environments. Our findings show…
Addendum to GPT-5.2 System Card: GPT-5.2-Codex
[Read on openai.com](https://openai.com/index/gpt-5-2-codex-system-card)
Evaluating AI’s ability to perform scientific research tasks
OpenAI introduces FrontierScience, a benchmark testing AI reasoning in physics, chemistry, and biology to measure progress toward real scientific research.
Measuring AI’s capability to accelerate biological research
OpenAI introduces a real-world evaluation framework to measure how AI can accelerate biological research in the wet lab. Using GPT-5 to optimize a molecular …
Advancing science and math with GPT-5.2
GPT-5.2 is OpenAI’s strongest model yet for math and science, setting new state-of-the-art results on benchmarks like GPQA Diamond and FrontierMath. This pos…
Update to GPT-5 System Card: GPT-5.2
GPT-5.2 is the latest model family in the GPT-5 series. The comprehensive safety mitigation approach for these models is largely the same as that described i…
How confessions can keep language models honest
OpenAI researchers are testing “confessions,” a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, tr…
Early experiments in accelerating science with GPT-5
OpenAI introduces the first research cases showing how GPT-5 accelerates scientific progress across math, physics, biology, and computer science. Explore how…
How evals drive the next chapter in AI for businesses
Learn how evals help businesses define, measure, and improve AI performance—reducing risk, boosting productivity, and driving strategic advantage.
GPT-5.1-Codex-Max System Card
This system card outlines the comprehensive safety measures implemented for GPT‑5.1-CodexMax. It details both model-level mitigations, such as specialized sa…
Understanding neural networks through sparse circuits
OpenAI is exploring mechanistic interpretability to understand how neural networks reason. Our new sparse model approach could make AI systems more transpare…
GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum
This GPT-5 system card addendum provides updated safety metrics for GPT-5.1 Instant and Thinking, including new evaluations for mental health and emotional r…
Introducing IndQA
OpenAI introduces IndQA, a new benchmark for evaluating AI systems in Indian languages. Built with domain experts, IndQA tests cultural understanding and rea…
Defining and evaluating political bias in LLMs
Learn how OpenAI evaluates political bias in ChatGPT through new real-world testing methods that improve objectivity and reduce bias.
Sora 2 is here
Our latest video generation model is more physically accurate, realistic, and controllable than prior systems. It also features synchronized dialogue and sou…
Sora 2 System Card
Sora 2 is our new state of the art video and audio generation model. Building on the foundation of Sora, this new model introduces capabilities that have bee…
Measuring the performance of our models on real-world tasks
OpenAI introduces GDPval, a new evaluation that measures model performance on real-world economically valuable tasks across 44 occupations.
Detecting and reducing scheming in AI models
Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across…
How people are using ChatGPT
New research from the largest study of ChatGPT use shows how the tool creates economic value through both personal and professional use. Adoption is broadeni…
Why language models hallucinate
OpenAI’s new research explains why language models hallucinate. The findings show how improved evaluations can enhance AI reliability, honesty, and safety.
Collective alignment: public input on our Model Spec
OpenAI surveyed over 1,000 people worldwide on how AI should behave and compared their views to our Model Spec. Learn how collective alignment is shaping AI …
Accelerating life sciences research
Discover how a specialized AI model, GPT-4b micro, helped OpenAI and Retro Bio engineer more effective proteins for stem cell therapy and longevity research.
GPT-5 System Card
This GPT-5 system card explains how a unified model routing system powers fast and smart responses using gpt-5-main, gpt-5-thinking, and lightweight versions…
gpt-oss-120b & gpt-oss-20b Model Card
We introduce gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models available under the Apache 2.0 license and our gpt-oss usage policy.
Pioneering an AI clinical copilot with Penda Health
OpenAI and Penda Health debut an AI clinical copilot that cuts diagnostic errors by 16% in real-world use—offering a new path for safe, effective AI in healt…
ChatGPT agent System Card
ChatGPT agent System Card: OpenAI’s agentic model unites research, browser automation, and code tools with safeguards under the Preparedness Framework.
Toward understanding and preventing misalignment generalization
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one tha…
Introducing HealthBench
HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims …
OpenAI o3 and o4-mini System Card
OpenAI o3 and OpenAI o4-mini combine state-of-the-art reasoning with full tool capabilities—web browsing, Python, image and file analysis, image generation, …
Our updated Preparedness Framework
Sharing our updated framework for measuring and protecting against severe harm from frontier AI capabilities.
BrowseComp: a benchmark for browsing agents
BrowseComp: a benchmark for browsing agents.
PaperBench: Evaluating AI’s Ability to Replicate AI Research
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
Addendum to GPT-4o System Card: 4o image generation
4o image generation is a new, significantly more capable image generation approach than our earlier DALL·E 3 series of models. It can create photorealistic o…
Early methods for studying affective use and emotional well-being on ChatGPT
An OpenAI and MIT Media Lab Research collaboration.
Detecting misbehavior in frontier reasoning models
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing…
OpenAI GPT-4.5 System Card
We’re releasing a research preview of OpenAI GPT‑4.5, our largest and most knowledgeable model yet.
Introducing the SWE-Lancer benchmark
Can frontier LLMs earn $1 million from real-world freelance software engineering?
Introducing deep research
An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you. Available to Pro users today, …
OpenAI o3-mini System Card
This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external red teaming, and Preparedness Framework…
OpenAI o3-mini
[Read on openai.com](https://openai.com/index/openai-o3-mini)
Computer-Using Agent
[Read on openai.com](https://openai.com/index/computer-using-agent)
Trading inference-time compute for adversarial robustness
Trading Inference-Time Compute for Adversarial Robustness
OpenAI o1 System Card
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations accor…
Advancing red teaming with people and AI
Advancing red teaming with people and AI
Introducing SimpleQA
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Simplifying, stabilizing, and scaling continuous-time consistency models
We’ve simplified, stabilized, and scaled continuous-time consistency models, achieving comparable sample quality to leading diffusion models, while using onl…
Evaluating fairness in ChatGPT
We’ve analyzed how ChatGPT responds to users based on their name, using AI research assistants to protect privacy.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Learning to reason with LLMs
[Read on openai.com](https://openai.com/index/learning-to-reason-with-llms)
OpenAI o1-mini
Advancing cost-efficient reasoning
Introducing SWE-bench Verified
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
Improving Model Safety Behavior with Rule-Based Rewards
We’ve developed and applied a new method leveraging Rule-Based Rewards (RBRs) that aligns models to behave safely without extensive human data collection.
GPT-4o mini: advancing cost-efficient intelligence
[Read on openai.com](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence)
Prover-Verifier Games improve legibility of language model outputs
Discover how prover-verifier games improve the legibility of language model outputs, making AI solutions clearer, easier to verify, and more trustworthy for …
OpenAI and Los Alamos National Laboratory announce research partnership
OpenAI and Los Alamos National Laboratory are working to develop safety evaluations to assess and measure biological capabilities and risks associated with f…
A Holistic Approach to Undesired Content Detection in the Real World
We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation.
Consistency Models
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes s…
Improved Techniques for Training Consistency Models
Consistency models are a nascent family of generative models that can sample high quality data in one step without the need for adversarial training.
Extracting Concepts from GPT-4
Using new techniques for scaling sparse autoencoders, we automatically identified 16 million patterns in GPT-4's computations.
Hello GPT-4o
We’re announcing GPT-4 Omni, our new flagship model which can reason across audio, vision, and text in real time.
Understanding the source of what we see and hear online
Today we’re introducing new technology to help researchers identify content created by our tools and joining the Coalition for Content Provenance and Authent…
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with the…
Video generation models as world simulators
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of …
Building an early warning system for LLM-aided biological threat creation
We’re developing a blueprint for evaluating the risk that a large language model (LLM) could aid someone in creating a biological threat. In an evaluation in…
Improving mathematical reasoning with process supervision
We’ve trained a model to achieve a new state-of-the-art in mathematical problem solving by rewarding each correct step of reasoning (“process supervision”) i…
Democratic inputs to AI
Our nonprofit organization, OpenAI, Inc., is launching a program to award ten $100,000 grants to fund experiments in setting up a democratic process for deci…
GPTs are GPTs: An early look at the labor market impact potential of large language models
[Read on openai.com](https://openai.com/index/gpts-are-gpts)
GPT-4
We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs,…
Point-E: A system for generating 3D point clouds from complex prompts
[Read on openai.com](https://openai.com/index/point-e)
Scaling laws for reward model overoptimization
[Read on openai.com](https://openai.com/index/scaling-laws-for-reward-model-overoptimization)
Introducing Whisper
We’ve trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition.
Efficient training of language models to fill in the middle
[Read on openai.com](https://openai.com/index/efficient-training-of-language-models-to-fill-in-the-middle)
DALL·E 2 pre-training mitigations
In order to share the magic of DALL·E 2 with a broad audience, we needed to reduce the risks associated with powerful image generation models. To this end, w…
Learning to play Minecraft with Video PreTraining
We trained a neural network to play Minecraft by Video PreTraining (VPT) on a massive unlabeled video dataset of human Minecraft play, while using only a sma…
Evolution through large models
[Read on openai.com](https://openai.com/index/evolution-through-large-models)
Techniques for training large neural networks
Large neural networks are at the core of many recent advances in AI, but training them is a difficult engineering and research challenge which requires orche…
Teaching models to express their uncertainty in words
[Read on openai.com](https://openai.com/index/teaching-models-to-express-their-uncertainty-in-words)
Hierarchical text-conditional image generation with CLIP latents
[Read on openai.com](https://openai.com/index/hierarchical-text-conditional-image-generation-with-clip-latents)
A research agenda for assessing the economic impacts of code generation models
[Read on openai.com](https://openai.com/index/economic-impacts-research)
Solving (some) formal math olympiad problems
We built a neural theorem prover for Lean that learned to solve a variety of challenging high-school olympiad problems, including problems from the AMC12 and…
Text and code embeddings by contrastive pre-training
[Read on openai.com](https://openai.com/index/text-and-code-embeddings-by-contrastive-pre-training)
WebGPT: Improving the factual accuracy of language models through web browsing
We’ve fine-tuned GPT-3 to more accurately answer open-ended questions using a text-based web browser.
Solving math word problems
We’ve trained a system that solves grade school math problems with nearly twice the accuracy of a fine-tuned GPT-3 model. It solves about 90% as many problem…
TruthfulQA: Measuring how models mimic human falsehoods
[Read on openai.com](https://openai.com/index/truthfulqa)
Introducing Triton: Open-source GPU programming for neural networks
We’re releasing Triton 1.0, an open-source Python-like programming language which enables researchers with no CUDA experience to write highly efficient GPU c…
Evaluating large language models trained on code
[Read on openai.com](https://openai.com/index/evaluating-large-language-models-trained-on-code)
Multimodal neurons in artificial neural networks
We’ve discovered neurons in CLIP that respond to the same concept whether presented literally, symbolically, or conceptually. This may explain CLIP’s accurac…
Understanding the capabilities, limitations, and societal impact of large language models
[Read on openai.com](https://openai.com/index/understanding-the-capabilities-limitations-and-societal-impact-of-large-language-models)
Scaling Kubernetes to 7,500 nodes
We’ve scaled Kubernetes clusters to 7,500 nodes, producing a scalable infrastructure for large models like GPT-3, CLIP, and DALL·E, but also for rapid small-…
CLIP: Connecting text and images
We’re introducing a neural network called CLIP which efficiently learns visual concepts from natural language supervision. CLIP can be applied to any visual …
DALL·E: Creating images from text
We’ve trained a neural network called DALL·E that creates images from text captions for a wide range of concepts expressible in natural language.
Generative language modeling for automated theorem proving
[Read on openai.com](https://openai.com/index/generative-language-modeling-for-automated-theorem-proving)
Image GPT
We find that, just as a large transformer model trained on language can generate coherent text, the same exact model trained on pixel sequences can generate …
Language models are few-shot learners
[Read on openai.com](https://openai.com/index/language-models-are-few-shot-learners)
AI and efficiency
We’re releasing an analysis showing that since 2012 the amount of compute needed to train a neural net to the same performance on ImageNet classification has…
Jukebox
We’re introducing Jukebox, a neural net that generates music, including rudimentary singing, as raw audio in a variety of genres and artist styles. We’re rel…
Improving verifiability in AI development
We’ve contributed to a multi-stakeholder report by 58 co-authors at 30 organizations, including the Centre for the Future of Intelligence, Mila, Schwartz Rei…
OpenAI Microscope
We’re introducing OpenAI Microscope, a collection of visualizations of every significant layer and neuron of eight vision “model organisms” which are often s…
Scaling laws for neural language models
[Read on openai.com](https://openai.com/index/scaling-laws-for-neural-language-models)
Dota 2 with large scale deep reinforcement learning
[Read on openai.com](https://openai.com/index/dota-2-with-large-scale-deep-reinforcement-learning)
Deep double descent
We show that the double descent phenomenon occurs in CNNs, ResNets, and transformers: performance first improves, then gets worse, and then improves again wi…
Procgen Benchmark
We’re releasing Procgen Benchmark, 16 simple-to-use procedurally-generated environments which provide a direct measure of how quickly a reinforcement learnin…
GPT-2: 1.5B release
As the final model release of GPT-2’s staged release, we’re releasing the largest version (1.5B parameters) of GPT-2 along with code and model weights to fac…
Solving Rubik’s Cube with a robot hand
We’ve trained a pair of neural networks to solve the Rubik’s Cube with a human-like robot hand. The neural networks are trained entirely in simulation, using…
Emergent tool use from multi-agent interaction
We’ve observed agents discovering progressively more complex tool use while playing a simple game of hide-and-seek. Through training in our new simulated hid…
GPT-2: 6-month follow-up
We’re releasing the 774 million parameter GPT-2 language model after the release of our small 124M model in February, staged release of our medium 355M model…
MuseNet
We’ve created MuseNet, a deep neural network that can generate 4-minute musical compositions with 10 different instruments, and can combine styles from count…
Generative modeling with sparse transformers
We’ve developed the Sparse Transformer, a deep neural network which sets new records at predicting what comes next in a sequence—whether text, images, or sou…
OpenAI Five defeats Dota 2 world champions
OpenAI Five is the first AI to beat the world champions in an esports game, having won two back-to-back games versus the world champion Dota 2 team, OG, at F…
Implicit generation and generalization methods for energy-based models
We’ve made progress towards stable and scalable training of energy-based models (EBMs) resulting in better sample quality and generalization ability than exi…
Neural MMO: A massively multiagent game environment
We’re releasing a Neural MMO, a massively multiagent game environment for reinforcement learning agents. Our platform supports a large, variable number of ag…
Better language models and their implications
We’ve trained a large-scale unsupervised language model which generates coherent paragraphs of text, achieves state-of-the-art performance on many language m…
Computational limitations in robust classification and win-win results
[Read on openai.com](https://openai.com/index/computational-limitations-in-robust-classification-and-win-win-results)
How AI training scales
We’ve discovered that the gradient noise scale, a simple statistical metric, predicts the parallelizability of neural network training on a wide range of tas…
Quantifying generalization in reinforcement learning
We’re releasing CoinRun, a training environment which provides a metric for an agent’s ability to transfer its experience to novel situations and has already…
Spinning Up in Deep RL
We’re releasing Spinning Up in Deep RL, an educational resource designed to let anyone learn to become a skilled practitioner in deep reinforcement learning.…
Learning concepts with energy functions
We’ve developed an energy-based model that can quickly learn to identify and generate instances of concepts, such as near, above, between, closest, and furth…
Plan online, learn offline: Efficient learning and exploration via model-based control
[Read on openai.com](https://openai.com/index/plan-online-learn-offline)
Reinforcement learning with prediction-based rewards
We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments thro…
FFJORD: Free-form continuous dynamics for scalable reversible generative models
[Read on openai.com](https://openai.com/index/ffjord)
The International 2018: Results
OpenAI Five lost two games against top Dota 2 players at The International in Vancouver this week, maintaining a good chance of winning for the first 20–35 m…
Large-scale study of curiosity-driven learning
[Read on openai.com](https://openai.com/index/large-scale-study-of-curiosity-driven-learning)
OpenAI Five Benchmark: Results
Yesterday, OpenAI Five won a best-of-three against a team of 99.95th percentile Dota players: Blitz, Cap, Fogged, Merlini, and MoonMeander—four of whom have …
Learning dexterity
We’ve trained a human-like robot hand to manipulate physical objects with unprecedented dexterity.
Variational option discovery algorithms
[Read on openai.com](https://openai.com/index/variational-option-discovery-algorithms)
Glow: Better reversible generative models
We introduce Glow, a reversible generative model which uses invertible 1x1 convolutions. It extends previous work on reversible generative models and simplif…
Learning Montezuma’s Revenge from a single demonstration
We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published resul…
OpenAI Five
Our team of five neural networks, OpenAI Five, has started to defeat amateur human teams at Dota 2.
Retro Contest: Results
The first run of our Retro Contest—exploring the development of algorithms that can generalize from previous experience—is now complete.
Learning policy representations in multiagent systems
[Read on openai.com](https://openai.com/index/learning-policy-representations-in-multiagent-systems)
GamePad: A learning environment for theorem proving
[Read on openai.com](https://openai.com/index/gamepad)
Gym Retro
We’re releasing the full version of Gym Retro, a platform for reinforcement learning research on games. This brings our publicly-released game count from aro…
AI and compute
We’re releasing an analysis showing that since 2012, the amount of compute used in the largest AI training runs has been increasing exponentially with a 3.4-…
Evolved Policy Gradients
We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can …
Gotta Learn Fast: A new benchmark for generalization in RL
[Read on openai.com](https://openai.com/index/gotta-learn-fast)
Retro Contest
We’re launching a transfer learning contest that measures a reinforcement learning algorithm’s ability to generalize from previous experience.
Variance reduction for policy gradient with action-dependent factorized baselines
[Read on openai.com](https://openai.com/index/variance-reduction-for-policy-gradient-with-action-dependent-factorized-baselines)
Improving GANs using optimal transport
[Read on openai.com](https://openai.com/index/improving-gans-using-optimal-transport)
On first-order meta-learning algorithms
[Read on openai.com](https://openai.com/index/on-first-order-meta-learning-algorithms)
Reptile: A scalable meta-learning algorithm
We’ve developed a simple meta-learning algorithm called Reptile which works by repeatedly sampling a task, performing stochastic gradient descent on it, and …
Some considerations on learning to explore via meta-reinforcement learning
[Read on openai.com](https://openai.com/index/some-considerations-on-learning-to-explore-via-meta-reinforcement-learning)
Ingredients for robotics research
We’re releasing eight simulated robotics environments and a Baselines implementation of Hindsight Experience Replay, all developed for our research over the …
Multi-Goal Reinforcement Learning: Challenging robotics environments and request for research
[Read on openai.com](https://openai.com/index/multi-goal-reinforcement-learning)
Interpretable machine learning through teaching
We’ve designed a method that encourages AIs to teach each other with examples that also make sense to humans. Our approach automatically selects the most inf…
Discovering types for entity disambiguation
We’ve built a system for automatically figuring out which object is meant by a word by having a neural network decide if the word belongs to each of about 10…
Requests for Research 2.0
We’re releasing a new batch of seven unsolved problems which have come up in the course of our research at OpenAI.
Scaling Kubernetes to 2,500 nodes
[Read on openai.com](https://openai.com/index/scaling-kubernetes-to-2500-nodes)
Block-sparse GPU kernels
We’re releasing highly-optimized GPU kernels for an underexplored class of neural network architectures: networks with block-sparse weights. Depending on the…
Learning sparse neural networks through L₀ regularization
[Read on openai.com](https://openai.com/index/learning-sparse-neural-networks-through-l0-regularization)
Interpretable and pedagogical examples
[Read on openai.com](https://openai.com/index/interpretable-and-pedagogical-examples)
Learning a hierarchy
We’ve developed a hierarchical reinforcement learning algorithm that learns high-level actions useful for solving a range of tasks, allowing fast solving of …
Generalizing from simulation
Our latest robotics techniques allow robot controllers, trained entirely in simulation and deployed on physical robots, to react to unplanned changes in the …
Asymmetric actor critic for image-based robot learning
[Read on openai.com](https://openai.com/index/asymmetric-actor-critic-for-image-based-robot-learning)
Sim-to-real transfer of robotic control with dynamics randomization
[Read on openai.com](https://openai.com/index/sim-to-real-transfer-of-robotic-control-with-dynamics-randomization)
Domain randomization and generative models for robotic grasping
[Read on openai.com](https://openai.com/index/domain-randomization-and-generative-models-for-robotic-grasping)
Competitive self-play
We’ve found that self-play allows simulated AIs to discover physical skills like tackling, ducking, faking, kicking, catching, and diving for the ball, witho…
Meta-learning for wrestling
We show that for the task of simulated robot wrestling, a meta-learning agent can learn to quickly defeat a stronger non-meta-learning agent, and also show t…
Nonlinear computation in deep linear networks
[Read on openai.com](https://openai.com/index/nonlinear-computation-in-deep-linear-networks)
Learning to model other minds
We’re releasing an algorithm which accounts for the fact that other agents are learning too, and discovers self-interested yet collaborative strategies like …
Learning with opponent-learning awareness
[Read on openai.com](https://openai.com/index/learning-with-opponent-learning-awareness)
OpenAI Baselines: ACKTR & A2C
We’re releasing two new OpenAI Baselines implementations: ACKTR and A2C. A2C is a synchronous, deterministic variant of Asynchronous Advantage Actor Critic (…
More on Dota 2
Our Dota 2 result shows that self-play can catapult the performance of machine learning systems from far below human level to superhuman, given sufficient co…
Dota 2
We’ve created a bot which beats the world’s top professionals at 1v1 matches of Dota 2 under standard tournament rules. The bot learned the game from scratch…
Gathering human feedback
RL-Teacher is an open-source implementation of our interface to train AIs via occasional human feedback rather than hand-crafted reward functions. The underl…
Better exploration with parameter noise
We’ve found that adding adaptive noise to the parameters of reinforcement learning algorithms frequently boosts performance. This exploration method is simpl…
Proximal Policy Optimization
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-ar…
Robust adversarial inputs
We’ve created images that reliably fool neural network classifiers when viewed from varied scales and perspectives. This challenges a claim from last week th…
Hindsight Experience Replay
[Read on openai.com](https://openai.com/index/hindsight-experience-replay)
Teacher–student curriculum learning
[Read on openai.com](https://openai.com/index/teacher-student-curriculum-learning)
Faster physics in Python
We’re open-sourcing a high-performance Python library for robotic simulation using the MuJoCo engine, developed over our past year of robotics research.
Learning to cooperate, compete, and communicate
Multiagent environments where agents compete for resources are stepping stones on the path to AGI. Multiagent environments have two useful properties: first,…
UCB exploration via Q-ensembles
[Read on openai.com](https://openai.com/index/ucb-exploration-via-q-ensembles)
OpenAI Baselines: DQN
We’re open-sourcing OpenAI Baselines, our internal effort to reproduce reinforcement learning algorithms with performance on par with published results. We’l…
Robots that learn
We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.
Roboschool
We are releasing Roboschool: open-source software for robot simulation, integrated with OpenAI Gym.
Equivalence between policy gradients and soft Q-learning
[Read on openai.com](https://openai.com/index/equivalence-between-policy-gradients-and-soft-q-learning)
Stochastic Neural Networks for hierarchical reinforcement learning
[Read on openai.com](https://openai.com/index/stochastic-neural-networks-for-hierarchical-reinforcement-learning)
Unsupervised sentiment neuron
We’ve developed an unsupervised system which learns an excellent representation of sentiment, despite being trained only to predict the next character in the…
Spam detection in the physical world
We’ve created the world’s first Spam-detecting AI trained entirely in simulation and deployed on a physical robot.
Evolution strategies as a scalable alternative to reinforcement learning
We’ve discovered that evolution strategies (ES), an optimization technique that’s been known for decades, rivals the performance of standard reinforcement le…
One-shot imitation learning
[Read on openai.com](https://openai.com/index/one-shot-imitation-learning)
Learning to communicate
In this post we’ll outline new OpenAI research in which agents develop their own language.
Emergence of grounded compositional language in multi-agent populations
[Read on openai.com](https://openai.com/index/emergence-of-grounded-compositional-language-in-multi-agent-populations)
Prediction and control with temporal segment models
[Read on openai.com](https://openai.com/index/prediction-and-control-with-temporal-segment-models)
Third-person imitation learning
[Read on openai.com](https://openai.com/index/third-person-imitation-learning)
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
[Read on openai.com](https://openai.com/index/pixelcnn-plus-plus)
Universe
We’re releasing Universe, a software platform for measuring and training an AI’s general intelligence across the world’s supply of games, websites and other …
#Exploration: A study of count-based exploration for deep reinforcement learning
[Read on openai.com](https://openai.com/index/exploration)
On the quantitative analysis of decoder-based generative models
[Read on openai.com](https://openai.com/index/on-the-quantitative-analysis-of-decoder-based-generative-models)
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models
[Read on openai.com](https://openai.com/index/a-connection-between-generative-adversarial-networks-inverse-reinforcement-learning-and-energy-based-models)
RL²: Fast reinforcement learning via slow reinforcement learning
[Read on openai.com](https://openai.com/index/rl2)
Variational lossy autoencoder
[Read on openai.com](https://openai.com/index/variational-lossy-autoencoder)
Extensions and limitations of the neural GPU
[Read on openai.com](https://openai.com/index/extensions-and-limitations-of-the-neural-gpu)
Transfer from simulation to real world through learning deep inverse dynamics model
[Read on openai.com](https://openai.com/index/transfer-from-simulation-to-real-world-through-learning-deep-inverse-dynamics-model)
Infrastructure for deep learning
Deep learning is an empirical science, and the quality of a group’s infrastructure is a multiplier on progress. Fortunately, today’s open-source ecosystem ma…
Generative models
This post describes four projects that share a common theme of enhancing or using generative models, a branch of unsupervised learning techniques in machine …
OpenAI Gym Beta
We’re releasing the public beta of OpenAI Gym, a toolkit for developing and comparing reinforcement learning (RL) algorithms. It consists of a growing suite …
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
[Read on openai.com](https://openai.com/index/weight-normalization)