Paper Lineage
Esc
MethodJul 2017arXiv 1707.06347cs.LG

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski, Prafulla Dhariwal and 2 others

We propose a new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using stochastic gradient ascent. Whereas standard policy gradient methods perform one gradient update per data sample, we propose a novel objective function that enables multiple epochs of minibatch updates.

From the abstract

Built on

0 papers · 0 verifiedSee as graph

PPO has no earlier papers in this dataset.

Led to

  • MnasNet2018 · cited 1×, 1 in Method
    “At the end of each step, the parameters θ\theta of the controller are updated by maximizing the expected reward defined by equation 5 using Proximal Policy Optimization schulman2017proximal.”
    From MnasNet · §Mobile Neural Architecture Search
  • RLHF for LMs (Ziegler)2019 · cited 1×, 1 in Method
    “Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. 2017) with reward RR from eq. 2 on x∼𝒟x\simD.”
    From RLHF for LMs (Ziegler) · §Methods
  • “Finally, we train a policy via reinforcement learning (RL) to maximize the score given by the RM; the policy generates a token of text at each ‘time step’, and is updated using the PPO algorithm [58] based on the RM ‘rew…”
    From Learning to summarize from human feedbac · §unknown section
  • InstructGPT2022 · cited 3×, 2 in Method
    “Once again following Stiennon et al., 2020, we fine-tuned the SFT model on our environment using PPO (Schulman et al., 2017).”
    From InstructGPT · §Methods and experimental details
Abstract

We propose a new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using stochastic gradient ascent. Whereas standard policy gradient methods perform one gradient update per data sample, we propose a novel objective function that enables multiple epochs of minibatch updates. The new methods, which we call proximal policy optimization (PPO), have some of the benefits of trust region policy optimization (TRPO), but they are much simpler to implement, more general, and have better sample complexity (empirically). Our experiments test PPO on a collection of benchmark tasks, including simulated robotic locomotion and Atari game playing, and we show that PPO outperforms other online policy gradient methods, and overall strikes a favorable balance between sample complexity, simplicity, and wall-time.