Skip to content

Latest commit

 

History

History
98 lines (64 loc) · 3.01 KB

File metadata and controls

98 lines (64 loc) · 3.01 KB

Proximal Policy Optimization

In our implementation of PPO, we mainly follow the original paper [1].

Loss functions

We denote probability ratio as:

$$r_t(\theta) = \frac{\pi_{\theta}(a_t | s_t)}{\pi_{\theta_{\text{old}}}(a_t | s_t)},$$

so that $r_t(\theta_{\text{old}}) = 1$, and $\theta$ is the policy parameter vector.

Objective is defined as:

$$L^{\text{CLIP}}(\theta) = \hat{\mathbb{E}}_t \Big[ \min\bigl( r_t(\theta)\,\hat{A}_t,\ \mathrm{clip}\bigl(r_t(\theta), 1 - \epsilon, 1 + \epsilon\bigr)\,\hat{A}_t \bigr) \Big],$$

where $\hat{A}_t$ is the advantage estimate, and $\epsilon = 0.2$ is the clipping parameter.

We also use the value network to estimate the advantages. With $\phi$ being the vector of value network parameters, the value function loss is defined as:

$$L^{\text{VF}}(\phi) = \frac{1}{2} \Big[ V_{\phi}(s_t) - V_t^{\text{target}} \Big]^2,$$

where:

  • $V_{\phi}(s_t)$ is the value function estimate at time $t$,
  • $V_t^{\text{target}}$ is the target value function, defined as
$$V_t^{\text{target}} = \hat{A}_t + V_{\phi}(s_t).$$

We also use the entropy loss to encourage exploration, as we found it to be helpful for overall performance. The entropy loss is defined as:

$$L^{\text{entropy}}(\theta) = S[\pi_{\theta}](s_t),$$

where $S\pi_{\theta}$ is the entropy of the parametric action distribution at time $t$.

Total loss is:

$$L(\theta, \phi) = \hat{\mathbb{E}}_t \Big[ -L^{\text{CLIP}}(\theta) + c_1 L^{\text{VF}}(\phi) - c_2 S[\pi_{\theta}](s_t) \Big]$$

where $c_1 = 0.5$ and $c_2 = 0.01$.

Generalized Advantage Estimation

We use Generalized Advantage Estimation (GAE) to estimate the advantage function. The advantage estimate is defined as:

$$\hat{A}_t = \delta_t + (\gamma \lambda) \delta_{t+1} + \cdots + (\gamma \lambda)^{T-t} \delta_{T-1},$$

where $\delta_t = r_t + \gamma V_{\phi}(s_{t+1}) - V_{\phi}(s_t)$ is the TD error. In our experiments, we used $\gamma = 0.99$ and $\lambda = 0.95$.

$T$ here is the length of the trajectory until truncation or termination. For more details on handling truncation and termination, see Implementation details.

Algorithm

General formulation for PPO is:

General PPO

Here,

  • $N_e$ is the number of training epochs,
  • $B$ is the batch size,
  • $U$ is the unroll length,
  • $K$ is the number of optimizer steps per collected batch of trajectories,
  • $M$ is the minibatch size.

We empirically found that using $K = 2$ works better than other values. Increasing batch size and minibatch size while keeping number of minibatches small improves training stability. Unroll length is set to $U = 16$ (which is less than number of steps required to reach the target), and we found that using larger values did not improve performance.

Full algorithm listing:

PPO

References

  1. Schulman, John, et al. "Proximal policy optimization algorithms." arXiv preprint arXiv:1707.06347 (2017). Link