You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
where $\hat{A}_t$ is the advantage estimate, and $\epsilon = 0.2$ is the clipping parameter.
We also use the value network to estimate the advantages. With $\phi$ being the vector of value network parameters, the value function loss is defined as:
where $\delta_t = r_t + \gamma V_{\phi}(s_{t+1}) - V_{\phi}(s_t)$ is the TD error. In our experiments, we used $\gamma = 0.99$ and $\lambda = 0.95$.
$T$ here is the length of the trajectory until truncation or termination. For more details on handling truncation and termination, see Implementation details.
Algorithm
General formulation for PPO is:
Here,
$N_e$ is the number of training epochs,
$B$ is the batch size,
$U$ is the unroll length,
$K$ is the number of optimizer steps per collected batch of trajectories,
$M$ is the minibatch size.
We empirically found that using $K = 2$ works better than other values. Increasing batch size and minibatch size while keeping number of minibatches small improves training stability. Unroll length is set to $U = 16$ (which is less than number of steps required to reach the target), and we found that using larger values did not improve performance.
Full algorithm listing:
References
Schulman, John, et al. "Proximal policy optimization algorithms." arXiv preprint arXiv:1707.06347 (2017). Link