This repo implements PPO algorithm for 3D drone control with collision avoidance using JAX.
We describe the details of the implementation in separate files:
- Environment description - general RL env description, quadrotor dynamics.
- PPO algorithm - detailed description of used PPO algorithm.
- Implementation details - implementation details of the PPO algorithm.
- Results - results of the training.
Note
In implementation details, we describe particular choices which allow us to perform fast and efficient PPO training using JAX.
We use conda to create a new environment and install the necessary packages.
Use
conda env create -f environment.ymlfor a CPU jax installation, or
conda env create -f environment_cuda.ymlfor a GPU jax installation.
See:
train.ipynb- for training PPO.plot_metrics.ipynb- for plotting training metrics.run.ipynb- for infering trained policy from saved training state.
On our particular setup, trained agent achieves the target in around 70% of all cases.
For more details, see results.