Repository files navigation
ExpectiMinimax Optimal strategy in games with chance nodes , MelkΓ³ E., Nagy B. (2007).
Sparse sampling A sparse sampling algorithm for near-optimal planning in large Markov decision processes , Kearns M. et al. (2002).
MCTS Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search , RΓ©mi Coulom, SequeL (2006).
UCT Bandit based Monte-Carlo Planning , Kocsis L., SzepesvΓ‘ri C. (2006).
Bandit Algorithms for Tree Search , Coquelin P-A., Munos R. (2007).
OPD Optimistic Planning for Deterministic Systems , Hren J., Munos R. (2008).
OLOP Open Loop Optimistic Planning , Bubeck S., Munos R. (2010).
OPSS Optimistic planning for sparsely stochastic systems , L. BuΕoniu, R. Munos, B. De Schutter, and R. Babuska (2011).
LGP Logic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning , Toussaint M. (2015). ποΈ
AlphaGo Mastering the game of Go with deep neural networks and tree search , Silver D. et al. (2016).
AlphaGo Zero Mastering the game of Go without human knowledge , Silver D. et al. (2017).
AlphaZero Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm , Silver D. et al. (2017).
TrailBlazer Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning , Grill J. B., Valko M., Munos R. (2017).
MCTSnets Learning to search with MCTSnets , Guez A. et al. (2018).
ADI Solving the Rubik's Cube Without Human Knowledge , McAleer S. et al. (2018).
OPC/SOPC Continuous-action planning for discounted inο¬nite-horizon nonlinear optimal control with Lipschitz values , Busoniu L., Pall E., Munos R. (2018).
Minimax analysis of stochastic problems , Shapiro A., Kleywegt A. (2002).
Robust DP Robust Dynamic Programming , Iyengar G. (2005).
Robust Planning and Optimization , Laumanns M. (2011). (lecture notes)
Robust Markov Decision Processes , Wiesemann W., Kuhn D., Rustem B. (2012).
Safe and Robust Learning Control with Gaussian Processes , Berkenkamp F., Schoellig A. (2015). ποΈ
Coarse-Id On the Sample Complexity of the Linear Quadratic Regulator , Dean S., Mania H., Matni N., Recht B., Tu S. (2017).
Tube-MPPI Robust Sampling Based Model Predictive Control with Sparse Objective Information , Williams G. et al. (2018). ποΈ
ICS Will the Driver Seat Ever Be Empty? , Fraichard T. (2014).
SafeOPT Safe Controller Optimization for Quadrotors with Gaussian Processes , Berkenkamp F., Schoellig A., Krause A. (2015). ποΈ
SafeMDP Safe Exploration in Finite Markov Decision Processes with Gaussian Processes , Turchetta M., Berkenkamp F., Krause A. (2016).
RSS On a Formal Model of Safe and Scalable Self-driving Cars , Shalev-Shwartz S. et al. (2017).
HJI-reachability Safe learning for control: Combining disturbance estimation, reachability analysis and reinforcement learning with systematic exploration , Heidenreich C. (2017).
CPO Constrained Policy Optimization , Achiam J., Held D., Tamar A., Abbeel P. (2017).
RCPO Reward Constrained Policy Optimization , Tessler C., Mankowitz D., Mannor S. (2018).
BFTQ A Fitted-Q Algorithm for Budgeted MDPs , Carrara N. et al. (2018).
MPC-HJI On Infusing Reachability-Based Safety Assurance within Probabilistic Planning Frameworks for Human-Robot Vehicle Interactions , Leung K. et al. (2018).
LTL-RL Reinforcement Learning with Probabilistic Guarantees for Autonomous Driving , Bouton M. et al. (2019).
Safe Reinforcement Learning with Scene Decomposition for Navigating Complex Urban Environments , Bouton M. et al. (2019).
Batch Policy Learning under Constraints , Le H., Voloshin C., Yue Y. (2019).
Uncertain Dynamical Systems
TS On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples , Thompson W. (1933).
UCB1 / UCB2 Finite-time Analysis of the Multiarmed Bandit Problem , Auer P., Cesa-Bianchi N., Fischer P. (2002).
Empirical Bernstein / UCB-V Exploration-exploitation tradeoff using variance estimates in multi-armed bandits , Audibert J-Y, Munos R., Szepesvari C. (2009).
Empirical Bernstein Bounds and Sample Variance Penalization , Maurer A., Ponti M. (2009).
An Empirical Evaluation of Thompson Sampling , Chapelle O., Li L. (2011).
kl-UCB The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond , Garivier A., CappΓ© O. (2011).
KL-UCB Kullback-Leibler Upper Confidence Bounds for Optimal Sequential Allocation , CappΓ© O. et al. (2013).
LUCB PAC Subset Selection in Stochastic Multi-armed Bandits , Kalyanakrishnan S. et al. (2012).
Track-and-Stop Optimal Best Arm Identification with Fixed Confidence , Garivier A., Kaufmann E. (2016).
M-LUCB / M-Racing Maximin Action Identification: A New Bandit Framework for Games , Garivier A., Kaufmann E., Koolen W. (2016).
LUCB-micro Structured Best Arm Identification with Fixed Confidence , Huang R. et al. (2017).
Black-box Optimization β¬
Reinforcement Learning π€
Dyna Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming , Sutton R. (1990).
UCRL2 Near-optimal Regret Bounds for Reinforcement Learning , Jaksch T. (2010).
PILCO PILCO: A Model-Based and Data-Efficient Approach to Policy Search , Deisenroth M., Rasmussen C. (2011). (talk )
DBN Probabilistic MDP-behavior planning for cars , Brechtel S. et al. (2011).
GPS End-to-End Training of Deep Visuomotor Policies , Levine S. et al. (2015). ποΈ
DeepMPC DeepMPC: Learning Deep Latent Features for Model Predictive Control , Lenz I. et al. (2015). ποΈ
SVG Learning Continuous Control Policies by Stochastic Value Gradients , Heess N. et al. (2015). ποΈ
Optimal control with learned local models: Application to dexterous manipulation , Kumar V. et al. (2016). ποΈ
BPTT Long-term Planning by Short-term Prediction , Shalev-Shwartz S. et al. (2016). ποΈ 1 | 2
Deep visual foresight for planning robot motion , Finn C., Levine S. (2016). ποΈ
VIN Value Iteration Networks , Tamar A. et al (2016). ποΈ
VPN Value Prediction Network , Oh J. et al. (2017).
An LSTM Network for Highway Trajectory Prediction , AltchΓ© F., de La Fortelle A. (2017).
DistGBP Model-Based Planning with Discrete and Continuous Actions , Henaff M. et al. (2017). ποΈ 1 | 2
Prediction and Control with Temporal Segment Models , Mishra N. et al. (2017).
Predictron The Predictron: End-To-End Learning and Planning , Silver D. et al. (2017). ποΈ
MPPI Information Theoretic MPC for Model-Based Reinforcement Learning , Williams G. et al. (2017). ποΈ
Learning Real-World Robot Policies by Dreaming , Piergiovanni A. et al. (2018).
Coupled Longitudinal and Lateral Control of a Vehicle using Deep Learning , Devineau G., Polack P., AlchtΓ© F., Moutarde F. (2018) ποΈ
PlaNet Learning Latent Dynamics for Planning from Pixels , Hafner et al. (2018). ποΈ
Hierarchy and Temporal Abstraction π
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning , Sutton R. et al. (1999).
Intrinsically motivated learning of hierarchical collections of skills , Barto A. et al. (2004).
OC The Option-Critic Architecture , Bacon P-L., Harb J., Precup D. (2016).
Learning and Transfer of Modulated Locomotor Controllers , Heess N. et al. (2016). ποΈ
Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving , Shalev-Shwartz S. et al. (2016).
FuNs FeUdal Networks for Hierarchical Reinforcement Learning , Vezhnevets A. et al. (2017).
Combining Neural Networks and Tree Search for Task and Motion Planning in Challenging Environments , Paxton C. et al. (2017). ποΈ
DeepLoco DeepLoco: Dynamic Locomotion Skills Using Hierarchical Deep Reinforcement Learning , Peng X. et al. (2017). ποΈ | ποΈ
Hierarchical Policy Design for Sample-Efficient Learning of Robot Table Tennis Through Self-Play , Mahjourian R. et al (2018). [ποΈ](https://sites.google.com/view/
DAC DAC: The Double Actor-Critic Architecture for Learning Options , Zhang S., Whiteson S. (2019).
Partial Observability ποΈ
PBVI Point-based Value Iteration: An anytime algorithm for POMDPs , Pineau J. et al. (2003).
cPBVI Point-Based Value Iteration for Continuous POMDPs , Porta J. et al. (2006).
POMCP Monte-Carlo Planning in Large POMDPs , Silver D., Veness J. (2010).
A POMDP Approach to Robot Motion Planning under Uncertainty , Du Y. et al. (2010).
Solving Continuous POMDPs: Value Iteration with Incremental Learning of an Efficient Space Representation , Brechtel S. et al. (2013).
Probabilistic Decision-Making under Uncertainty for Autonomous Driving using Continuous POMDPs , Brechtel S. et al. (2014).
MOMDP Intention-Aware Motion Planning , Bandyopadhyay T. et al. (2013).
The value of inferring the internal state of traffic participants for autonomous freeway driving , Sunberg Z. et al. (2017).
Belief State Planning for Autonomously Navigating Urban Intersections , Bouton M., Cosgun A., Kochenderfer M. (2017).
Probabilistic Decision-Making at Road Intersections: Formulation and Quantitative Evaluation , Barbier M., Laugier C., Simonin O., Ibanez J. (2018).
IT&E Robots that can adapt like animals , Cully A., Clune J., Tarapore D., Mouret J-B. (2014). ποΈ
MAML Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks , Finn C., Abbeel P., Levine S. (2017). ποΈ
Virtual to Real Reinforcement Learning for Autonomous Driving , Pan X. et al. (2017). ποΈ
Sim-to-Real: Learning Agile Locomotion For Quadruped Robots , Tan J. et al. (2018). ποΈ
ME-TRPO Model-Ensemble Trust-Region Policy Optimization , Kurutach T. et al. (2018). ποΈ
Kickstarting Deep Reinforcement Learning , Schmitt S. et al. (2018).
Learning Dexterous In-Hand Manipulation , OpenAI (2018). ποΈ
GrBAL / ReBAL Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning , Nagabandi A. et al. (2018). ποΈ
Learning agile and dynamic motor skills for legged robots , Hwangbo J. et al. (ETH Zurich / Intel ISL) (2019). ποΈ
Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning , Lee J., Hwangbo J., Hutter M. (ETH Zurich RSL) (2019)
IT&E Learning and adapting quadruped gaits with the "Intelligent Trial & Error" algorithm , Dalin E., Desreumaux P., Mouret J-B. (2019). ποΈ
Minimax-Q Markov games as a framework for multi-agent reinforcement learning , M. Littman (1994).
Autonomous Agents Modelling Other Agents: A Comprehensive Survey and Open Problems , Albrecht S., Stone P. (2017).
MILP Time-optimal coordination of mobile robots along specified paths , AltchΓ© F. et al. (2016). ποΈ
MIQP An Algorithm for Supervised Driving of Cooperative Semi-Autonomous Vehicles , AltchΓ© F. et al. (2017). ποΈ
SA-CADRL Socially Aware Motion Planning with Deep Reinforcement Learning , Chen Y. et al. (2017). ποΈ
Multipolicy decision-making for autonomous driving via changepoint-based behavior prediction: Theory and experiment , Galceran E. et al. (2017).
Online decision-making for scalable autonomous systems , Wray K. et al. (2017).
MAgent MAgent: A Many-Agent Reinforcement Learning Platform for Artificial Collective Intelligence , Zheng L. et al. (2017). ποΈ
Cooperative Motion Planning for Non-Holonomic Agents with Value Iteration Networks , Rehder E. et al. (2017).
MPPO Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement Learning , Long P. et al. (2017). ποΈ
COMA Counterfactual Multi-Agent Policy Gradients , Foerster J. et al. (2017).
FTW Human-level performance in first-person multiplayer games with population-based deep reinforcement learning , Jaderberg M. et al. (2018). ποΈ
Variable Resolution Discretization in Optimal Control , Munos R., Moore A. (2002). ποΈ
DeepDriving DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving , Chen C. et al. (2015). ποΈ
On the Sample Complexity of End-to-end Training vs. Semantic Abstraction Training , Shalev-Shwartz S. et al. (2016).
Learning sparse representations in reinforcement learning with sparse coding , Le L., Kumaraswamy M., White M. (2017).
World Models , Ha D., Schmidhuber J. (2018). ποΈ
Learning to Drive in a Day , Kendall A. et al. (2018). ποΈ
MERLIN Unsupervised Predictive Memory in a Goal-Directed Agent , Wayne G. et al. (2018). ποΈ 1 | 2 | 3 | 4 | 5 | 6
Variational End-to-End Navigation and Localization , Amini A. et al. (2018). ποΈ
Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks , Lee M. et al. (2018). ποΈ
Deep Neuroevolution of Recurrent and Discrete World Models , Risi S., Stanley K.O. (2019). ποΈ
Learning from Demonstrations π
QMDP-RCNN Reinforcement Learning via Recurrent Convolutional Neural Networks , Shankar T. et al. (2016). (talk )
DQfD Learning from Demonstrations for Real World Reinforcement Learning , Hester T. et al. (2017). ποΈ
Find Your Own Way: Weakly-Supervised Segmentation of Path Proposals for Urban Autonomy , Barnes D., Maddern W., Posner I. (2016). ποΈ
GAIL Generative Adversarial Imitation Learning , Ho J., Ermon S. (2016).
From perception to decision: A data-driven approach to end-to-end motion planning for autonomous ground robots , Pfeiffer M. et al. (2017). ποΈ
Branched End-to-end Driving via Conditional Imitation Learning , Codevilla F. et al. (2017). ποΈ | talk
UPN Universal Planning Networks , Srinivas A. et al. (2018). ποΈ
DeepMimic DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills , Peng X. B. et al. (2018). ποΈ
R2P2 Deep Imitative Models for Flexible Inference, Planning, and Control , Rhinehart N. et al. (2018). ποΈ
Applications to Autonomous Driving π
Inverse Reinforcement Learning
Projection Apprenticeship learning via inverse reinforcement learning , Abbeel P., Ng A. (2004).
MMP Maximum margin planning , Ratliff N. et al. (2006).
BIRL Bayesian inverse reinforcement learning , Ramachandran D., Amir E. (2007).
MEIRL Maximum Entropy Inverse Reinforcement Learning , Ziebart B. et al. (2008).
LEARCH Learning to search: Functional gradient techniques for imitation learning , Ratliff N., Siver D. Bagnell A. (2009).
CIOC Continuous Inverse Optimal Control with Locally Optimal Examples , Levine S., Koltun V. (2012). ποΈ
MEDIRL Maximum Entropy Deep Inverse Reinforcement Learning , Wulfmeier M. (2015).
GCL Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization , Finn C. et al. (2016). ποΈ
RIRL Repeated Inverse Reinforcement Learning , Amin K. et al. (2017).
Bridging the Gap Between Imitation Learning and Inverse Reinforcement Learning , Piot B. et al. (2017).
Applications to Autonomous Driving π
Apprenticeship Learning for Motion Planning, with Application to Parking Lot Navigation , Abbeel P. et al. (2008).
Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior , Ziebart B. et al. (2008).
Planning-based Prediction for Pedestrians , Ziebart B. et al. (2009). ποΈ
Learning for autonomous navigation , Bagnell A. et al. (2010).
Learning Autonomous Driving Styles and Maneuvers from Expert Demonstration , Silver D. et al. (2012).
Learning Driving Styles for Autonomous Vehicles from Demonstration , Kuderer M. et al. (2015).
Learning to Drive using Inverse Reinforcement Learning and Deep Q-Networks , Sharifzadeh S. et al. (2016).
Watch This: Scalable Cost-Function Learning for Path Planning in Urban Environments , Wulfmeier M. (2016). ποΈ
Planning for Autonomous Cars that Leverage Effects on Human Actions , Sadigh D. et al. (2016).
A Learning-Based Framework for Handling Dilemmas in Urban Automated Driving , Lee S., Seo S. (2017).
Learning Trajectory Prediction with Continuous Inverse Optimal Control via Langevin Sampling of Energy-Based Models , Xu Y. et al. (2019).
Motion Planning πββοΈ
Dijkstra A Note on Two Problems in Connexion with Graphs , Dijkstra E. W. (1959).
A* A Formal Basis for the Heuristic Determination of Minimum Cost Paths , Hart P. et al. (1968).
Planning Long Dynamically-Feasible Maneuvers For Autonomous Vehicles , Likhachev M., Ferguson D. (2008).
Optimal Trajectory Generation for Dynamic Street Scenarios in a Frenet Frame , Werling M., Kammel S. (2010). ποΈ
3D perception and planning for self-driving and cooperative automobiles , Stiller C., Ziegler J. (2012).
Motion Planning under Uncertainty for On-Road Autonomous Driving , Xu W. et al. (2014).
Monte Carlo Tree Search for Simulated Car Racing , Fischer J. et al. (2015). ποΈ
Architecture and applications
You canβt perform that action at this time.