hEART 2019 conference papers

Learning online combinatorial stochastic policies with deep reinforcement

Teo Stocco, Alexandre Alahi

Conference
hEART 2019: 8th Symposium of the European Association for Research in Transportation (2019)
Publication year
2019

Abstract

We study on-demand delivery services in stochastic environments and propose one representation of the online vehicle routing problem that can be learned with supervised and reinforcement learning. Using a simulation framework with topology projections, we evaluate different models and show that stochastic policies can be learned. Fine-tuning the models through the use of the REINFORCE rule and with a deterministic critic suggests that this approach can lead to promising policies to solve the problem.

How to cite

Teo Stocco; Alexandre Alahi (2019). Learning online combinatorial stochastic policies with deep reinforcement. In: hEART 2019: 8th Symposium of the European Association for Research in Transportation.