RG-CQL: A Reward-Guided Conservative Q-Learning Framework for the Coordination of Ride-Pooling and Public Transit Services
Yulong Hu, Sen Li
- Conference
- hEART 2025: 13th Symposium of the European Association for Research in Transportation (2025)
- Publication year
- 2025
Abstract
This paper presents the Reward-Guided Conservative Q-learning (RG-CQL), a novel Reinforcement Learning (RL) framework designed to optimize the coordination between ride-pooling and public transit in multimodal transportation systems. By modeling the problem as a Markov Decision Process (MDP), RG-CQL employs a two-phase approach: offline learning and online fine-tuning. In the offline phase, the Conservative Double Deep Q Network (CDDQN) as potential action executor and a Guider Network as rewards estimator are trained directly from past noisy data trajectories. During the online fine-tuning phase, the Guider Network assists the CDDQN in exploring new stateaction pairs, merging conservative training with optimistic online strategies. Extensive numerical experiments in Manhattan demonstrate that RG-CQL improves the operational performance of multi-modal transportation systems by 4.3%, reduces training sample complexity by 81.3%, and effectively addresses the Offline to Online (O2O) RL challenge in large-scale ride-pooling systems.
How to cite
Yulong Hu; Sen Li (2025). RG-CQL: A Reward-Guided Conservative Q-Learning Framework for the Coordination of Ride-Pooling and Public Transit Services. In: hEART 2025: 13th Symposium of the European Association for Research in Transportation.