Adaptive fine-tuning of feedback perimeter controllers for multi-region urban networks with MFDs
Anastasios Kouvelas, Mohammadreza Saeedmanesh, Nikolas Geroliminis
- Conference
- hEART 2015: 4th Symposium of the European Association for Research in Transportation (2015)
- Publication year
- 2015
Abstract
Real-time traffic management is deemed to be an efficient and cost effective way to ameliorate traffic conditions and prevent gridlock phenomena in cities. An approach for real-time network-wide control for heterogeneous urban networks that has recently gain a lot of interest is the perimeter control (or gating). The basic concept of such an approach is to partition the heterogeneous network into a small number of homogeneous regions (i.e. areas with compact shape that have small variance of link densities) and apply perimeter control to the inter-transferring flows along the boundaries of each region. The key modelling tool that is used for the design of the control strategy is the Macroscopic Fundamental Diagram (MFD), which provides a concave, low-scatter relationship between network vehicle accumulations (veh) or density (veh/km) and network circulating flow (veh/h) if the region of a city is homogeneously congested (in terms of space). The concept of a network MFD was firstly introduced in Godfrey (1969), but the empirical verification of its existence with dynamic features is recent (Geroliminis and Daganzo, 2008). Since then, a vast literature on MFD modeling and control has been developed. Partitioning algorithms have been recently proposed in Ji and Geroliminis (2012) and Saeedmanesh and Geroliminis (2015). Perimeter MFD-based control policies have been introduced for multi-region heterogeneous networks (Geroliminis et al., 2013, Aboudolas and Geroliminis, 2013 and elsewhere) using different control methodologies. However, none of these works deal with parameter uncertainties in the model or short-term and long-term variations in the dynamics of the system. In this work a multivariable proportional integral (PI) feedback regulator is implemented to control the multi-region system. The structure of the PI controller is similar to the one used in Aboudolas and Geroliminis (2013), however the gain matrices and the targets (set-points) of the controller are updated in real-time based on performance measurements by an adaptive optimization algorithm. Especially when origin-destination tables are very asymmetric and there are strong directional flows to some regions of a city, an equal distribution of congestion might not be the optimal state of the system. Some regions should be penalized with more vehicles (and higher values of set points) so that regions with high attraction of destinations can operate at the critical value of accumulation that maximizes the regional outflow. This is a challenging estimation for complex situations with multiple pockets of congestion that should be performed through an optimization framework and not with simple heuristics or engineering principles. Moreover, in the proposed approach there are no control variables at the boundaries of the network but only at the borders between regions. As a consequence,
there are no vehicles kept outside of the network in order to protect the congestion of the regions and all the (gating) queues created by the controllers are internal to the network and thus affecting other movements. The overall control scheme (PI controller and adaptive optimization algorithm) is tested in microsimulation for the urban network of Barcelona, Spain. A detailed literature review will be provided in the full version of the paper. Consider an urban network partitioned in N homogeneous regions with well-defined MFDs. The index i € N = {1,2,...,N} denotes the region of the system and ni(k,) the total accumulation (number of vehicles) in region i at the discrete time k, (k; =0,1,2,...,K. —1). Let \j be the set of all regions that are directly reachable from the borders of region i, i.e. adjacent regions to region i. We assume that for each region i there exists a production MFD between accumulation nj;(k,) and total production P;(n;(k,)) (vehicle kilometers travelled per unit time), which describes the performance of the system in an aggregated way. This MFD can be easily estimated using measurements from loop detectors and/or GPS trajectories. The control variables ujj(k.), Vi € N’,j € M denote the fraction of the flow that is allowed to transfer from region i to region j at time k,. The values of the control variables are constrained by physical or operational constraints as follows
0< Ujmin < Uyj(km) < Uijmax <1, VieN,ji EM (1) The N-region MFDs system can be controlled in real-time by defining the values of the variables uij(ke), Vi € N,j € M. The control goal is to keep the traffic state of each region around a set value, so that the throughput is maximized and the region does not enter the saturated regime of the MFD. In this work, the following classical multivariable proportional-integral-type (PI) state feedback regulator is applied
u(k,) = u(k, — 1) — Kp [n(k,) — n(k, — 1)] — Ky [n(k,) — fi] (2) where u(k,) is the control vector of uj(k.), Vi € N,j € Ni, nk.) € RN the state vector of region accumulations n;(k,), Vi € NV, fi € RN the vector of the set points fi; for each region i and Kp, Ky € RMXN are the proportional and integral gains, respectively. The number of control variables M depends on the network partition and the sets Ni,i € NV. The values of the set points fi; can be defined by observing the production MFDs Pi(ni(k,)) for each region i. Note that a well-known property of the PI regulator is that it provides zero steady-state error (due to the existence of the integral term), i.e. n(k,) = fi under stationary conditions. The gain matrices Kp, K; as well as the vector with the set points fi are optimized in real-time by the use of AFT (Adaptive Fine-Tuning) algorithm. AFT is a recently developed iterative algorithm (see Kouvelas et al. (2011) for details) that is based on machine learning techniques and adaptive optimization principles and adjusts the control parameters to the variations of the process under control in order to optimize performance. The N-region MFDs system is controlled in real-time by the PI regulator (2) which includes a number of tunable parameters @ = vec (Kp, Kp, fi) € R20M*N)+N_ At the end of appropriately defined periods (e.g. at the end of each day), AFT algorithm receives the value of the real (measured) performance index J (e.g. total delay of the system), as well as the values of the most significant measurable external disturbances x (e.g. aggregated demand). Note that the scalar performance index J(0,x) is a (generally unknown) function of the external factors x and the tunable parameters 0. Using the measured quantities (the samples of which increase iteration by iteration), AFT calculates new tunable parameter values to be applied at the next period (e.g. the next day) in an attempt to improve the system performance. This (iterative) procedure is continued over many periods (e.g. days) until the algorithm converges and an optimal performance is reached; then,
AFT algorithm may remain active for continuous adaptation or can be switched off and re-activated at a later stage. The main component of the employed algorithm is a universal approximator J (0,x) (e.g., a polynomial-like approximator or a neural network) that is used in order to obtain an approximation of the nonlinear mapping J (0,x), based on all previous samples. At each algorithm iteration ko, the following main steps are taking place (the reader is referred to Kouvelas et al. (2011) for a complete description of the algorithm):
1. A new polynomial approximator with L, regressor terms is produced, which has the following structure J") (8,x) = 9T(ko) op) (8, x) (3) where L, = min {2 (k, — 1) , Lg max} (i.e. the number increases with iterations up to a maximum value), O(ko) € R'« are the weights of the approximator for iteration k, (equivalent to the synaptic connections in neural networks) and '**) (@,x) is a vector with L, sigmoidal functions of polynomials constructed using the elements of vectors 0, x (nonlinear activation functions or neurons).
2. The values of the weights 9(k,) are obtained by the solution of the following optimization problem ko (ko) = arg min5 (Ji —9(ko)"(*)” (4) 2m
3. Many randomly chosen candidate perturbations AG”) (ke), p €{1,2,...,K} are generated. The effect of all candidate new vectors 0'?)(k, +1) = 0*(k,) + A®'?)(k,) (where 0*(k,) is the “best” set of tunable parameters until the k,-th experiment, i.e. the one with the best performance so far), as well as e'*)(k, +1) = 0*(k,) — A@)(k,) to the system performance is estimated by using the approximator mentioned above, i.e.
F (0 (ko +1), R(ko + 1) = (ko) Tp"! (Chaos +1), &(ko + 0) (5) where &(k, +1) is an estimate of the external disturbances x for the next experiment k, +1.
4. The vector 0(k, + 1) that corresponds to the best estimate, i.e.
O(ko+1)=arg min F (OP (ky +1), &(ko + ) (6) 0!*P)(ko+1)
is selected to determine the new values for the tunable parameters 0(k, + 1) to be applied at the next period k, + 1 (e.g. the next day).
The efficiency of the adaptive system described above is tested in microsimulation experiments. The urban network of Barcelona, Spain is used as the test site, which is modeled and calibrated via the AIMSUN microscopic environment (Figure fifa). The duration of the simulation is 2 hours including a 15 minutes warm-up period. In the no control case (where the real fixed-time plans are applied to the intersections) the network faces some serious congestion problems, with queues spilling back to upstream intersections. The network is first partitioned into 4 homogeneous regions. The results of the partitioning algorithm are presented in Figure [i{b). This partitioning derives M = 6 control variables (ie. 144, U24, U34, Ug], W472, U43) and N = 4 state variables (nj, n2,n3, 14). The PI regulator is applied
Figure 1: The test site of Barcelona, Spain: (a) simulation model with four regions; (b) results of the partitioning algorithm and controlled intersections. Blue circles correspond to intersections belonging to uy4, red to uz4, green to u34 and black to uyy,h = 1, 2,3.
every T; = 90sec and the control decisions (after modified to satisfy the operational constraints) are forwarded for application to 28 signalized intersections which are all across the boundaries of region 4 (Figure [ifb)). AFT runs for many iterations starting from an initial point where Kp = K; = 0™N. For these values the regulator operates as a fixed-time policy and this point is equivalent to the no control (NC) case (ie. the actual fixed-time plans of the city are applied). The initial values for the set points fi are obtained from the MFDs of the NC case and are equal to fi = [1700,600, 600, 2400)". The performance index of AFT (i.e. the objective function J that tries to minimize) is selected to be the total delay of the system, which is available after the end of the simulation. In each iteration the whole simulation of 2 hours is run with the same parameters for the controller. At the end of the simulation AFT is called to calculate the new values of Kp, Kj, fi to be used in the next iteration. Figure [2{a) illustrates the control decisions u = [uy4, U24, Us4, U41, U42, U43]' for every k., for the simulation with the best results in terms of total delay (BC). The controller is activated at k, = 24 and stays active until the end of the simulation, as the accumulations are continuously increasing. Figure [2{b) presents the time series of the accumulations over the simulation time for all regions. At the last 30 minutes of the simulation the accumulations of regions 1, 3 and 4 for the BC case are quite lower than the NC case. Figure Bla) and (b) show the production MFDs of all regions for NC and BC respectively. For the BC case, the conditions in regions 1, 3 and 4 are improved as they only have a few states in the congested regime, whereas region 2 remains uncongested in both cases. The improvement of BC case versus NC is about 14% for the total delay and about 11% for the average speed of the network.
How to cite
Anastasios Kouvelas; Mohammadreza Saeedmanesh; Nikolas Geroliminis (2015). Adaptive fine-tuning of feedback perimeter controllers for multi-region urban networks with MFDs. In: hEART 2015: 4th Symposium of the European Association for Research in Transportation.