The initial condition problem with complete history dependency in learning models for travel choices
C Angelo Guevara, Yue Tang, Song Gao
- Conference
- hEART 2017: 6th Symposium of the European Association for Research in Transportation (2017)
- Publication year
- 2017
Abstract
2 Learning-based models that capture travelers’ day-to-day learning processes in repeated travel choices could benefit 3 from ubiquitous sensors such as smartphones, which provide individual-level longitudinal data to help validate and 4 improve such models. However, the common problem of missing initial observations in longitudinal data collection 5 can lead to inconsistent estimates of perceived value of attributes in question, and thus inconsistent parameter esti6 mates. In this paper, the stated problem is addressed by treating the missing observations as latent variables. The 7 proposed method is implemented in practice as maximum simulated likelihood (MSL) correction with two sampling 8 methods in an instance-based learning model for travel choice, and the finite sample bias and efficiency of the esti9 mators are investigated. Monte Carlo experimentation based on synthetic data shows that both the MSL with random 10 sampling (MSLrs) and MSL with importance sampling (MSLis) are effective in correcting for the endogeneity prob11 lem in that the percent error and empirical coverage of the estimators are greatly improved after correction. Compared 12 to the MSLrs method, the MSLis method is superior in both effectiveness and computational efficiency. Furthermore, 13 MSLis passes a formal statistical test for the recovery of the population values up to a scale with a large number of 14 missing observations, while MSLrs systematically fails due to the curse of dimensionality. The impacts of sampling 15 size in MSLrs and number of high probability choice sequences in MSLis on the methods’ performances are inves16 tigated. The methods are applied to an experimental route-choice dataset to demonstrate their empirical application. 17 Hausman-McFadden tests show that the estimators after correction are statistically equal to the estimators of the full 18 dataset without missing observations, confirming that the proposed methods are practical and effective for addressing 19 the stated problem.
C. A. Guevara, Y. Tang and S. Gao 3
1 1. INTRODUCTION 2 Learning-based models for travel choice capture travelers’ learning process in repeated choices (e.g., Ben-Elia and 3 Shiftan, 2010; Lu et al., 2014; Tang and Gao, conditionally accepted). In a learning model, a traveler’ s perception 4 of an alternative’s attribute (e.g., travel time) evolves over time based on all her past experience with the alternative. 5 When forming the perception, each past experience with the alternative takes a weight in memory and the perception 6 is a weighted average of all past experience. The weighting scheme of past experience is specific to the learning model 7 in use. Compared to non-learning models where the perception of an alternative is static over time, estimation of a 8 learning model requires data of travelers’ complete past experience with the alternatives. Longitudinal data collection 9 in real life, however, inevitably starts midstream, and rarely includes subjects’ complete choice histories. Specialized 10 data collection targeted at newcomers (e.g., new employees or students) to a region might provide the needed data, 11 but such efforts are difficult to implement. In the case of incomplete data, the missing initial observations can lead to 12 biased estimate of the perceived value of the attribute in question, and thus inconsistent parameter estimates. Note that 13 the majority of empirical studies on learning models for travel choice are based on experimental data in a laboratory, 14 where subjects make choices from “day” and thus the stated problem does not exist. 15 An econometric model is said to suffer from endogeneity when the systematic part of the utility is correlated 16 with the error term. The variables that cause the correlation are called the endogenous variables. Endogeneity can 17 lead to inconsistent estimation of model parameters, since changes in the error term are misinterpreted as changes 18 of the endogenous variable. Endogeneity is common in discrete choice models (e.g., probit, logit, nested logit) as 19 the assumption that the explanatory variables are independent from the error term is often violated. Guevara (2010) 20 classifies endogeneity into three types based on their causes: (1) Omission of the variables that are correlated with some 21 observed variables; (2) Simultaneous determination of multiple variables; and (3) The propagation of measurement 22 errors in explanatory variables to the error term. Several correction methods have been developed to solve endogeneity 23 problems (e.g., Berry et al., 1995; Brownstone, 1991; Fernández-Antolín et al., 2016; Guevara, 2010; Guevara and 24 Polanco, 2016; Heckman, 1978; Schenker and Welsh, 1988) . The endogeneity problem this paper tackles can be 25 classified within the third group, a special case in which endogeneity arises because the researcher has an incorrect 26 measure of the attributes of the alternatives perceived by the decision makers. 27 Solving the initial observation problem for dynamic panel data discrete choice models is known to be a 28 difficult task. Most existing studies deal with first-order Markov process where the dependent variable is only lagged 29 once. The major focus of these studies is that the initial condition is not exogenous due to correlation of error terms 30 over time. Therefore, if there is no serial correlation, first-order Markov process model would not suffer from the 31 problem. For example, Heckman (1981) and Lee (1997) examined the problem of initial conditions in a time-discrete 32 data stochastic process when serially correlated unobservable variables generate the process. Correction methods 33 were proposed and tested with Monte Carlo experiments. More of such studies can be found in the reference list (e.g. 34 Blundell and Bond, 1998; Carro, 2007; Honore and Kyriazidou, 2000; Wooldridge, 2005). In the learning models 35 for travel choice, a current decision depends on the entire history of past experience, defined as a Polya process in 36 Heckman (1981a). The complete history dependence makes the initial condition problem more challenging than those 37 in the existing studies. The model will suffer from the initial observation problem even without serial correlation. To 38 the best of our knowledge, no solution has been developed to date. 39 In this paper, the proposed method is based on noting that the likelihood function of this problem can be 40 written as a sequence of integrals over the conditional distribution of the possible choices on the missing days. This 41 multifold integral is then maximized using a variation of the maximum simulated likelihood (MSL), which is described 42 in detail by Train (2009). The MSL numerical estimation method has reached great popularity in the past 15 years, 43 thanks to the significant improvement in computational power. This method has been mainly used for the estimation 44 of Logit Mixture models aimed to account for random coefficients or different error component. The application of 45 the method in this paper is different from the usual ones, although all the conditions for consistency described in Train 46 (2009) are extendable, e.g., the need for having the number of draws growing faster than the square root of the sample 47 size. Despite its popularity, the MSL is not exempt from drawbacks. For example, MSL estimators have a downward 48 bias for a finite number of draws, and they may suffer from empirical identification problems, both in the form of false 49 empirical identification and lack of empirical identification. More importantly for this application, MSL may suffer 50 from the problem known as the curse of dimensionality, which in this case implies that the number of draws required 51 for estimation grows exponentially with the number of missing days, quickly making estimation impractical. This 52 problem is shared by all estimation methods based on simulation. This issue will be illustrated and investigated with 53 Monte Carlo experimentation. 54 Two sampling methods are proposed for the correction. The MSL random sampling (MSLrs) method ran-
C. A. Guevara, Y. Tang and S. Gao 4
1 domly draws a set of missing choice sequences following the learning model and a simple average of the simulated 2 choice probabilities is used in the simulated likelihood. This sampling method is expected to suffer from the curse 3 of dimensionality as the number of missing days grows. To overcome this limitation, the MSL importance sampling 4 (MSLis) method is proposed. It can be seen as a variation of the kernel conditional density nonparametric estimator 5 proposed by Rosenblatt (1969) and enhanced by Hyndman et al. (1996). In this case, instead of randomly simulating a 6 large enough number of missing choice sequences to evaluate the Logit Kernel function, a small number of sequences 7 with high probability of occurrence are sampled and the kernel, conditioning on the said probability are evaluated. 8 The main contribution of this paper is that a practical and theoretically sound correction method is developed 9 and assessed to address the endogeneity problem due to missing initial observations in learning models with complete 10 history dependency. To the best of our knowledge, the stated problem is tackled for the first time. Two sampling 11 methods are proposed for the correction method, with the aim of avoiding the problem of the curse of dimensionality 12 that arises as the number of missing days grows. The sample bias and effectiveness of the proposed method is inves13 tigated using a learning model proposed in recent literature (Tang and Gao, conditionally accepted) where perceived 14 attribute values are non-linear functions of a memory decay parameter. The suitability of the proposed method is con15 firmed using Monte Carlo experimentation on synthetic data, and its applicability is demonstrated using a laboratory 16 experimental dataset. 17 The remainder of the paper is organized as follows. The next section introduces the instance-based learning 18 (IBL) model for travel choice and presents the endogeneity problem due to missing initial observations. The MSL 19 method and two sampling methods are then proposed. The effectiveness and applicability of MSL with the two 20 sampling methods are demonstrated using both synthetic data and empirical data. Lastly, conclusions are presented 21 and future research directions are discussed.
22 2. AN INSTANCE-BASED LEARNING MODEL FOR TRAVEL CHOICE 23 The IBL model developed by Tang and Gao (conditionally accepted) is utilized to investigate the finite sample bias 24 and effectiveness of the proposed methods in correcting the endogeneity problem due to missing initial observations, 25 since: (1) The model is developed based on mainstream psychological findings of the power law of forgetting and 26 reinforcement and is shown to be able to capture various psychological effects that reside in travelers’ repeated choice 27 behaviors (Anderson and Schooler, 1991; Gonzalez et al., 2003; Newell and Rosenbloom, 1981; Rubin and Wenzel, 28 1996; Wickelgren, 1976). (2) Learning in the IBL model resides in the nonlinear memory decay parameter and is 29 based on complete history. The complexity of the model presents challenges to the effectiveness and efficiency of 30 the proposed method. For illustrative purpose, in this study the IBL model is introduced within a repeated binary 31 route-choice context. 32 A traveler n chooses one alternative from a choice set with two alternatives on each day t from day 1 to 33 K. Each alternative has an underlying random travel time whose realizations are independent from day to day, and 34 independent across alternatives. The traveler experiences the realized travel times of the chosen alternative on a given 35 day, and has no knowledge of the realized travel time on un-chosen alternatives. An instance is defined as a past 36 experience of a chosen alternative i on day t0 and its associated outcome (realized travel time), xi (t0 ). The realized 37 travel time is not indexed by traveler n, since it is sampled from the nature’s process and does not differ depending 38 on who is experiencing it. The index of a past day t0 ranges from 0 to t − 1, where t0 = 0 is a special time index 39 of the traveler’s initial perception of the alternative prior to her first experience. The traveler’s initial perception is 40 unobserved and reasonable assumptions can be made to represent its value, e.g., free flow travel time or a personal 41 trip planner’s information (such as Google Maps). An instance is stored in the declarative memory of the traveler, and 42 its activation decays over time following a power law. Specifically, on day t, its activation is (t − t0 )−d , where the 43 decay parameter d captures the rate of forgetting in that a smaller d value translates into higher activation in memory 44 and t − t0 measures the recency of the experienced travel times (smaller t − t0 values represent higher recency). The 45 weight of an instance in the traveler’ s memory is directly related to its activation. 46 Eq.(1) shows the weight of the experience from a past day t0 for traveler n, where the denominator is the 47 summation of activations over all past experiences on alternative i. The binary indicator ani (t) indicates whether 48 traveler n chose alternative i on day t. By definition ani (0) = 1, ∀i. The weight function shows that recency and 49 frequency jointly define the weight, i.e. more recent and frequent experienced travel times are more active in memory.
ani (t0 )(t − t0 )−d wni (t0 , t) = Pt−1 (1) τ =0 ani (τ )(t − τ ) −d
C. A. Guevara, Y. Tang and S. Gao 5
where
t : index of the current day, t = 1, . . . , K t0 : index of a previous day, t0 = 0, . . . , t − 1 wni (t0 , t) : weight of the experienced travel time on day t0 for the perceived travel time on day t for alternative i, traveler n d : decay parameter that captures the rate of forgetting, d > 0 ani (t0 ) : a binary indicator. It is 1 if traveler n chose alternative i on day t0 and 0 otherwise
On day t, the perceived travel time of alternative i is the weighted average of experienced travel times of all past days when alternative i is experienced, shown in Eq. (2). It depends on the entire choice history {ani (1), . . . , ani (t− 1)}. t−1 X bni (t) = wni (t0 , t)xi (t0 ) (2) t0 =0
where
bni (t) : perceived travel time of alternative i on day t for traveler n xi (t0 ) : realized travel time of alternative i on day t0 , t0 = 0, . . . , t − 1
Eq. (3) shows the utility function with the parameter vector φ = {d, βtime , α}, where the random residual ε is assumed to be i.i.d. Gumbel distributed. The systematic utility is linear in the perceived travel time bni (t) that varies from day to day and other attributes zi of the alternative that are constant over time (e.g., bus fare and number of traffic lights). It is straightforward to extend the utility function to include other attributes that vary from day to day, such as perceived fuel consumption or perceived crowdedness of public transit. Eq. (4) and Eq. (5) specify the choice probability of choosing path 1 and log-likelihood of observing all travelers’ choice sequences from day 1 respectively. The binary network assumption can be generalized to a larger number of alternatives, where issues such as overlapping alternatives and choice set generation need to be addressed.
Uni (t; φ) = Vni (t; φ) + εni (t) = βtime bni (t) + α0 z i + εni (t) (3)
where
Uni (t) : random utility of alternative i for traveler n on day t Vni (t) : systematic utility of alternative i for traveler n on day t βtime : coefficient to perceived travel time zi : explanatory variables for alternative i and traveler n that do not vary from day to day α : a vector of coefficients for attributes zi φ : parameter vector, φ = {d, βtime , α} εni (t) : random residuals that are i.i.d. Gumbel distributed at location 0 and scale 1
eVn1 (t) Pn (1|t, {1, 2}) = (4) eVn1 (t) + eVn2 (t)
where
Pn (1|t, {1, 2}) : choice probability of path 1 for traveler n on day t
K N X X 1−an1 (t) `N 1 = log Pn (1|t, {1, 2})an1 (t) 1 − Pn (1|t, {1, 2}) (5) n=1 t=1
How to cite
C Angelo Guevara; Yue Tang; Song Gao (2017). The initial condition problem with complete history dependency in learning models for travel choices. In: hEART 2017: 6th Symposium of the European Association for Research in Transportation.