{"id":"0fc6e02a-3a3f-4828-ac35-0c38bc6780bd","arxiv_id":"2506.18068","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":12,"one_line_summary":"Linking gaze and stress measurements to decision field theory's process parameters improves in-sample fit over standard logit models in two travel-choice datasets.","lead":"This paper adds eye-tracking, heart-rate and skin-conductance data to two kinds of travel-choice models, including decision field theory, and tests them on accommodation choices and driving gap-acceptance. It asks whether physiological measures of attention and stress improve predictions and behavioural insight beyond standard utility models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gaze endogeneity is the load-bearing risk: gaze shares are computed over the whole deliberation or the seconds before a gap appears, so the in-sample fit gains may reflect gaze predicting an already-formed choice rather than attention shaping it; a first-half-of-trial fixations test would settle…","rationale":"The reader's weakest assumption identifies the same load-bearing concern: gaze is treated as exogenous when it is plausibly partly a consequence of an emerging preference. I agree with this assessment. The paper explicitly acknowledges the ambiguity in Section 3.2, which makes the omission a real limitation rather than an unnoticed subtlety. The static case is the cleanest place to test it because gaze shares are aggregated over the whole trial; if gaze simply tracks the choice as it forms, the large in-sample gains and the behavioral alpha coefficients are not identified as attention effects. The dynamic case has the same issue in a less pure form, since the 5-second pre-gap window may include anticipatory scanning that is influenced by the upcoming decision. A split-window test using only early fixations would distinguish causal attention from predictive gaze. I also considered whether the absence of a heteroskedastic logit competitor undermines the stress-related claim, but the eye-tracking endogeneity is more central because it affects both the static and dynamic results and the overarching 'Beyond utility' interpretation. The nested likelihood-ratio tests and BIC comparisons are legitimate for in-sample fit, but they cannot settle the causal direction. The concern is addressable, so the conditional verdict stands; no change to the reader's recommendation is needed.","tokens_in":22474,"tokens_out":7857,"duration_ms":85453,"concrete_test":"Re-estimate the static models (MNL-E and DFT-E5/E6) using gaze shares computed only from fixations in the first half of each choice trial, holding all other specification choices fixed. If the log-likelihood gains over the base models shrink dramatically or the alpha_gaze coefficients change sign, the reported improvements are driven by gaze as an outcome of the already-forming choice rather than by attention shaping it. The same split-window logic can be applied to the dynamic task by using only the first 2.5 seconds of the 5-second pre-gap gaze window; if the improvement of DFT-E3 over DFT-S2 disappears, the attention-weight interpretation is not identified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim treats gaze fixations as exogenous information about attention and enters them directly into DFT attention weights or scaling parameters (Eq. 7; Sections 4.3.1 and 5.4.1). The paper itself acknowledges in Section 3.2 that a decision-maker may have 'already made their decision' while looking, but no specification accounts for this. In the static SP task, gaze shares are computed over the entire choice deliberation, so late fixations are partly a consequence of the preference that is forming; feeding them back into attention weights or attribute scaling creates a mechanical correlation between gaze and choice that inflates in-sample log-likelihood. In the dynamic gap-acceptance task, the gaze measures come from the 5 seconds before the gap appears, which may similarly include anticipation of rejection or acceptance. The reported gains (86-93 LL units in the static task; 6.08 LL units for DFT-E3 vs DFT-S2) cannot be interpreted causally as attention shaping choice; the behavioral coefficients, such as positive alpha_gaze_left in DFT-E3, may just reflect gaze predicting the eventual choice. All comparisons are in-sample, so BIC and likelihood-ratio tests do not resolve this ambiguity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for incorporating physiological process data (eye-tracking, heart rate, and skin conductance responses) into decision field theory (DFT) models of travel behaviour, and compares these against multinomial logit (MNL) benchmarks. Two empirical settings are studied: a static stated-preference accommodation choice task with eye-tracking, and a dynamic gap-acceptance task in a driving simulator with heart rate, skin conductance, and gaze measures. The authors estimate DFT variants in which gaze or stress indicators enter attention weights, attribute scaling parameters, initial preferences, or process noise, and they evaluate models using in-sample log-likelihood, BIC, and likelihood-ratio tests. The headline claim is that physiological data linked to DFT process parameters yields larger improvements in fit than adding the same data to utility functions in econometric models, with the static results showing larger DFT improvements for eye-tracking and the dynamic results showing that stress-linked process noise and gaze-linked attention weights improve DFT fit.","tokens_in":22887,"tokens_out":6786,"duration_ms":69968,"significance":"If the central claim held, this would be a valuable step toward using physiological sensor data to estimate cognitive process parameters in applied travel behaviour models, bridging two literatures that rarely meet. The paper is transparent in reporting model specifications, parameter tables, and likelihood-ratio test statistics across a large number of model variants, and it uses two genuinely different choice contexts. The authors also deserve credit for acknowledging in Section 5.5 that their process data are incorporated only in aggregated, static form. However, the load-bearing behavioural interpretation of the gaze coefficients is threatened by the endogeneity of gaze with respect to an emerging choice, and the headline comparison in the dynamic case is not fully supported by the reported fit statistics. The paper's contribution is therefore conditional on additional robustness analysis.","major_comments":[{"comment":"The eye-tracking covariates are aggregated over the entire deliberation in the static task and over the 5 seconds before each gap in the dynamic task, and are entered directly into attention weights or scaling parameters. The paper itself acknowledges in Section 3.2 that a decision-maker may have 'already made their decision' while looking, but no specification or robustness test addresses the resulting reverse causality. Late fixations are partly a consequence of the preference that is forming, so the 86-93 log-likelihood gains in Tables 2-3 and the significant gaze coefficients in Table 6 may be mechanical: gaze predicting an already-formed choice rather than attention shaping it. A test based only on first-half-of-deliberation fixations, or on fixations before some fixed threshold, would be needed to support the claim that gaze measures attention that shapes choice.","section":"Section 3.2, Eq. (7); Section 4.3.1, Eq. (13); Section 5.4.1, Eqs. (20)-(22)"},{"comment":"All model comparisons are in-sample. BIC and likelihood-ratio tests are computed on the same data used to motivate the specifications, and the gaze measures are high-dimensional aggregated summaries of the choice process. Neither BIC nor likelihood-ratio tests protect against overfitting when the process data are endogenous to the choice being modelled. Without holdout validation by task, individual, or scenario, the claims that physiological data add 'substantial' value and that DFT gives 'larger improvements' are overstated. I ask the authors to add out-of-sample or cross-validated log-likelihood comparisons, at least for the main competing models in Tables 5 and 6.","section":"Tables 1-6"},{"comment":"The abstract states that in the dynamic scenarios, linking stress and eye-tracking data to DFT process parameters results in 'larger improvements in comparison to simpler methods for incorporating this data in either DFT or econometric models.' The reported numbers do not support this for eye-tracking: MNL-E improves by 10.31 log-likelihood units over MNL-S (p=0.00013), whereas DFT-E3 improves by 6.08 over DFT-S2 (p=0.00686). Only for the stress data is the DFT improvement (10.34 for DFT-S2 over DFT-B) larger than the MNL improvement (2.98 for MNL-S over MNL-B). The comparative claim in the abstract and conclusions should be revised, or a more appropriate comparison should be provided that accounts for the different base models and parameter counts.","section":"Abstract and Table 6"},{"comment":"In DFT-E5, the gaze coefficient alpha_gazecount is estimated as 14.55 with a robust t-ratio of 1.65, i.e., not significant at conventional levels, yet the model is reported to improve by 93.08 log-likelihood units over the base model. This combination of a statistically insignificant gaze parameter with a very large fit gain is suspicious. It suggests that the fit improvement may be driven by changes in other parameters, notably sigma_epsilon = 125.56, rather than by the intended attention mechanism. The behavioural interpretation of alpha as the 'relative importance of gaze' is therefore not cleanly identified in this specification. Please report a profile likelihood over alpha or a comparison in which alpha is fixed at zero to clarify the source of the gain.","section":"Section 4.3.2, Table 3"}],"minor_comments":[{"comment":"The word 'accomodation' should be spelled 'accommodation' in the section heading.","section":"Section 4 heading"},{"comment":"The base model for MNL-E is labelled 'MNL-B' in the table header, but the 'Improvement over MNL-S/DFT-S2' row indicates that the comparison is against MNL-S. This is inconsistent and should be corrected.","section":"Table 6"},{"comment":"The footnote says 'to avoid overcomplicating Table 12', but no Table 12 exists; the reference should be to the table or appendix actually used.","section":"Section 4.3.2, footnote 3"},{"comment":"Equation (7) gives a generic function for the attention weight but no explicit functional form. Since the empirical work uses specific linear-in-gaze forms (e.g., Eq. (13) and Eq. (22)), please state the functional form used in estimation.","section":"Eq. (7)"},{"comment":"The parameter delta_bias is described as a bias that 'increases over the course of deliberating', but in the specification it enters the scaling vector for the constant 'alternative factors' attribute and is therefore active at every updating step. Please clarify how this implements a time-increasing bias.","section":"Section 5.2.1, Eq. (17)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a choice-modelling or transport-behaviour journal, and the dataset and model variations are presented with unusual transparency. The main risk is that the central 'process parameter' interpretation of gaze is not identified given the aggregated, end-of-deliberation gaze measures and the purely in-sample evaluation. I see no issue with the citation pattern beyond the expected reliance on the authors' own prior DFT work. The abstract and conclusions should be brought in line with the Table 6 numbers after the robustness analysis is added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first paper I know of that puts eye-tracking data inside a decision field theory (DFT) choice model and links skin-conductance and heart-rate data to DFT's process-noise parameter. That is a useful integration, and it is honestly executed. Second, the central empirical claim — process data helps DFT more than it helps logit — is supported only in-sample, and the gaze variables are treated as if they were exogenous snapshots of attention. That is the main soft spot, but it is not fatal for a first demonstration.\n\nWhat is actually new is the systematic comparison of where to put the sensor data: attention weights, scaling parameters, initial preference, or process noise, tested in both a static stated-preference task and a dynamic gap-acceptance task. The model comparison is careful. Likelihood-ratio tests, BIC, adjusted rho-squared, and robust t-ratios are reported clearly. Base DFT beats MNL in both contexts, consistent with prior DFT work. The static result that eye-tracking fits best in the scaling parameters, and the dynamic result that gaze belongs in attention weights while stress belongs in process noise, are genuinely informative for people building these models. The authors also explicitly note in Section 5.5 that their process data is aggregated rather than truly dynamic, which is a fair limitation to state.\n\nSoft spots, in proportion. All comparisons are in-sample; no holdout or cross-validation. That matters because the gaze variables are aggregated over the whole deliberation in the static task and over the five seconds before the gap appears in the dynamic task. The paper itself admits a decision-maker may have \"already made their decision\" while looking, but no specification allows for that. So part of the large log-likelihood gains may be gaze predicting the choice rather than attention shaping it. A first-half-of-trial gaze test would settle that. Some headline coefficients are only marginally significant, and the best specification flips between contexts. That is fine as a descriptive result — context dependence is claimed explicitly — but it undercuts any generic \"DFT is better with process data\" conclusion. The self-citations to Hancock et al. are appropriate here, not padding; the DFT machinery is genuinely theirs.\n\nWho this is for: choice modellers working with eye-tracking or physiological data, especially in transport and driving behaviour. The paper deserves a serious referee. I would not cite it this year as evidence of a causal attention mechanism, but I would cite it as the first integration of physiological data into DFT process parameters. Recommend engaging with it and asking for an out-of-sample or time-split robustness check before publication.","headline":"A credible first pass at wiring eye-tracking and stress data into DFT's process parameters; the empirical gains are real but in-sample, and the gaze endogeneity means the causal attention story still needs a cleaner test.","tokens_in":23315,"tokens_out":2158,"would_cite":true,"duration_ms":23207,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Physiological sensors improve cognitive choice models more than logit","keywords":["choice modelling","decision field theory","eye-tracking","physiological data","stress","gap acceptance","stated preference","process parameters"],"falsifier":"A decisive test would be out-of-sample prediction: estimate the models on one half of the choice tasks and compare predictive accuracy on the other half. If the DFT-with-process-data gains shrink or vanish relative to the logit models, the in-sample improvements come from gaze and stress tracing the choice rather than from a structural process link. A second test would manipulate gaze exogenously, for example by cueing attention, and see whether the alpha-driven probability shifts follow the manipulation.","tokens_in":22298,"feed_emoji":"👀","tokens_out":8444,"duration_ms":71420,"temperature":0.7,"pith_summary":"This paper tries to establish that physiological sensor data—eye fixations, heart rate, and skin conductance—can be meaningfully wired into the process parameters of decision field theory (DFT), a psychological model of how preferences accumulate over time. In a static stated-preference experiment on accommodation choice, adding eye-tracking data improved both standard logit and DFT models, with the best DFT specifications placing gaze information on attribute scaling parameters. In a dynamic driving-simulator gap-acceptance task, linking stress indicators to DFT's process noise gave a significantly better fit than adding them to initial preference, and linking gaze patterns to the attention weight on gap size outperformed simpler insertions in either framework. The paper argues that cognitive models with explicit process parameters are a more natural home for process data than utility-only econometric models, and that the gains are context-dependent.","feed_headline":"Eye-tracking and stress data lift cognitive models over logit","feed_subtitle":"Two travel experiments show gaze and stress measures work best when tied to decision field theory's process parameters.","key_machinery":"The central object is decision field theory (DFT) as implemented in the framework this paper builds on: a dynamic, stochastic accumulation model in which preference for each alternative evolves through a feedback matrix and a random valence vector driven by attribute attention. Its process parameters—attribute attention weights $w_k$, attribute scaling factors $\\beta$, sensitivity $\\phi_1$, memory $\\phi_2$, process noise $\\sigma_\\varepsilon$, and the number of preference-updating steps $\\tau$—are the slots into which physiological data are plugged. Eye-tracking enters either by adjusting the logit-transformed attention weights or by adjusting the scaling parameters, with an $\\alpha$ coefficient estimating the strength of the gaze link; stress enters either through the initial preference bias or by reparameterising process noise as $\\sigma_\\varepsilon = \\exp(\\alpha_{\\mathrm{stress}})$. The machinery works because DFT separates 'how often an attribute is considered' from 'how much it matters,' so gaze can inform the former and stress can inform the latter.","core_discovery":"On the paper's own terms, the discovery is that physiological measurements can be tied to the mechanisms of decision field theory rather than merely tacked onto a utility function, and that this yields measurable improvements in fit and interpretable behavioural parameters. In the static accommodation-choice task, the best model moved the eye-tracking effect onto the attribute scaling parameters, improving log-likelihood by 93 units over the base DFT model; in the dynamic gap-acceptance task, the best model moved gaze information onto the attention weight for gap size, and the stress data onto process noise. Significant positive alpha coefficients meant that more time looking at an attribute increased its influence, while higher measured stress increased choice unpredictability. The paper interprets this as evidence that process data can separate attention from preference within DFT, something choice data alone cannot do.","pith_inferences":["Because gaze is known to drift toward the eventual favourite, the large in-sample fits may partly be gaze predicting choice; a decisive out-of-sample or penultimate-fixation test would separate the structural attention effect from this reverse-causal path.","The same parameter-plugging strategy should transfer to other sequential sampling models, such as the attentional drift-diffusion or leaky competing accumulator, with gaze entering drift or boundary parameters; the static-versus-dynamic contrast suggests the best entry point will differ by model.","If stress genuinely inflates process noise, DFT predicts that choice consistency and response-time stability should co-vary across individuals; matching model-predicted preference trajectories to continuous simulator data would be a strong external check."],"forward_implications":["DFT models with eye-tracking can separately estimate attention and attribute importance, a decomposition that choice data alone cannot identify.","Stress-linked process noise implies stressed decision-makers do not simply shift their bias; they become less predictable, which is testable through response-time or steering variability.","The context-dependence of the best entry point—scaling parameters in the static task, attention weights in the dynamic task—means no single integration rule will work; applications must tailor where gaze enters.","If validated out of sample, the approach gives real-time driver-state prediction a structural home: gaze and heart-rate streams could feed DFT parameters for gap acceptance or lane-change models."],"supporting_citations":[{"why":"Supplies the multinomial logit baseline against which all physiological-data models are compared.","marker":"McFadden, 1974"},{"why":"Provides the original decision field theory whose preference-accumulation equations the paper extends.","marker":"Busemeyer and Townsend, 1992, 1993"},{"why":"Extends DFT to multiple alternatives and defines the contrast and feedback structure used here.","marker":"Roe et al., 2001"},{"why":"Offers the operationalised DFT framework with attention weights and scaling parameters that the paper reparameterises with physiological data.","marker":"Hancock et al., 2021"},{"why":"Adapts DFT for large-scale choice-model estimation, providing the identification and probability calculation approach.","marker":"Hancock et al., 2018"},{"why":"Provides the stated-preference eye-tracking dataset used for the static accommodation choice case.","marker":"Cohen et al., 2017"},{"why":"Supplies the driving-simulator dataset and stress-in-utility baseline that the dynamic gap-acceptance models extend.","marker":"Paschalidis et al., 2018"},{"why":"Reviews eye-tracking in discrete choice experiments, framing the gap this paper addresses.","marker":"Bansal et al., 2024"}],"fun_headline_variants":["Gaze and stress data boost decision field theory fits","Physiological signals sharpen cognitive choice models","Eye-tracking and stress metrics improve travel choice models","Attention and stress data refine travel behavior models","Process data outperforms utility in travel choice models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that gaze fixations and physiological stress readings can be treated as noisy but exogenous observations of the deliberation process, entered directly into attention weights or utility functions; if gaze is instead a consequence of a preference that has already formed, the model's alpha coefficients would be measuring gaze predicting choice, not attention shaping it.","fun_headline_variants_meta":{"raw":{"variants":["Gaze and stress data boost decision field theory fits","Physiological signals sharpen cognitive choice models","Eye-tracking and stress metrics improve travel choice models","Attention and stress data refine travel behavior models","Process data outperforms utility in travel choice models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1374,"prompt_tokens":993,"completion_tokens":381,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":311}},"tokens_in":609,"tokens_out":381,"duration_ms":4111,"temperature":1.0,"reasoning_tokens":311,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:55:24.795590+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be out-of-sample prediction: estimate the models on one half of the choice tasks and compare predictive accuracy on the other half. If the DFT-with-process-data gains shrink or vanish relative to the logit models, the in-sample improvements come from gaze and stress tracing the choice rather than from a structural process link. A second test would manipulate gaze exogenously, for example by cueing attention, and see whether the alpha-driven probability shifts follow the manipulation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the multinomial logit baseline against which all physiological-data models are compared."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the original decision field theory whose preference-accumulation equations the paper extends."},{"cited_title":"M., Busemeyer, J","cited_arxiv_id":null,"evidence_quote":"Extends DFT to multiple alternatives and defines the contrast and feedback structure used here."},{"cited_title":"O., Hess, S., Marley, A., and Choudhury, C","cited_arxiv_id":null,"evidence_quote":"Offers the operationalised DFT framework with attention weights and scaling parameters that the paper reparameterises with physiological data."},{"cited_title":"O., Hess, S., and Choudhury, C","cited_arxiv_id":null,"evidence_quote":"Adapts DFT for large-scale choice-model estimation, providing the identification and probability calculation approach."},{"cited_title":"L., Kang, N., and Leise, T","cited_arxiv_id":null,"evidence_quote":"Provides the stated-preference eye-tracking dataset used for the static accommodation choice case."},{"cited_title":"F., and Hess, S","cited_arxiv_id":null,"evidence_quote":"Supplies the driving-simulator dataset and stress-in-utility baseline that the dynamic gap-acceptance models extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reviews eye-tracking in discrete choice experiments, framing the gap this paper addresses."}],"review_version":1}