REVIEW 3 major objections 5 minor 26 references
Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper derives a three-state Markov model of driver takeover behavior from a 206-person online study and reports environment and sex effects.
desk verdict The paper has a new 206-person dataset and honestly reports null effects, but the abstract claims a transparency effect the data contradict, so the central finding as advertised is unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a first-order Markov chain over three driver states — Takeover (T), Alert (A), and Normal (N) — with a Start state fixed at Normal. The mapping from data is given by thresholds on the participant's chosen level of autonomy: $\text{LoA} \le 0.5$ maps to T, $0.5 < \text{LoA} < 2.0$ maps to A, and $\text{LoA} \ge 2.0$ maps to N. Transition probabilities between successive scenes are the observed proportions of participants who moved from one state to another, and chains computed under different experimental conditions (environment, information level, sex) are compared with chi-square tests. This chain is the compact mathematical object that converts the study's multiple-choice answers into a predictive model of driver behavior.
What would settle it
A controlled driving-simulator study presenting the same highway and suburban scenarios with real vehicle motion and a genuine takeover task would settle the claim. Computing the same transition matrices with the same state mapping ($\text{LoA} \le 0.5$, $0.5 < \text{LoA} < 2.0$, $\text{LoA} \ge 2.0$), if the simulator data show no environment or sex differences — or a robust transparency effect that the online study missed — the model would fail as a predictor of real driving.
Extended reading notes
Core claim
The central discovery the authors are trying to establish is that a driver's willingness to delegate or retake control can be captured by a compact Markov chain whose transition probabilities are not universal. Using thresholds on the participant's chosen level of autonomy, responses fall into Takeover (LoA ≤ 0.5), Alert (0.5 < LoA < 2.0), or Normal (LoA ≥ 2.0), and transitions between scenes are estimated from the observed proportions. The authors report statistically significant differences between highway and suburban environments and between male and female participants, and they interpret the Takeover state as highly persistent. They also put forward transparency — the amount of information the vehicle displays — as one of the factors under investigation. The end goal is a model that can be embedded in a multi-agent situation-awareness architecture, letting the vehicle foresee driver state changes and adapt its behavior accordingly.
Load-bearing premise
The whole model rests on the assumption that what people say they would do while looking at staged images in an online survey is a good stand-in for what they would actually do behind the wheel.
Editorial extensions
If this is right
- If the reported environment and sex effects replicate, autonomous vehicles should not rely on a single generic driver model; transition expectations should be conditioned on road context and driver demographics.
- The high persistence of the Takeover state (75–100% probability of staying in it) implies that once a driver intervenes, they are unlikely to hand control back, so the AV should treat an intervention as the start of a sustained manual-driving phase.
- The better second-order fit for highway data implies that a first-order Markov chain is too simple for highway settings; predicting driver behavior there requires remembering more than the immediately previous state.
- The significant non-stationarity between scene pairs means transition probabilities drift as a scenario progresses, so practical use of the model would need time- or phase-dependent transition matrices rather than one global chain.
Reading between the lines
- The paper's own chi-square and homogeneity tests do not show a statistically significant effect of information transparency on transition probabilities at the conventional threshold, so the abstract's framing of transparency as a factor goes beyond what the reported analyses support.
- Because the data come from self-reported choices on static AI-generated vignettes rather than from a simulator or real vehicle, the transition probabilities are estimates of stated hypothetical behavior; a simulator-based replication is the natural next test of their predictive value.
- The thresholds that split the continuous level-of-autonomy response into T, A, and N are hand-chosen; using data-driven boundaries or keeping the continuous response could reveal finer-grained condition differences.
- The stickiness of the Takeover state suggests an untested design implication: automated vehicles should include explicit handover protocols and monitor driver state after any intervention, since drivers are unlikely to return control quickly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a large-scale online user study (N=206) on human interaction with an autonomous vehicle, manipulating the amount of information (transparency) provided by the AV interface and collecting participants' self-reported Level of Autonomy choices in two scenarios (highway and suburbs). From these choices, the authors discretize behavior into Takeover, Alert, and Normal states and estimate first-order Markov chain transition probabilities. They then use chi-square tests to compare transitions across environment, information level, and participant sex, and report additional stationarity, homogeneity, and goodness-of-fit tests. The abstract and introduction claim that significant differences in the model's transitions depend on the AV's transparency, the environment, and the users' demographics, while the conclusion describes the resulting models as 'statistically validated Markov models.'
Significance. If the modeling approach were sound, a data-driven Markov model of driver state transitions could be a valuable component for situational-awareness architectures and for studying trust calibration in human-AV interaction. The paper deserves credit for transparently describing its methodology, for reporting null results explicitly rather than hiding them, and for grounding the interface design in a prior participatory workshop. However, the central advertised finding—that transparency yields significant differences in transitions—is contradicted by the paper's own statistical tests, and the first-order Markov model is rejected for the highway scenario by the authors' own likelihood-ratio test. The contribution is therefore not established as stated; the paper would need either a corrected analysis, a second-order model, and a substantive reframing before the claims could be supported.
major comments (3)
- [Abstract; Section IV.C.2; Section IV.D.2] The abstract states that 'depending on the AV's transparency, the scenario's environment, and the users' demographics, we can obtain significant differences in the model's transitions,' but the reported chi-square tests show no significant information-level effect in either environment (highway, chi-square = 2.24, p = 0.326; suburbs, chi-square = 1.04, p = 0.593), and the homogeneity test for information-level subgroups is also non-significant (chi-square = 7.31, p = 0.120). Since the transparency manipulation is one of the three named factors in the paper's central claim, the paper's own statistics directly refute that component of the headline result. Table VII shows numerically larger Alert-to-Takeover transitions under low information, but this is an uncorrected comparison of selected transitions and cannot substitute for the overall non-significant chi-square test.
- [Section IV.D.3; Section V] The likelihood-ratio goodness-of-fit test in Section IV.D.3 rejects the first-order Markov model for the highway scenario (G2 = 14.29, p = 0.027), yet Section V describes the presented models as 'statistically validated Markov models.' Section IV.E.3 acknowledges the higher-order effect but the paper does not estimate, report, or validate a second-order model; therefore the central modeling contribution is not supported for one of the two scenarios and the conclusion overstates what the data show.
- [Section III.C; Section III.D; Section IV.B] The behavioral data consist of self-reported multiple-choice Level-of-Autonomy selections made in response to static AI-generated images and roleplay narration, and these choices are mapped to Takeover, Alert, and Normal states using hand-chosen thresholds (LoA <= 0.5, 0.5 < LoA < 2.0, LoA >= 2.0). All transition probabilities, chi-square tests, and model-validation results therefore rest on the assumption that such vignette choices are a valid proxy for real driving behavior, and on threshold choices that are not subjected to any sensitivity analysis. Without evidence for either the proxy validity or the robustness of the thresholds, the ecological validity of the resulting 'human driver behavior' model is not established.
minor comments (5)
- [Section III.C] One of the multiple-choice answers contains the typo 'approach the roa' in the Highway scenario; it should read 'approach the road.'
- [Section I, Section III.C] The text contains small typographical errors such as 'artifical' and 'T o'; these should be corrected in revision.
- [Section III.B] The paper says participants were 'equally distributed between male and female,' but the reported self-reported gender ratio is 104:101:1 (F:M:X); please clarify which variable was used for the sex-based analyses and why the counts differ from the legal-sex distribution.
- [Section III.D] The description of the trust score is unclear: three items coded as 1 or -1 cannot produce a range from -2 to 4 unless more than three items or a different coding scheme is used; please specify the exact coding and number of items.
- [References] Reference [18] is incomplete and should be corrected to a full citation of Norris's 'Markov Chains'.
Circularity Check
No circularity: the Markov-chain transition probabilities are empirical counts from the study data, and the only self-citations are non-load-bearing background.
full rationale
The paper's derivation chain is self-contained. Section IV.B defines three Markov states by fixed, pre-specified LoA thresholds (T: LoA <= 0.5; A: 0.5 < LoA < 2.0; N: LoA >= 2.0) and then computes transition probabilities as empirical counts from the collected data, e.g., 25/57 = 43.9% for Alert to Takeover in the highway scenario. These estimated matrices are subsequently compared across environment, information level, and sex using chi-square tests. No step defines its target conclusion into an input: the thresholds are stated before the comparisons and are not optimized to produce the reported effects, and the transition probabilities are directly estimated rather than fitted to force a claimed prediction. The goodness-of-fit tests use the same data that generated the model, which is standard model checking rather than definitional circularity. The paper cites its own prior work, notably [14] for the SymAware architecture and the participatory workshop that informed the study design, but that citation is background context and does not carry the statistical derivation; no uniqueness theorem, fitted ansatz, or unpublished claim is imported to force the Markov-chain results. The abstract's claim that AV transparency yields significant transition differences is contradicted by the paper's own chi-square results (Section IV.C.2 reports chi-square = 2.24, p = 0.326 for highway and chi-square = 1.04, p = 0.593 for suburbs; Section IV.D.2 reports chi-square = 7.31, p = 0.120 for information-level subgroups), but that is an internal-consistency or correctness problem, not a circularity. No specific reduction of a result to its own input can be exhibited, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- LoA state thresholds =
T<=0.5, 0.5<A<2.0, N>=2.0
- Trust score coding =
-1/1 per item, sum range -2 to 4
- Comfort coding =
relaxed=1, anxious=-1, neutral/unsure=0
- Start state assumption =
Normal (LoA=3)
assumptions (5)
- domain assumption First-order Markov property holds for driver behavior
- domain assumption AI-generated image and roleplay vignettes elicit realistic driver behavior
- standard math Chi-square tests on transition matrices assume independent observations
- domain assumption SAE levels of autonomy form an ordinal scale for driver intent
- domain assumption Adapted DBQ and AV-NARS questionnaires are valid measures
Cite this review
Pith. "Pith review of Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles." pith.science (2026). https://pith.science/paper/SCNNOUVQ
@misc{pith2026250607633,
author = {Pith},
title = {Pith review of: Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles},
year = {2026},
howpublished = {\url{https://pith.science/paper/SCNNOUVQ}},
note = {Machine review of arXiv:2506.07633}
}
read the original abstract
Current robotic agents, such as autonomous vehicles (AVs) and drones, need to deal with uncertain real-world environments with appropriate situational awareness (SA), risk awareness, coordination, and decision-making. The SymAware project strives to address this issue by designing an architecture for artificial awareness in multi-agent systems, enabling safe collaboration of autonomous vehicles and drones. However, these agents will also need to interact with human users (drivers, pedestrians, drone operators), which in turn requires an understanding of how to model the human in the interaction scenario, and how to foster trust and transparency between the agent and the human. In this work, we aim to create a data-driven model of a human driver to be integrated into our SA architecture, grounding our research in the principles of trustworthy human-agent interaction. To collect the data necessary for creating the model, we conducted a large-scale user-centered study on human-AV interaction, in which we investigate the interaction between the AV's transparency and the users' behavior. The contributions of this paper are twofold: First, we illustrate in detail our human-AV study and its findings, and second we present the resulting Markov chain models of the human driver computed from the study's data. Our results show that depending on the AV's transparency, the scenario's environment, and the users' demographics, we can obtain significant differences in the model's transitions.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Dignum, Responsible artificial intelligence: How to develop and use AI in a responsible way
V . Dignum, Responsible artificial intelligence: How to develop and use AI in a responsible way . Springer, 2019
work page 2019
-
[2]
J. Ferber and G. Weiss, Multi-agent systems: An introduction to distributed artificial intelligence . Addison-wesley Reading, 1999, vol. 1
work page 1999
-
[3]
A security-aware metamodel for multi-agent systems (MAS),
G. Beydoun, G. Low, H. Mouratidis, and B. Henderson-Sellers, “A security-aware metamodel for multi-agent systems (MAS),” Informa- tion and software technology , vol. 51, no. 5, pp. 832–845, 2009
work page 2009
-
[4]
C. Zhang, A. Hammad, and H. Bahnassi, “Collaborative multi-agent systems for construction equipment based on real-time field data cap- turing,” Journal of Information Technology in Construction (ITcon) , vol. 14, no. 16, pp. 204–228, 2009
work page 2009
-
[5]
Toward a theory of situation awareness in dynamic systems,
M. R. Endsley, “Toward a theory of situation awareness in dynamic systems,” Human factors, vol. 37, no. 1, pp. 32–64, 1995
work page 1995
-
[6]
Coalition formation in manufacturing multi-agent systems,
M. Pechoucek, V . Marik, and O. Stepankova, “Coalition formation in manufacturing multi-agent systems,” inProceedings 11th International Workshop on Database and Expert Systems Applications. IEEE, 2000, pp. 241–246
work page 2000
-
[7]
Multiagent systems: A survey from a machine learning perspective,
P. Stone and M. Veloso, “Multiagent systems: A survey from a machine learning perspective,” Autonomous Robots, vol. 8, pp. 345– 383, 2000
work page 2000
-
[8]
A socially adaptable framework for human-robot interaction,
A. Tanevska, F. Rea, G. Sandini, L. Ca ˜namero, and A. Sciutti, “A socially adaptable framework for human-robot interaction,” Frontiers in Robotics and AI , vol. 7, p. 121, 2020
work page 2020
Show all 26 references
-
[9]
Understanding the intentions of others: re-enactment of intended acts by 18-month-old children
A. N. Meltzoff, “Understanding the intentions of others: re-enactment of intended acts by 18-month-old children.” Developmental psychol- ogy, vol. 31, no. 5, p. 838, 1995
1995
-
[10]
Leador: A method for end- to-end participatory design of autonomous social robots,
K. Winkle, E. Senft, and S. Lemaignan, “Leador: A method for end- to-end participatory design of autonomous social robots,” Frontiers in Robotics and AI , vol. 8, p. 704119, 2021
2021
-
[11]
Modeling trust in human-robot interaction: A survey,
Z. R. Khavas, S. R. Ahmadzadeh, and P. Robinette, “Modeling trust in human-robot interaction: A survey,” in Social Robotics: 12th International Conference, ICSR 2020, Golden, CO, USA, November 14–18, 2020, Proceedings 12 . Springer, 2020, pp. 529–541
2020
-
[12]
Fast adaptation with meta-reinforcement learning for trust modelling in human-robot interaction,
Y . Gao, E. Sibirtseva, G. Castellano, and D. Kragic, “Fast adaptation with meta-reinforcement learning for trust modelling in human-robot interaction,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 305–312
2019
-
[13]
Ethics guidelines for trustworthy ai,
E. Commission, “Ethics guidelines for trustworthy ai,” 2019, accessed: 2024-06-05. [Online]. Available: https://digital-strategy.ec.europa.eu/ en/library/ethics-guidelines-trustworthy-ai
2019
-
[14]
Communicating awareness: Designing a framework for cognitive human-agent interaction for autonomous vehicles,
A. Tanevska, A. Ghosh, K. Winkle, G. Castellano, and S. Soudjani, “Communicating awareness: Designing a framework for cognitive human-agent interaction for autonomous vehicles,” in Cars As Social Agents (CarSA): A Perspective Shift in Human-Vehicle Interaction. , 2023
2023
-
[15]
Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,
S. International, “Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,” SAE international , vol. 4970, no. 724, pp. 1–5, 2018
2018
-
[16]
Errors and violations on the roads: a real distinction?
J. Reason, A. Manstead, S. Stradling, J. Baxter, and K. Campbell, “Errors and violations on the roads: a real distinction?” Ergonomics, vol. 33, no. 10-11, pp. 1315–1332, 1990
1990
-
[17]
Psychology in human- robot communication: An attempt through investigation of negative attitudes and anxiety toward robots,
T. Nomura, T. Kanda, T. Suzuki, and K. Kato, “Psychology in human- robot communication: An attempt through investigation of negative attitudes and anxiety toward robots,” in RO-MAN 2004. 13th IEEE in- ternational workshop on robot and human interactive communication (IEEE cata...
2004
-
[18]
Jr; markov chains, cambridge uni,
S. Norris, “Jr; markov chains, cambridge uni,” 1997
1997
-
[19]
PRISM 4.0: Verifica- tion of probabilistic real-time systems,
M. Kwiatkowska, G. Norman, and D. Parker, “PRISM 4.0: Verifica- tion of probabilistic real-time systems,” in Proc. 23rd International Conference on Computer Aided Verification (CAV’11) , ser. LNCS, G. Gopalakrishnan and S. Qadeer, Eds., vol. 6806. Springer, 2011, pp. 585–591
2011
-
[20]
Adaptive and sequential gridding pro- cedures for the abstraction and verification of stochastic processes,
S. Soudjani and A. Abate, “Adaptive and sequential gridding pro- cedures for the abstraction and verification of stochastic processes,” SIAM Journal on Applied Dynamical Systems , vol. 12, no. 2, pp. 921– 956, 2013
2013
-
[21]
G. N. Theocharous, Hierarchical learning and planning in partially observable Markov decision processes . Michigan State University, 2002
2002
-
[22]
Modeling and prediction of human behavior,
A. Pentland and A. Liu, “Modeling and prediction of human behavior,” Neural computation, vol. 11, no. 1, pp. 229–242, 1999
1999
-
[23]
Pedestrian models for autonomous driving part ii: high-level models of human behaviour,
F. Camara, N. Bellotto, S. Cosar, F. Weber, D. Nathanael, M. Althoff, J. Wu, J. Ruenz, A. Dietrich, G. Markkula, et al., “Pedestrian models for autonomous driving part ii: high-level models of human behaviour,” IEEE Transactions on Intelligent Transportation Systems , vol. 22,...
2020
-
[24]
The hierarchical hidden markov model: Analysis and applications,
S. Fine, Y . Singer, and N. Tishby, “The hierarchical hidden markov model: Analysis and applications,” Machine learning, vol. 32, pp. 41– 62, 1998
1998
-
[25]
Policy recognition in the abstract hidden markov model,
H. H. Bui, S. Venkatesh, and G. West, “Policy recognition in the abstract hidden markov model,” Journal of Artificial Intelligence Research, vol. 17, pp. 451–499, 2002
2002
-
[26]
Using knowledge awareness to improve safety of autonomous driving,
A. Calvagna, A. Ghosh, and S. Soudjnai, “Using knowledge awareness to improve safety of autonomous driving,” in 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 2023, pp. 2997–3002
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.