REVIEW 4 major objections 4 minor 20 references
"I don't like things where I do not have control": Participants' Experience of Trustworthy Interaction with Autonomous Vehicles
T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Road type, not dashboard, determines how much control drivers give AVs
desk verdict Honest preliminary AV trust study whose conclusion overstates what the correlations actually show; deserves a serious referee who pushes for the missing inferential tests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a 2x2 between-subjects online vignette study with two factors: interface information level (High versus Low, derived from an earlier participatory design workshop) and scenario order (Highway-first versus Suburbs-first). Each scenario is broken into three scenes, and after each scene participants choose an action from a multiple-choice list whose options map onto SAE Levels of Autonomy 0 to 3; when multiple options are selected, the levels are averaged to produce a continuous LoA score from 0 to 3. Trust is measured by six yes/no-style statements about the car, each coded $+1$ or $-1$ and summed into a total trust score. The load-bearing correlations are between scenario and LoA, between scenario/information and trust, and between AV-NARS scores and both LoA and trust, while free-text answers are categorized using an existing set of trustworthy HRI/HAI design guidelines.
What would settle it
Run a replication where participants rate their preferred autonomy directly on a validated scale instead of inferring it from action choices; if the scenario effect on that rating disappears, the paper's central claim is falsified. Alternatively, hold the perceived risk constant across two different road contexts and show that trust no longer tracks the information level, which would undercut the claim that transparency independently shapes trust.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that when people choose their action in a critical driving situation, the type of scenario (highway versus suburbs) is the most important factor in the level of autonomy they select, whereas trust in the AV is influenced by both the scenario and the amount of information shown on the AV's interface. This is supported by preliminary correlations: the average level of autonomy rating showed no correlation across the two scenarios ($r(204)=-0.002$, $p=0.980$), while trust ratings correlated positively across scenarios ($r(204)=0.580$, $p<0.001$), suggesting that the willingness to hand over control is heavily situation-dependent while trust is more stable and intrinsic to the person. Negative attitudes toward AVs (measured by an adapted NARS) correlated negatively with both autonomy level ($r=-0.369$, $p<0.001$) and trust ($r=-0.583$, $p<0.001$), whereas driving style (measured by an adapted DBQ) did not correlate with either. The qualitative analysis of free-text answers, grouped under published guidelines for trustworthy human-agent interaction, shows that users focus on safety and performance, transparency, and their own ability to take back control, and that they rarely mention privacy or positive experience spontaneously.
Load-bearing premise
The study assumes that averaging the SAE autonomy levels attached to participants' action choices yields a valid continuous measure of how much autonomy a person would accept.
Editorial extensions
If this is right
- If the scenario is the dominant factor in autonomy choices, designers of AV systems should adapt the level of automation to the driving context rather than offering a single fixed autonomy mode.
- Trust is shaped by both the situation and the information displayed, so interface transparency is a lever for trust even when it does not directly change the autonomy level a user selects.
- Negative attitudes toward AVs are a stronger correlate of low trust and low autonomy acceptance than actual driving style, pointing toward attitude-focused interventions or familiarization experiences.
- Users spontaneously voice concerns about safety, transparency, and the ability to retake control, but not about privacy or positive experience in these scenarios, suggesting those topics may need explicit design attention in other contexts.
- The eventual goal of modeling the human driver for multi-agent systems would combine these scenario-dependent autonomy preferences with trust measures to predict when take-over is likely.
Reading between the lines
- The LoA score relies on averaging multiple action choices mapped to SAE levels, an assumption that could be tested in a replication using a single continuous autonomy preference scale; if the scenario effect disappears under direct measurement, the core claim weakens.
- The two scenarios differ not only in road type but also in the nature of the critical event (an acute collision risk versus a slow, ambiguous slowdown), so the 'scenario effect' may actually be a mix of risk type and context; vary the event type within a fixed road setting to separate these.
- The absence of privacy and positive-experience mentions in free text may reflect the particular scenarios rather than a general lack of concern; a study involving shared mobility or long-term ownership might surface those dimensions.
- The trust statements were summed as an interval score from $+1$/$-1$ items, which is a strong measurement assumption; future work could validate the trust scale factorially before relying on its correlations.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a preliminary analysis of a 206-participant online study in which drivers viewed two AV scenarios (highway and suburbs) under high or low interface information, chose actions mapped to SAE Levels 0-3, and rated confidence, comfort, and trust. Section 5 presents Pearson correlations among DBQ, AV-NARS, LoA, and Trust scores, and groups 14 free-text responses according to Trustworthy HAI guidelines. The stated conclusions are that scenario type is the most important factor for choosing a level of autonomy and that trust is influenced by both the scenario and the amount of information presented on the AV's interface.
Significance. If the conclusions were supported by the analyses, the paper would provide useful empirical grounding for scenario-sensitive and information-sensitive design of AV interfaces and would connect HRI/HAI results to the Trustworthy AI guideline framework. The study has concrete strengths: it builds on a participatory design workshop, uses established external instruments (DBQ, NARS, SAE levels), reports df and p-values transparently, and includes qualitative free-text data. However, the headline claims are not derived from the statistical analyses actually reported, and one measurement assumption is load-bearing for the LoA results. The data may well be able to support the conclusions after additional inferential analysis or substantially weakened conclusions.
major comments (4)
- [Section 5 and Section 6] The conclusion in Section 6 that the type of scenario is 'the most important factor' for LoA, and that trust is influenced by both scenario and information level, is not supported by Section 5. Section 5 explicitly states that the authors 'have yet to conduct a full statistical analysis' and reports only Pearson correlations. No test compares LoA between the Highway and Suburbs scenarios, and no test compares Trust between the High and Low Information conditions. The only scenario-related evidence is r(204)=-0.002, p=0.980 between the LoA ratings of the two scenarios; a zero between-subjects correlation of two within-subject measurements is compatible with a strong scenario effect, with individual differences in scenario sensitivity, or with low reliability of the LoA measure, so it cannot establish scenario dependence. The Trust correlations reported in Section 5 involve AV-NARS, DBQ, and LoA, not the information-level condition. A mixed model with scenario, information level, and order as fixed effects and participant as a random effect, or at minimum paired tests and condition comparisons with effect sizes, is needed before the Section 6 claims can be made.
- [Section 4.3] The mapping from multiple-choice action options to SAE Levels of Autonomy is load-bearing but unvalidated. Section 4.3 states that the action choices 'corresponded to a level of autonomy between 0 and 3' and that the score 'was averaged in the cases where participants chose more than one option, so that also the final LoA score ranged from 0 to 3.' This assumes that the options are on an equal-interval scale and that averaging multiple selections yields a meaningful continuous autonomy score. The illustrative options in Section 4.2, such as checking a phone or focusing one's eyes on the road, are not self-evidently SAE levels. The paper should provide the full option-to-level mapping, justify the averaging procedure, and report the distribution and internal consistency of the resulting LoA scores; otherwise the non-correlation between the two scenario LoA ratings is difficult to interpret.
- [Section 4.3] The Trust score is constructed by coding six statements as +1 or -1 and summing them into a score from -2 to 4. This assumes equal item weights and interval-level measurement, but no reliability or validity evidence is reported. Since Trust is one of the two central dependent variables, correlations involving it (for example r(204)=-0.583 with AV-NARS and r(204)=0.576 with LoA) are hard to interpret without evidence that the six items form a coherent scale. The authors should report Cronbach's alpha or item-level analyses, and consider treating the score as ordinal or modeling it with an item response approach.
- [Section 5] The statement that participants' driving style 'did not impact' their attitudes toward AVs is based on non-significant correlations (e.g., DBQ-AV-NARS r(204)=-0.075, p=0.286; DBQ-Trust r(204)=0.120, p=0.087). Absence of statistical significance is not evidence of absence, and the paper reports no confidence intervals or effect sizes for these null results. Additionally, multiple correlations are reported without correction for multiplicity. This is less central than the scenario/information claims, but the framing should be qualified.
minor comments (4)
- [Section 4.2] The design is described as a '2×2 between-subjects design' for information level and scenario order, but scenario is a within-subject factor since every participant experiences both Highway and Suburbs. The wording should distinguish the between-subjects conditions from the repeated-measures factor.
- [Section 4.1] The text says participants were 'equally distributed between male and female' and the footnote reports a self-reported gender ratio of 104:101:1 (F:M:X). This conflates sex and gender; the terminology should be made consistent and precise.
- [Section 4.3 and Section 5] The comfort score is described in data collection as rated -1, 0, or 1, but no comfort results are reported in Section 5. Either report the analysis or state explicitly that comfort data are reserved for future work.
- [Abstract and Section 5] The abstract says the paper analyzes preliminary findings 'within existing guidelines on Trustworthy HAI/HRI,' but in Section 5 the guidelines are applied only to the 14 free-text responses. The scope of the qualitative analysis should be stated more precisely.
Circularity Check
No circularity: the paper's empirical correlations are computed from the collected survey data, and its self-citations are design provenance and consistency checks, not load-bearing derivations.
full rationale
The paper makes no parametric fit and derives no quantity from its own prior work. The LoA score is defined by participants' multiple-choice action selections against SAE Levels 0-3 (Section 4.3); the Trust score is a sum of coded trust statements; both are direct measurement constructs, not outputs of a model trained on the same data. The central reported results — AV-NARS correlates with LoA (r=-0.369) and Trust (r=-0.583), and Trust/LoA correlate (r=0.576) — are Pearson correlations calculated on the current 206-participant sample. The self-citation to the authors' workshop [19] is used to explain the provenance of the High/Low Information conditions and the scenario stimuli, not to justify the magnitude or existence of any reported effect; the statement in Section 5 that the correlation pattern is 'in line with' the earlier workshop is a consistency note, not an evidential input. The external instruments (DBQ, NARS, SAE levels, Trustworthy HAI guidelines [3]) are independent of the paper's claims. The main weakness is inferential: Section 6 asserts that scenario drives LoA and information drives trust, whereas Section 5 reports only correlations and explicitly defers full statistical analysis. That is an over-interpretation, but it is a validity/correctness concern, not a circular reduction. Accordingly, no circular step can be quoted and exhibited, and the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- Action-to-Level-of-Autonomy mapping =
0 to 3
- Trust score coding weights =
+1 and -1
assumptions (3)
- domain assumption SAE Levels of Autonomy are an ordinal scale suitable for measuring participants' preferred control in the vignettes.
- domain assumption Self-report measures of trust, confidence, and comfort capture real-world trust in AVs.
- domain assumption Free-text responses can be reliably grouped into the five Trustworthy HAI guidelines without a formal coding procedure or inter-rater reliability.
Cite this review
Pith. "Pith review of "I don't like things where I do not have control": Participants' Experience of Trustworthy Interaction with Autonomous Vehicles." pith.science (2026). https://pith.science/paper/3AUDUYKF
@misc{pith2026250315522,
author = {Pith},
title = {Pith review of: "I don't like things where I do not have control": Participants' Experience of Trustworthy Interaction with Autonomous Vehicles},
year = {2026},
howpublished = {\url{https://pith.science/paper/3AUDUYKF}},
note = {Machine review of arXiv:2503.15522}
}
read the original abstract
With the rapid advancement of autonomous vehicle (AV) technology, AVs are progressively seen as interactive agents with some level of autonomy, as well as some context-dependent social features. This introduces new challenges and questions, already relevant in other areas of human-robot interaction (HRI) - namely, if an AV is perceived as a social agent by the human with whom it is interacting, how are the various facets of its design and behaviour impacting its human partner? And how can we foster a successful human-agent interaction (HAI) between the AV and the human, maximizing the human's comfort, acceptance, and trust in the AV? In this work, we attempt to understand the various factors that could influence na\"ive participants' acceptance and trust when interacting with an AV in the role of a driver. Through a large-scale online study, we investigate the effect of the AV's autonomy on the human driver, as well as explore which parameters of the interaction have the highest impact on the user's sense of trust in the AV. Finally, we analyze our preliminary findings from the user study within existing guidelines on Trustworthy HAI/HRI.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Ghassan Beydoun, Graham Low, Haralambos Mouratidis, and Brian Henderson- Sellers. 2009. A security-aware metamodel for multi-agent systems (MAS). Information and software technology 51, 5 (2009), 832–845
work page 2009
-
[2]
Andrea Calvagna, Arabinda Ghosh, and Sadegh Soudjnai. 2023. Using Knowledge Awareness to improve Safety of Autonomous Driving. In 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 2997–3002
work page 2023
-
[3]
N. Calvo-Barajas, A. Rouchitsas, and D.G. Broo. In press. Examining Human- Robot Interactions: Design Guidelines for Trust and Acceptance.Interdisciplinary Dimensions of Human-Technology Interaction, Springer (In press)
-
[4]
Patrick Ferdinand Christ, Florian Lachner, Axel Hösl, Bjoern Menze, Klaus Diepold, and Andreas Butz. 2016. Human-drone-interaction: A case study to investigate the relation between autonomy and user experience. In Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part II 14 . Springer, 238–253
work page 2016
-
[5]
European Commission. 2019. Ethics Guidelines for Trustworthy AI. https: //digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai Ac- cessed: 2024-06-05
work page 2019
-
[6]
Henrik Detjen, Sarah Faltaous, Bastian Pfleging, Stefan Geisler, and Stefan Schneegass. 2021. How to increase automated vehicles’ acceptance through in- vehicle interaction design: A review. International Journal of Human–Computer Interaction 37, 4 (2021), 308–330
work page 2021
-
[7]
Jacques Ferber and Gerhard Weiss. 1999. Multi-agent systems: An introduction to distributed artificial intelligence. Vol. 1. Addison-wesley Reading
work page 1999
-
[8]
Yuan Gao, Elena Sibirtseva, Ginevra Castellano, and Danica Kragic. 2019. Fast adaptation with meta-reinforcement learning for trust modelling in human-robot interaction. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 305–312
work page 2019
Show all 20 references
-
[9]
Sae International. 2018. Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles. SAE international 4970, 724 (2018), 1–5
2018
-
[10]
Zahra Rezaei Khavas, S Reza Ahmadzadeh, and Paul Robinette. 2020. Modeling trust in human-robot interaction: A survey. In Social Robotics: 12th International Conference, ICSR 2020, Golden, CO, USA, November 14–18, 2020, Proceedings 12 . Springer, 529–541
2020
-
[11]
Yee Mun Lee, Ruth Madigan, Oscar Giles, Laura Garach-Morcillo, Gustav Markkula, Charles Fox, Fanta Camara, Markus Rothmueller, Signe Alexandra Vendelbo-Larsen, Pernille Holm Rasmussen, et al. 2021. Road users rarely use explicit communication when interacting in today’s traffi...
2021
-
[12]
Sumbal Malik, Manzoor Ahmed Khan, and Hesham El-Sayed. 2021. Collaborative autonomous driving—A survey of solution approaches and future challenges. Sensors 21, 11 (2021), 3783
2021
-
[13]
Tatsuya Nomura, Takayuki Kanda, Tomohiro Suzuki, and Kennsuke Kato. 2004. Psychology in human-robot communication: An attempt through investigation of negative attitudes and anxiety toward robots. In RO-MAN 2004. 13th IEEE international workshop on robot and human interactive ...
2004
-
[14]
Kazuo Okamura and Seiji Yamada. 2020. Calibrating trust in human-drone cooperative navigation. In 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . IEEE, 1274–1279
2020
-
[15]
Kaspar Raats, Vaike Fors, and Sarah Pink. 2020. Trusting autonomous vehicles: An interdisciplinary approach. Transportation Research Interdisciplinary Perspectives 7 (2020), 100201
2020
-
[16]
James Reason, Antony Manstead, Stephen Stradling, James Baxter, and Karen Campbell. 1990. Errors and violations on the roads: a real distinction?Ergonomics 33, 10-11 (1990), 1315–1332
1990
-
[17]
Michael Rettenmaier and Klaus Bengler. 2021. The matter of how and when: comparing explicit and implicit communication strategies of automated vehicles in bottleneck scenarios. IEEE Open Journal of Intelligent Transportation Systems 2 (2021), 282–293
2021
-
[18]
Anirudh Sripada, Pavlo Bazilinskyy, and Joost de Winter. 2021. Automated vehicles that communicate implicitly: examining the use of lateral position within the lane. Ergonomics 64, 11 (2021), 1416–1428
2021
-
[19]
Ana Tanevska, Arabinda Ghosh, Katie Winkle, Ginevra Castellano, and Sadegh Soudjani. 2023. Communicating Awareness: Designing a Framework for Cogni- tive Human-Agent Interaction for Autonomous Vehicles. InCars As Social Agents (CarSA): A Perspective Shift in Human-Vehicle Interaction
2023
-
[20]
Katie Winkle, Emmanuel Senft, and Séverin Lemaignan. 2021. LEADOR: A method for end-to-end participatory design of autonomous social robots. Fron- tiers in Robotics and AI 8 (2021), 704119
2021
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.