{"id":"e83af99c-c0e3-45de-9295-47e49a741c7d","arxiv_id":"2503.15522","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 206-person online vignette study finds that negative attitudes toward autonomous vehicles correlate with lower trust and lower accepted autonomy, while chosen autonomy varies strongly by scenario and trust is more stable.","lead":"This paper reports a large online study of how people decide how much control to keep when riding in an autonomous vehicle. It finds that people with more negative attitudes toward AVs trust them less and accept less autonomy, and that trust is more stable across situations than the chosen level of autonomy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5 reports only correlations; no inferential test supports the Section 6 claims that scenario drives LoA and information level drives trust.","rationale":"The reader's weakest assumption concerned the validity of the LoA mapping and averaging, which is a real measurement concern. My stress-test identifies a different but complementary load-bearing issue: even if the LoA score is valid, the reported analyses do not test the paper's central claims. The conclusion attributes importance to scenario and information level, but Section 5 contains only correlations, and the one correlation used to infer scenario dependence (LoA Scenario 1 vs Scenario 2, r=-0.002) is ambiguous because it could indicate measurement unreliability rather than a scenario effect. The trust-information claim is never analyzed at all. This is not a question of internal inconsistency or fraud; the authors candidly label the work preliminary and state that full statistical analysis is pending. However, the conclusion overstates what the current results show. The proposed check--fitting mixed models with scenario, information level, and order--would directly test both headline claims and settle whether the conclusion is justified. Since the reader already assigned CONDITIONAL based on missing inferential tests, my concern reinforces that verdict rather than changing it.","tokens_in":7492,"tokens_out":2684,"duration_ms":26290,"concrete_test":"Obtain the raw per-participant responses and fit linear mixed models: LoA ~ Scenario + InfoLevel + Order + (1|Participant) and Trust ~ Scenario + InfoLevel + Order + (1|Participant). Report fixed-effect estimates, 95% confidence intervals, and p-values. If the Scenario effect on LoA is not significant, or if the InfoLevel effect on Trust is not significant, then the central claims in Section 6 are not supported by the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion (Section 6) claims scenario type is 'the most important factor' when selecting LoA and that trust is influenced by both scenario and interface information. Section 5, however, explicitly states that the authors 'have yet to conduct a full statistical analysis' and reports only Pearson correlations. No analysis compares LoA between the Highway and Suburbs scenarios, and no analysis compares Trust between the High and Low Information conditions. The 'scenario dependence' inference rests entirely on the non-correlation between LoA ratings across scenarios (r(204)=-0.002, p=0.980), but a zero correlation is equally compatible with a strong scenario effect and with low reliability of the LoA measure. The same participants' two scenario ratings are never compared in a paired test, so the absence of correlation cannot distinguish scenario-driven variation from measurement noise. The claim that trust is influenced by the amount of information is not tested anywhere in Section 5; the only trust correlations reported involve AV-NARS, DBQ, and LoA. Without a model that includes scenario, information level, and order as fixed factors, the paper's headline findings are unsupported by the analyses presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a preliminary analysis of a 206-participant online study in which drivers viewed two AV scenarios (highway and suburbs) under high or low interface information, chose actions mapped to SAE Levels 0-3, and rated confidence, comfort, and trust. Section 5 presents Pearson correlations among DBQ, AV-NARS, LoA, and Trust scores, and groups 14 free-text responses according to Trustworthy HAI guidelines. The stated conclusions are that scenario type is the most important factor for choosing a level of autonomy and that trust is influenced by both the scenario and the amount of information presented on the AV's interface.","tokens_in":7676,"tokens_out":6247,"duration_ms":56457,"significance":"If the conclusions were supported by the analyses, the paper would provide useful empirical grounding for scenario-sensitive and information-sensitive design of AV interfaces and would connect HRI/HAI results to the Trustworthy AI guideline framework. The study has concrete strengths: it builds on a participatory design workshop, uses established external instruments (DBQ, NARS, SAE levels), reports df and p-values transparently, and includes qualitative free-text data. However, the headline claims are not derived from the statistical analyses actually reported, and one measurement assumption is load-bearing for the LoA results. The data may well be able to support the conclusions after additional inferential analysis or substantially weakened conclusions.","major_comments":[{"comment":"The conclusion in Section 6 that the type of scenario is 'the most important factor' for LoA, and that trust is influenced by both scenario and information level, is not supported by Section 5. Section 5 explicitly states that the authors 'have yet to conduct a full statistical analysis' and reports only Pearson correlations. No test compares LoA between the Highway and Suburbs scenarios, and no test compares Trust between the High and Low Information conditions. The only scenario-related evidence is r(204)=-0.002, p=0.980 between the LoA ratings of the two scenarios; a zero between-subjects correlation of two within-subject measurements is compatible with a strong scenario effect, with individual differences in scenario sensitivity, or with low reliability of the LoA measure, so it cannot establish scenario dependence. The Trust correlations reported in Section 5 involve AV-NARS, DBQ, and LoA, not the information-level condition. A mixed model with scenario, information level, and order as fixed effects and participant as a random effect, or at minimum paired tests and condition comparisons with effect sizes, is needed before the Section 6 claims can be made.","section":"Section 5 and Section 6"},{"comment":"The mapping from multiple-choice action options to SAE Levels of Autonomy is load-bearing but unvalidated. Section 4.3 states that the action choices 'corresponded to a level of autonomy between 0 and 3' and that the score 'was averaged in the cases where participants chose more than one option, so that also the final LoA score ranged from 0 to 3.' This assumes that the options are on an equal-interval scale and that averaging multiple selections yields a meaningful continuous autonomy score. The illustrative options in Section 4.2, such as checking a phone or focusing one's eyes on the road, are not self-evidently SAE levels. The paper should provide the full option-to-level mapping, justify the averaging procedure, and report the distribution and internal consistency of the resulting LoA scores; otherwise the non-correlation between the two scenario LoA ratings is difficult to interpret.","section":"Section 4.3"},{"comment":"The Trust score is constructed by coding six statements as +1 or -1 and summing them into a score from -2 to 4. This assumes equal item weights and interval-level measurement, but no reliability or validity evidence is reported. Since Trust is one of the two central dependent variables, correlations involving it (for example r(204)=-0.583 with AV-NARS and r(204)=0.576 with LoA) are hard to interpret without evidence that the six items form a coherent scale. The authors should report Cronbach's alpha or item-level analyses, and consider treating the score as ordinal or modeling it with an item response approach.","section":"Section 4.3"},{"comment":"The statement that participants' driving style 'did not impact' their attitudes toward AVs is based on non-significant correlations (e.g., DBQ-AV-NARS r(204)=-0.075, p=0.286; DBQ-Trust r(204)=0.120, p=0.087). Absence of statistical significance is not evidence of absence, and the paper reports no confidence intervals or effect sizes for these null results. Additionally, multiple correlations are reported without correction for multiplicity. This is less central than the scenario/information claims, but the framing should be qualified.","section":"Section 5"}],"minor_comments":[{"comment":"The design is described as a '2×2 between-subjects design' for information level and scenario order, but scenario is a within-subject factor since every participant experiences both Highway and Suburbs. The wording should distinguish the between-subjects conditions from the repeated-measures factor.","section":"Section 4.2"},{"comment":"The text says participants were 'equally distributed between male and female' and the footnote reports a self-reported gender ratio of 104:101:1 (F:M:X). This conflates sex and gender; the terminology should be made consistent and precise.","section":"Section 4.1"},{"comment":"The comfort score is described in data collection as rated -1, 0, or 1, but no comfort results are reported in Section 5. Either report the analysis or state explicitly that comfort data are reserved for future work.","section":"Section 4.3 and Section 5"},{"comment":"The abstract says the paper analyzes preliminary findings 'within existing guidelines on Trustworthy HAI/HRI,' but in Section 5 the guidelines are applied only to the 14 free-text responses. The scope of the qualitative analysis should be stated more precisely.","section":"Abstract and Section 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparently labeled as preliminary, and the data may well be able to support the conclusions after additional inferential analysis. The main risk is the gap between the correlational results in Section 5 and the causal and comparative claims in Section 6; this is addressable in revision. The paper is within the scope of a short HRI/HAI paper, but the authors should either add the necessary analyses or substantially weaken the conclusions to match what is actually tested."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read: the paper is an honest preliminary report of a 206-person online AV study, but the conclusion runs ahead of the statistics. The useful bits: they adapt NARS to AVs, use two driving scenarios with multiple scenes, and collect both action choices (mapped to SAE LoA) and trust ratings. The new empirical observation is that trust ratings correlate across scenarios while chosen LoA does not (r=-0.002), suggesting trust may be more stable and autonomy preference more context-dependent. That is worth following up.\n\nWhat the paper does well: the writing is plain, the limitations are partially acknowledged (they say they have yet to do full statistical analysis), and the qualitative free-text analysis is reported straightforwardly. The connection to the earlier participatory workshop is handled without claiming more than consistency.\n\nSoft spots, in order of severity. First, the central claims in the conclusion — that scenario is the most important factor for autonomy and that information level influences trust — are not supported by anything in Section 5. There is no paired comparison of LoA between Highway and Suburbs, and no test involving the information-level condition at all. The zero correlation between LoA across scenarios could just as easily reflect low reliability of the LoA measure as a scenario effect. Second, the construction of the LoA score is under-validated: mapping multiple-choice actions to SAE levels and averaging when multiple options are selected assumes interval-level properties that are not defended. Third, the trust score sums six +1/-1 items into a range of -2 to 4, which is another interval-level assumption. None of these are fatal for a preliminary report, but they need to be acknowledged or tested before the headline claims can stand.\n\nThe citation pattern looks fine; the self-citation to the earlier workshop is appropriate, and adapting NARS is a reasonable extension. The paper would benefit from at least a paired test or a mixed model with scenario, information level, and order as fixed factors. As is, it is a useful dataset description and a candidate for revision, not a finished result.\n\nRecommendation: send it to peer review, but the reviewers should push for the missing inferential tests and a validation or caveat for the LoA mapping. It is the kind of paper that can become solid after revision.\n\nBest,","headline":"Honest preliminary AV trust study whose conclusion overstates what the correlations actually show; deserves a serious referee who pushes for the missing inferential tests.","tokens_in":8192,"tokens_out":1717,"would_cite":false,"duration_ms":16049,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Road type, not dashboard, determines how much control drivers give AVs","keywords":["autonomous vehicles","trust","level of autonomy","scenario effect","interface transparency","online user study","driver attitudes","human-agent interaction"],"falsifier":"Run a replication where participants rate their preferred autonomy directly on a validated scale instead of inferring it from action choices; if the scenario effect on that rating disappears, the paper's central claim is falsified. Alternatively, hold the perceived risk constant across two different road contexts and show that trust no longer tracks the information level, which would undercut the claim that transparency independently shapes trust.","tokens_in":7294,"feed_emoji":"🚗","tokens_out":4851,"duration_ms":44899,"temperature":0.7,"pith_summary":"This paper reports a large online study in which 206 licensed drivers were shown two driving scenarios—a busy highway with a sudden merge and a quiet suburban street with an unexplained slowdown—and asked what action they would take, choosing from options mapped to SAE Levels of Autonomy 0 through 3, then rating their trust in the vehicle. The central finding is that the type of scenario is the most important factor in the level of autonomy participants were comfortable granting the AV, while trust was influenced by both the scenario and the amount of information the AV's interface displayed. Sympathetically read, the paper is trying to establish that user acceptance and trust in autonomous vehicles are context-dependent, and that transparency (low versus high information display) plays a distinct role from the situation in shaping trust. The authors also interpret participants' free-text comments using an existing set of trustworthy human-agent interaction guidelines, finding that safety, transparency, and control dominate users' concerns, while privacy and positive experience were never spontaneously raised. This work is a step toward a computational model of the human driver for multi-agent driving systems.","feed_headline":"Road type, not dashboard, determines how much control drivers give AVs","feed_subtitle":"A 206-driver study: trust tracks both situation and interface detail, while autonomy choices follow the road.","key_machinery":"The central object is a 2x2 between-subjects online vignette study with two factors: interface information level (High versus Low, derived from an earlier participatory design workshop) and scenario order (Highway-first versus Suburbs-first). Each scenario is broken into three scenes, and after each scene participants choose an action from a multiple-choice list whose options map onto SAE Levels of Autonomy 0 to 3; when multiple options are selected, the levels are averaged to produce a continuous LoA score from 0 to 3. Trust is measured by six yes/no-style statements about the car, each coded $+1$ or $-1$ and summed into a total trust score. The load-bearing correlations are between scenario and LoA, between scenario/information and trust, and between AV-NARS scores and both LoA and trust, while free-text answers are categorized using an existing set of trustworthy HRI/HAI design guidelines.","core_discovery":"On the paper's own terms, the discovery is that when people choose their action in a critical driving situation, the type of scenario (highway versus suburbs) is the most important factor in the level of autonomy they select, whereas trust in the AV is influenced by both the scenario and the amount of information shown on the AV's interface. This is supported by preliminary correlations: the average level of autonomy rating showed no correlation across the two scenarios ($r(204)=-0.002$, $p=0.980$), while trust ratings correlated positively across scenarios ($r(204)=0.580$, $p<0.001$), suggesting that the willingness to hand over control is heavily situation-dependent while trust is more stable and intrinsic to the person. Negative attitudes toward AVs (measured by an adapted NARS) correlated negatively with both autonomy level ($r=-0.369$, $p<0.001$) and trust ($r=-0.583$, $p<0.001$), whereas driving style (measured by an adapted DBQ) did not correlate with either. The qualitative analysis of free-text answers, grouped under published guidelines for trustworthy human-agent interaction, shows that users focus on safety and performance, transparency, and their own ability to take back control, and that they rarely mention privacy or positive experience spontaneously.","pith_inferences":["The LoA score relies on averaging multiple action choices mapped to SAE levels, an assumption that could be tested in a replication using a single continuous autonomy preference scale; if the scenario effect disappears under direct measurement, the core claim weakens.","The two scenarios differ not only in road type but also in the nature of the critical event (an acute collision risk versus a slow, ambiguous slowdown), so the 'scenario effect' may actually be a mix of risk type and context; vary the event type within a fixed road setting to separate these.","The absence of privacy and positive-experience mentions in free text may reflect the particular scenarios rather than a general lack of concern; a study involving shared mobility or long-term ownership might surface those dimensions.","The trust statements were summed as an interval score from $+1$/$-1$ items, which is a strong measurement assumption; future work could validate the trust scale factorially before relying on its correlations."],"forward_implications":["If the scenario is the dominant factor in autonomy choices, designers of AV systems should adapt the level of automation to the driving context rather than offering a single fixed autonomy mode.","Trust is shaped by both the situation and the information displayed, so interface transparency is a lever for trust even when it does not directly change the autonomy level a user selects.","Negative attitudes toward AVs are a stronger correlate of low trust and low autonomy acceptance than actual driving style, pointing toward attitude-focused interventions or familiarization experiences.","Users spontaneously voice concerns about safety, transparency, and the ability to retake control, but not about privacy or positive experience in these scenarios, suggesting those topics may need explicit design attention in other contexts.","The eventual goal of modeling the human driver for multi-agent systems would combine these scenario-dependent autonomy preferences with trust measures to predict when take-over is likely."],"supporting_citations":[{"why":"Supplies the two scenario designs (highway with a merge, suburbs with a slowdown) and the Low/High interface information conditions from the earlier participatory design workshop.","marker":"[19]"},{"why":"Provides the SAE Levels of Autonomy 0-3 used to map participants' action choices to numerical LoA scores.","marker":"[9]"},{"why":"The NARS questionnaire adapted as AV-NARS, whose scores correlate negatively with both LoA and Trust.","marker":"[13]"},{"why":"The Driver Behaviour Questionnaire adapted to test whether driving style predicts attitudes toward AVs; results show it does not.","marker":"[16]"},{"why":"The Trustworthy HRI/HAI design guidelines used to group and interpret participants' free-text answers.","marker":"[3]"}],"fun_headline_variants":["Autonomy choices depend on road type, trust on interface detail","Drivers grant autonomy by scenario, trust by dashboard","Why drivers hand over control: it's the road, not the screen","Trust in AVs is stable; giving up control is situational"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that averaging the SAE autonomy levels attached to participants' action choices yields a valid continuous measure of how much autonomy a person would accept.","fun_headline_variants_meta":{"raw":{"variants":["Autonomy choices depend on road type, trust on interface detail","Drivers grant autonomy by scenario, trust by dashboard","Why drivers hand over control: it's the road, not the screen","Trust in AVs is stable; giving up control is situational"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00054,"raw_usage":{"total_tokens":2625,"prompt_tokens":1020,"completion_tokens":1605,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":1534}},"tokens_in":636,"tokens_out":1605,"duration_ms":11581,"temperature":1.0,"reasoning_tokens":1534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:46:39.279184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a replication where participants rate their preferred autonomy directly on a validated scale instead of inferring it from action choices; if the scenario effect on that rating disappears, the paper's central claim is falsified. Alternatively, hold the perceived risk constant across two different road contexts and show that trust no longer tracks the information level, which would undercut the claim that transparency independently shapes trust.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the two scenario designs (highway with a merge, suburbs with a slowdown) and the Low/High interface information conditions from the earlier participatory design workshop."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the SAE Levels of Autonomy 0-3 used to map participants' action choices to numerical LoA scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The NARS questionnaire adapted as AV-NARS, whose scores correlate negatively with both LoA and Trust."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Driver Behaviour Questionnaire adapted to test whether driving style predicts attitudes toward AVs; results show it does not."},{"cited_title":"Calvo-Barajas, A","cited_arxiv_id":null,"evidence_quote":"The Trustworthy HRI/HAI design guidelines used to group and interpret participants' free-text answers."}],"review_version":1}