{"id":"8b13881f-2c92-4350-8f60-e65aa29c385b","arxiv_id":"2602.19107","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Real-world Chinese robotaxi users value driverless privacy and consistent driving but struggle with rigid stops, poor transparency, and unclear emergency controls; these findings support a four-phase, user-driven design framework.","lead":"This paper reports interviews with 18 robotaxi riders plus 22 first-person rides in China, finding that passengers see driverless cars as private-but-monitored spaces and want clear explanations and easy emergency controls. It turns those experiences into a four-stage design framework for safer, more trustworthy robotaxis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generalizability of the design framework: sample is self-selected, young, China-only, and 16/18 Apollo Go, yet the framework is presented as user-driven for robotaxis generally; the paper's own 'first real-world' claim is also contradicted by cited field studies.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing issue: a homogeneous, self-selected sample cannot support a general user-driven design framework. My stress-test refines this by adding platform concentration (16/18 Apollo Go) and the autoethnographic removal of 'invalid' participant codes as specific mechanisms by which the framework may be narrower than claimed. These are external-validity risks, not internal inconsistencies; the interview protocol, iterative coding, and reflexive analysis are sound for qualitative work. Section 7 does acknowledge some limits, which is why the paper is not fundamentally flawed. The 'first real-world' novelty claim is weaker than the empirical contribution, but it is secondary to the framework's generality. The appropriate verdict remains CONDITIONAL: the paper deserves publication if the claims are scoped to the studied population and validated with more diverse samples, which is exactly the reader's condition. No verdict change is needed.","tokens_in":32689,"tokens_out":4353,"duration_ms":46316,"concrete_test":"Run a stratified replication with a pre-registered protocol: recruit at least 30 participants balanced by age (including 40+), education, and platform (Apollo Go, Pony.ai, WeRide, plus a non-China service such as Waymo), apply the same phased interview protocol (Appendix A), and conduct an independent thematic analysis. Then compare theme recurrence by subgroup. If the core framework themes—semi-private transitional space, fixed-stop friction, explainability need, emergency-control uncertainty—do not recur across older, non-Apollo, or non-China users, the Sec. 5 framework should be rescoped to the studied population; if they do recur, the generalizability concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central deliverable is a 'user-driven design framework' for robotaxis (Sec. 5, Fig. 1) presented as broadly applicable. For that claim to hold, the evidence base must be representative enough to distinguish general robotaxi properties from platform- and cohort-specific ones. It is not. Recruitment is self-selected via social media and snowball sampling (Sec. 3.1.1); all 18 participants are aged 19–29, almost all hold at least a bachelor's degree, 16/18 use Apollo Go, and all ride in China (Table 1). Section 7 concedes these demographic limitations, but the prose elsewhere repeatedly frames the framework as a general user-driven contribution for robotaxi systems (Secs. 1, 5, 8). More specifically, the dominant platform's current interface—fixed pick-up/drop-off stops, opaque wait-time estimates, and buried emergency controls—drives many of the coded frictions, while the autoethnographic 'validation' step removed participant codes the researchers judged invalid (Sec. 3.2.3). This risks producing a framework that is partly an Apollo Go / young-adult / author-experience account labeled 'user-driven.' The internal qualitative analysis is careful and transparent, but external validity is the load-bearing hinge. Additionally, the 'first real-world' novelty claim (Secs. 1, 2.3) is contradicted by the paper's own cited field studies of real robotaxi or robo-taxi services (e.g., [25, 36, 84]), which is a scoping overreach independent of the empirical contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative field study of robotaxi user experience in China, combining 18 semi-structured interviews with users and 22 autoethnographic rides by the two authors. The findings identify motivations (low cost, recommendations, curiosity), perceived advantages (agency, behavioral consistency, transparency), and frictions (limited flexibility, insufficient transparency, management, robustness, and emergency-handling concerns), as well as perceptions of privacy, safety, ethics, and trust. The central claim is that users do not experience robotaxis merely as transport but as 'autonomous, semi-private transitional spaces,' and the paper derives a four-phase (hailing, pick-up, traveling, drop-off) user-driven design framework with eleven design directions. The method is reported in unusual detail, including a full interview protocol, pilot-inclusion policy, saturation check, and explicit limitations.","tokens_in":33035,"tokens_out":6071,"duration_ms":68035,"significance":"If read as a transferable qualitative account of current robotaxi use in a major deployment context, the paper is a useful contribution. Its strengths are the transparency of the qualitative protocol, the use of authentic ride experiences in addition to interviews, and the concrete, stage-based design directions. The 'semi-private transitional space' idea is a plausible and potentially generative conceptualization. However, the contribution is currently framed as a general 'user-driven design framework for robotaxis' and as the 'first real-world account,' both of which outrun the evidence. The stress-test concern about generalizability lands, and the circularity concern partially lands with respect to the autoethnographic validation step. These are fixable with scoping, but they affect the central claim and therefore require revision before acceptance.","major_comments":[{"comment":"The paper twice claims to be the 'first' real-world empirical investigation of robotaxi user experience ('this work presents the first empirical investigation…', §2.3; 'We provide the first real-world account…', contributions). This is contradicted by the manuscript's own cited literature: [36] is a field-test study of robo-taxi user acceptance, [25] reports a field study of robotaxi HMI and perceived safety, [84] uses field tests to study robotaxi anxiety and HMI relief, and [53] is a real-life WOZ study of robo-taxi passengers. Please either narrow the novelty claim to what is genuinely new (e.g., the combination of interviews with multi-platform autoethnography, or the specific analytical lens of semi-private transitional space) or remove the 'first' wording.","section":"§2.3 and contributions"},{"comment":"The autoethnographic analysis is described as both 'supplementing and validating' the interview-derived codebook, and the authors state that 'invalid codes were removed.' Two examples are given: participants who reported that heated seats were unavailable, and participants who misunderstood in-cabin cameras as non-recording. Since the object of study is user perception, a participant's factually incorrect belief is itself data, not noise. The authors do say that misunderstandings are 'treated as meaningful perceptions,' but the blanket removal of 'invalid' codes remains in tension with the user-driven framing. Please separate factual verification from perceptual coding: retain perceptual codes with an analytic annotation about objective feature availability, and specify how disagreements between autoethnographic and interview evidence were resolved, rather than using the researchers' rid","section":"§3.2.3"},{"comment":"The framework is presented as a general 'user-driven design framework for robotaxis' (Sections 1, 5, 8 and the abstract), but the empirical base is 18 self-selected participants aged 19–29, mostly bachelor's degree or above, 16/18 using Apollo Go, all in China. Section 7 acknowledges the demographic limits and calls the framework 'initial,' but the framing throughout the paper is much stronger. Many of the coded frictions (fixed stops, opaque wait estimates, emergency-button visibility, pricing practices) are platform- and cohort-specific, likely reflecting Apollo Go's current interface and the expectations of young, educated Chinese users. To make the central claim defensible, please (a) scope the framework explicitly as a transferable starting point grounded in China/Apollo Go use; (b) add an analysis of which findings are likely platform-dependent vs. common to current robotaxi system","section":"§3.1.1, Table 1, §5, §7"}],"minor_comments":[{"comment":"Typo: 'In our autoethnographic, we found…' should read 'In our autoethnographic study…' or 'In our autoethnography, we found…' Also, hyphenation of 'autoethnographic' is inconsistent across the manuscript.","section":"§4.4.1"},{"comment":"Participant labels are inconsistent: Table 1 uses P1–P18, but the text often uses P01, P06, P04, etc. Please standardize to a single convention.","section":"Throughout"},{"comment":"The sentence 'Data collection ceased when no new themes emerged' gives no evidence for saturation in the autoethnographic data. Please report how many rides were analyzed before this decision, or otherwise document the saturation judgment.","section":"§3.2.3"},{"comment":"The framework diagram lists four phases and eleven design dimensions, but the relationships among the dimensions are not explained. A sentence on how the dimensions were associated with phases (e.g., why 'In-ride Experience Enrichment' is only under Traveling) would improve readability.","section":"Figure 1 and §5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid qualitative UX study with unusually transparent methods and rich data. My main concern is scope: the 'first real-world' novelty claim is demonstrably too strong, and the framework's label overstates the diversity of the evidence base. These are fixable with careful rewording and an additional transferability analysis. I would not reject the paper; with the generalizability framing corrected and the autoethnographic validation step clarified, it would be a useful contribution to the robotaxi HCI literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful, unusually transparent qualitative study of real robotaxi use in China, and the design framework is a reasonable synthesis. But it is not the first real-world study of robotaxi UX, and the framework is overgeneralized relative to the homogeneous sample. Both are fixable.\n\nThe authors did 18 semi-structured interviews (all 19–29, mostly bachelor's degrees, 16/18 Apollo Go) plus 22 autoethnographic rides across several platforms and cities. The method section is a model: full interview protocol in the appendix, pilot inclusion explained, iterative coding with two researchers, saturation check, and a limitations section that does not pretend the sample is diverse. The findings—agency, behavioral consistency, fixed-stop friction, transparency gaps, emergency-control confusion, polarized safety views, the semi-private space framing—are plausible and well-supported by quotes. The framework (hailing/pickup/traveling/drop-off) is concrete enough for practitioners to react to.\n\nSoft spots, in order. First, the 'first real-world' claim in the Introduction and Section 2.3 is contradicted by the paper's own references: Lee et al. [36] analyzed field test data, Gu et al. [25] ran a field study, Yoo et al. [84] ran field tests, and Meurer et al. [53] ran rides in real-life settings with a hidden driver. The novelty should be 'first to combine interviews with researcher autoethnography on commercial robotaxis,' or something equally scoped. Second, the sample supports an Apollo Go, young-adult, urban-China story much better than a general 'user-driven design framework.' The authors concede this in Section 7, but the abstract and conclusion present the framework as broadly applicable. A scoped title or abstract would be honest. Third, Section 3.2.3 has the researchers' autoethnographic rides acting as arbiter of which participant codes are 'valid'; they say they kept misunderstandings as perceptions, but they also removed invalid codes. That is a mild circularity worth clarifying in the method.\n\nBottom line: the empirical core is solid, the paper deserves a real referee, and with scoped claims it would be a solid venue paper. I'd read it again for the method transparency; I wouldn't lean on the generality of the framework without replications. For HCI researchers and robotaxi designers it's a useful, citable study—just not the definitive general framework it claims to be.","headline":"Careful, transparent qualitative study of real robotaxi riders in China; the framework is useful but overgeneralized, and the 'first real-world' claim doesn't survive contact with its own references.","tokens_in":33526,"tokens_out":2493,"would_cite":true,"duration_ms":26550,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Real-world robotaxi users describe the cabin as an autonomous, semi-private space—not merely a ride—and this reframes robotaxi design.","keywords":["robotaxi","autonomous driving","user experience","design framework","semi-private space","trust","privacy","autoethnography"],"falsifier":"A larger, more diverse study in a different cultural context—for example, with older adults or users in North America or Europe—would falsify or sharply qualify the central claim if it found that passengers experience in-cabin monitoring as intrusive rather than discreet, and that the 'semi-private transitional space' framing does not hold. Concretely, a survey or interview study with more than 100 participants across age and culture groups, measuring perceived privacy intrusion and felt freedom to engage in personal activities, would settle whether the finding generalizes beyond the studied p","tokens_in":32553,"feed_emoji":"🚕","tokens_out":1883,"duration_ms":21745,"temperature":0.7,"pith_summary":"This paper claims that people who have actually ridden in robotaxis experience them as autonomous, semi-private transitional spaces where they feel freer to relax, work, or talk than in a human-driven taxi, even while knowing the vehicle records them. Through 18 interviews and 22 first-person test rides in China, it identifies the concrete frictions that keep this space from feeling trustworthy: rigid pick-up and drop-off points, over-cautious driving, unclear emergency controls, and thin explanations for unexpected behavior. The paper turns these findings into a user-driven design framework covering hailing, pick-up, traveling, and drop-off phases, with specific features such as pre-ride configuration, context-aware pickup facilitation, driving-behavior explainability, and accountable feedback. If correct, the framework gives designers and operators a practical, evidence-based starting point for making robotaxis feel safe and accountable rather than merely functional.","feed_headline":"Riders treat robotaxis as private spaces, not just rides","feed_subtitle":"Interviews and 22 test rides in China yield a four-stage design framework for trustworthy, transparent robotaxis.","key_machinery":"The central object is the user-driven design framework for robotaxis, organized around four journey phases (hailing, pick-up, traveling, drop-off) with twelve proposed design directions, from pre-ride configuration and system disclosure to context-aware drop-off facilitation and accountable feedback. The framework is carried by the mixed qualitative method: 18 semi-structured interviews with actual robotaxi users and 22 autoethnographic rides by the authors, analyzed through inductive thematic analysis. The interpretive lens tying everything together is the concept of the robotaxi as an 'autonomous, semi-private transitional space,' which is used to explain why agency, privacy indifference,","core_discovery":"The central claim is that, in real-world use, passengers perceive robotaxis not just as transport but as autonomous, semi-private transitional spaces: the absence of a human driver lowers the social intrusion of monitoring and in-vehicle presence, so users feel freer to engage in personal activities, while remaining aware of in-cabin recording and adjusting their behavior. This perception coexists with polarized safety views (some fear reduced control, others feel safer than with human drivers), persistent frictions around limited flexibility and transparency, and ethical concerns about accountability. The paper argues from these findings that robotaxi design should span the full journey—hai","pith_inferences":["A broader and more diverse sample, including older adults, families with children, and users outside China, might reveal that the 'semi-private space' perception is culturally specific: in societies with stronger privacy norms, in-cabin monitoring could produce more resistance than indifference, which would shift the design priorities.","The framework is testable in operational deployments: comparing satisfaction and trust metrics before and after implementing pre-ride disclosure, physical emergency buttons, or context-aware drop-off facilitation would isolate which of the twelve directions carries the most weight.","If the semi-private space perception becomes a design goal, it suggests new metrics for robotaxi UX—such as 'felt freedom to engage in personal activities' or 'perceived social intrusion of monitoring'—that current acceptance models do not capture.","The authors' autoethnographic finding that safety instructions are inconsistent across interfaces points to a regulatory opportunity: standardizing emergency signage and in-vehicle safety education could become a low-cost, high-impact policy intervention for robotaxi operators."],"forward_implications":["If accepted, the four-phase framework gives design teams a concrete checklist for robotaxi UX, directly prioritizing transparency (arrival-time and pricing breakdowns) and emergency controls (visible physical buttons and pre-ride safety education).","The finding that users feel less socially intruded-upon in a robotaxi suggests that interior design should deliberately exploit this semi-private space with entertainment, relaxation, and companion features, rather than treating the cabin as purely functional.","The polarized safety perceptions imply that trust is constructed through explanations and observable system behavior, so that explainability is not an optional extra but a core requirement for robotaxi acceptability.","The ethical concerns about accountability after customization imply that legal frameworks for autonomous driving must be developed in parallel with any user-facing customization features, not after them.","The empirical grounding in real deployment distinguishes this work from simulation- or wizard-of-oz studies, providing a practice-driven baseline for future robotaxi design research."],"fun_headline_variants":["Robotaxis: semi-private spaces that change rider behavior","Riders perceive robotaxis as private, not just transport","Passengers treat robotaxis like personal rooms","Robotaxis double as private, transitional spaces"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a self-selected group of 18 young, mostly highly educated Apollo Go users in China, plus 22 rides by the two authors, is enough to characterize robotaxi user experience generally and to validate the proposed design framework.","fun_headline_variants_meta":{"raw":{"variants":["Robotaxis: semi-private spaces that change rider behavior","Riders perceive robotaxis as private, not just transport","Passengers treat robotaxis like personal rooms","Robotaxis double as private, transitional spaces"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":2859,"prompt_tokens":683,"completion_tokens":2176,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":2116}},"tokens_in":427,"tokens_out":2176,"duration_ms":15114,"temperature":1.0,"reasoning_tokens":2116,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T21:43:23.809731+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A larger, more diverse study in a different cultural context—for example, with older adults or users in North America or Europe—would falsify or sharply qualify the central claim if it found that passengers experience in-cabin monitoring as intrusive rather than discreet, and that the 'semi-private transitional space' framing does not hold. Concretely, a survey or interview study with more than 100 participants across age and culture groups, measuring perceived privacy intrusion and felt freedom to engage in personal activities, would settle whether the finding generalizes beyond the studied p","supporting_citations":[],"review_version":1}