{"id":"7529d15d-15c6-49f2-b5fd-0cc99456d48a","arxiv_id":"2501.06089","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review, expert interviews, and a 90-respondent survey yield a five-module conceptual framework for developing socially compliant automated vehicles.","lead":"This paper reviews 68 studies on socially compliant automated vehicles, interviews ten experts, and surveys 90 professionals, then proposes a conceptual framework for AVs that respect human social expectations on the road. It matters because it maps the capabilities, such as anticipation, cultural alignment, and mutual adaptation, that AV developers must build for smooth mixed-traffic operation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey-based 'validation' is non-discriminating: Section 4.3.1 treats mean ratings >4.8 on the framework's own capability items as confirmation, but with no control items, no statistical tests, and rankings limited to preselected capabilities, such ratings cannot establish that the framework…","rationale":"The scoping review is genuinely useful: it follows a PRISMA-style protocol, identifies five methodological families, and maps them to scenarios and datasets. The proposed framework is coherent, and its modules do correspond to gaps identified in the literature and expert interviews. The weak link is the epistemic weight assigned to the survey. For the paper's central claim to hold, the survey must be able to disconfirm the framework or at least rank it against alternatives. It cannot: the importance items are the framework's own modules, there are no distractor items, no significance tests are reported, and the ranking tasks are constrained to preselected capabilities. This is not to say the framework is wrong; it is to say that the 'validation' is circular in practice. The reader's weakest assumption focuses on geographic underrepresentation. That limitation is real and explicitly acknowledged, but it is less damaging than the measurement problem: a perfectly global sample would not fix a non-discriminating instrument. I therefore keep the reader's CONDITIONAL verdict. The paper should be accepted conditional on reframing the survey as an exploratory expert-opinion elicitation rather than validation, and ideally on adding an independent check with control capabilities. The scoping review and framework stand on their own merits; the overstated validation claim should be softened.","tokens_in":31187,"tokens_out":4157,"duration_ms":44889,"concrete_test":"From the deposited anonymized data (https://doi.org/10.4121/3a46e61c-f5f0-4399-a4b8-4d146b62a4f7), check whether the survey included any distractor or control capability items (e.g., 'ability to select scenic routes' or 'adaptive infotainment'). If none exist, run a small preregistered follow-up expert panel (n>=30) that mixes the nine framework capabilities with three or four plausible but non-framework distractors, and test whether framework capabilities are rated significantly above the distractors (e.g., paired t-test or Wilcoxon signed-rank test). If framework items are not rated reliably above distractors, the Section 4 validation is non-discriminating and the abstract's 'affirming the significance' claim should be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the proposed framework 'captures the key elements needed' (Section 3.2, Fig. 5) rests heavily on the online survey in Section 4. The survey's design makes positive feedback nearly inevitable: respondents are asked to rate the importance of exactly the nine capabilities that compose Fig. 5 (Section 4.3.1), and all nine exceed 4.8 on a 1-7 scale. This pattern is reported as supporting and verifying the elements proposed in the conceptual framework. But no control capabilities are included, so a generic halo or acquiescence effect—or simply the fact that normatively plausible capabilities are hard to call unimportant—would produce the same result. The long-term ranking (Fig. 12) forces choice among four preselected framework capabilities, so it cannot reveal an omitted high-priority capability. The open-ended 'What else would you expect?' question is collected but not formally coded, and no threshold is reported for when a suggestion would count as a missing framework element. Consequently, the survey tests endorsement of the framework's components, not whether the component set is complete or whether an alternative framework would fare worse. Geographic imbalance (Section 5) is a separate, acknowledged limitation; the more serious problem is that even with a perfectly balanced sample, this instrument could not discriminate the proposed framework from a plausible rival.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a scoping review of research on socially compliant automated vehicles (SCAVs), supplemented by informal expert interviews, and proposes a conceptual framework comprising five elements: socially compliant decision-making, a safety constraint module, ego-versus-network trade-off management, bidirectional behavioral adaptation, and spatial-temporal memory. The framework is then 'evaluated' with an online survey of 90 experts who rated and ranked the importance of nine technical capabilities that correspond directly to the framework's components. The paper claims that the survey results provide valuable validation and affirm the significance of the framework. The scoping review and the qualitative synthesis are systematic, and the authors share their data and code. However, the survey-based validation is non-discriminating: it measures endorsement of the framework's own components without control items, statistical inference, or a formal analysis of open-ended suggestions, so it cannot substantiate the strong validation claims made in the Abstract and Section 4.3.1.","tokens_in":31439,"tokens_out":3565,"duration_ms":35142,"significance":"If the framework is viewed as a synthesis of the literature and expert opinion, it offers a useful structuring of an emerging field and a plausible research agenda. The scoping review is methodical, follows a PRISMA-based process, and covers a broad range of methods, datasets, and scenarios; the replication data and code are a concrete strength. The proposed capabilities (e.g., bidirectional adaptation, spatial-temporal memory) are grounded in gaps identified in the review and the expert interviews. The paper's main weakness is the evidential weight placed on the survey: the ratings and rankings are consistent with the framework by construction, so they add little beyond the initial synthesis. The significance of the paper therefore rests on the review and framework rather than on the claimed empirical validation. With a more modest interpretation of the survey and explicit acknowledgment of its limits, the contribution is valuable.","major_comments":[{"comment":"The survey asks respondents to rate exactly the nine capabilities that compose Fig. 5, and the average rating for every capability exceeds 4.8 on a 1–7 scale. The text states that this result 'supports and verifies the elements proposed in the conceptual framework.' This inference is not justified: because the items were derived from the framework itself, the high ratings show only that experts endorse the framework's components, not that the component set is correct, complete, or better than an alternative. No control items are included, no statistical tests are performed, and there is no comparison with a rival framework. Please either add a discriminant check (e.g., items that are intentionally less central, or a comparison with an alternative component list) or reframe the claim as 'expert endorsement' rather than 'validation.'","section":"Section 4.3.1, Fig. 10"},{"comment":"The ranking tasks restrict respondents to preselected capabilities from the framework, so they cannot reveal an omitted high-priority capability. The open-ended question 'What else would you expect for the socially compliant automated vehicles?' is collected but only summarized qualitatively, with no coding scheme, inter-rater reliability, or a specified threshold for when a suggestion would count as a missing framework element. The central claim that the framework 'captures the key elements needed' (Section 3.2) is therefore not tested for completeness. Please perform a formal content analysis of the open-ended responses and report whether any suggestions fall outside the framework, or explicitly state that completeness remains an open question.","section":"Section 4.3.1, Figs. 11–12; Section 4.3.3"},{"comment":"The authors acknowledge the geographic imbalance in the survey, but the Abstract and Section 4.3.1 still state that the survey provides 'valuable validation and insights, affirming the significance' of the framework. Beyond geography, the sample is self-selected and heavily weighted toward researchers (49 of 90 respondents), with only two OEM developers and seven policymakers. The survey also lacks any inferential statistics. These features severely constrain what the survey can establish, even beyond the acknowledged underrepresentation of the USA. Please moderate the validation claims throughout the manuscript to match what an exploratory, non-representative opinion survey can support, and explicitly list the sample's occupational skew and self-selection as limitations.","section":"Section 5, 'Limitations' paragraph; Abstract"}],"minor_comments":[{"comment":"The entry 'Bayesian ınference' contains a dotless 'ı'; correct this typo to 'inference.'","section":"Table 3"},{"comment":"The flow diagram indicates 1327 unique records after duplicate removal, but the text says 'Together with the manual examination of the titles, a total of 1327 valid unique records were included.' Please clarify whether title screening occurred before or after deduplication, so the reader can follow the PRISMA pipeline precisely.","section":"Section 2.2.1"},{"comment":"The explanation that the USA is grouped under 'Other' appears both in the main text and in the figure note. Move the full explanation to one place to reduce redundancy.","section":"Section 4.1 and Fig. 6 note"},{"comment":"Supplementary Attachment 1 (referenced in Section 2.2.1) and Supplementary Attachment 2 (referenced in Section 4) are both listed under the same shortened URL (https://lnkd.in/gpceU6gQ). Please verify that the link points to the correct documents or provide distinct URLs.","section":"Supplementary attachments"},{"comment":"The definition of socially compliant driving is clear, but consider adding a one-sentence definition in the Abstract or Nomenclature so that the term is immediately understandable to readers who skip the Introduction.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid scoping review and the proposed framework is plausible, but the survey-based validation is the weakest link and is overclaimed. The central claim can be made defensible by reframing the survey as an exploratory opinion survey and adding a formal analysis of the open-ended responses or a discriminant check. If the authors are unwilling to temper the validation language, the manuscript would be borderline for this journal. The shared data and code are a positive aspect that should be preserved. No concerns about citation patterns or novelty disclosure beyond the survey's circularity, which is already a major-comment issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. This is a scoping review plus a conceptual framework for socially compliant AVs, and the review part is genuinely useful. The authors went through 1,553 records, selected 68, grouped them into five methodological families, and did a decent job of charting methods, datasets, and scenarios. The framework itself—five modules: socially compliant decision-making, safety constraint, ego-vs-network trade-off, bidirectional behavioral adaptation, spatial-temporal memory—is a plausible synthesis. It recombines known ideas, but the specific combination, especially the bidirectional adaptation and memory modules, is a useful organizing device for a field that lacks shared structure. The informal expert interviews seem to have informed the framework, and the anonymized data and code are available, which is good practice.\n\nThe soft spot is the online survey. The stress-test you passed on is right. The survey asks respondents to rate exactly the nine capabilities that make up the framework, all nine average above 4.8 on a 1–7 scale, and the paper calls this 'validation' and says it 'affirms the significance' of the framework. There are no control items, no statistical tests, no pre-registered threshold, and the open-ended suggestions are summarized but not formally coded. The long-term ranking (Fig. 12) forces a choice among four preselected capabilities, so it can't reveal an omitted priority. In other words, even with a perfectly balanced sample, this instrument could not discriminate the proposed framework from a plausible rival. The geographic skew (one US respondent) is a real but secondary issue; the authors acknowledge it, but they don't acknowledge the design problem.\n\nI want to be clear: the weak survey doesn't sink the paper. The framework stands on the review and interviews, and the priorities it derives (anticipation near-term, mutual adaptation and memory long-term) are plausible hypotheses, not validated findings. If the authors had labeled the survey 'exploratory stakeholder feedback' rather than 'validation,' the concern would be minor. As written, the abstract and Section 4 overstate it.\n\nWho's this for? Researchers new to SCAV/mixed traffic who want a structured map of methods and a framework to argue against or build on. The survey sections are the ones to read as a cautionary example, not as evidence. I'd send it to peer review: the review itself is solid and the framework deserves to be examined and tested. I'd ask the authors to soften the validation language and add a critical limitations paragraph about the survey's non-discriminating design.","headline":"Useful organizing framework for SCAV research; the survey 'validation' is an endorsement poll and should be read as stakeholder feedback, not evidence.","tokens_in":31939,"tokens_out":2772,"would_cite":true,"duration_ms":26485,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a five-module framework for socially compliant automated vehicles and reports survey evidence from 90 professionals endorsing it.","keywords":["Automated vehicles (AVs)","Socially compliant driving","Mixed traffic","Conceptual framework","Scoping review","Expert survey","Human-driven vehicles","Bidirectional behavioral adaptation"],"falsifier":"If a replication survey with balanced representation from North America, Japan, and India returned an average importance rating at or below the neutral midpoint for any of the nine capabilities, the paper's claim that all framework elements are significant would be contradicted. A sharper test would be a controlled driving experiment in which an AV implementing the full framework is compared with a conventional AV; if human drivers do not predict the framework AV's actions more accurately in merging and left-turn scenarios, the claim that these modules enable social compliance would fail.","tokens_in":30955,"feed_emoji":"🚗","tokens_out":6509,"duration_ms":58859,"temperature":0.7,"pith_summary":"This paper sets out to establish that automated vehicles entering mixed traffic need a dedicated engineering agenda for social compliance, and that the authors' five-module conceptual framework is the right organising structure for that agenda. The paper's evidence chain runs from a scoping review of 68 studies, through interviews with ten experts, to a survey of 90 professionals who rate the framework's nine technical capabilities and rank their priorities. The authors use the survey results to claim that all framework elements are seen as significant, with anticipation of other road users' actions rated most urgent for the near term and bidirectional behavioural adaptation plus spatial-temporal memory rated most critical for the long term. If the framework is accepted, it would turn 'socially compliant driving' from a scattered research theme into a coordinated development roadmap for industry, academia, and policymakers.","feed_headline":"Experts back five-module framework for socially compliant AVs","feed_subtitle":"The expert survey endorses modules for cultural norms, mutual adaptation, safety, and network-wide trade-offs.","key_machinery":"The carrying object is the five-module conceptual framework shown in the paper's Fig. 5. Each module is tied to a specific limitation identified in current AVs: excessive conservatism, inability to read implicit communications, poor adaptation to driving styles, limited scenario anticipation, and cultural inflexibility. The framework also includes standard sensing/perception and communication/eHMI modules, and the survey instrument translates the framework into nine rated technical capabilities, giving the conceptual structure an operational form that experts could judge. That translation from framework modules to survey items is the mechanism by which the paper converts expert opinion into evidence for the framework.","core_discovery":"The central claim is that socially compliant automated driving is best understood and advanced through one integrated framework rather than through isolated algorithms. The proposed framework has five interdependent modules: a socially compliant decision-making module that embeds culture, norms, implicit cues, and driving styles; a safety constraint module that enforces hard safety boundaries on every planned action; an ego-versus-network trade-off module that balances the ego vehicle's benefits against network-level efficiency and societal outcomes; a bidirectional behavioural adaptation module through which AVs and human drivers adjust to each other over time; and a spatial-temporal memory module that stores short- and long-term interaction histories to refine future decisions. The paper treats the 90-respondent expert survey as validation of this structure: all nine rated technical capabilities score above the neutral midpoint of the rating scale, and the ranking exercises place anticipation capability first for medium-term development and bidirectional behavioural adaptation with spatial-temporal memory first for the long term. The contribution is therefore an integrated conceptual map and a prioritised research agenda, with the survey presented as evidence that the map matches expert expectations.","pith_inferences":["If the framework generalises, eHMI design and cultural calibration become first-class engineering requirements rather than optional extras; a testable extension would be matched on-road trials in two countries with different implicit driving norms to see whether the same module weights predict acceptance.","The survey's geographic imbalance means the priority ordering may shift with a more global sample; a direct test would re-run the ranking questions with balanced representation from North America, Japan, and India and compare rank orderings.","Because the framework is conceptual, its next test is computational instantiation; one concrete check is whether the spatial-temporal memory module improves long-horizon merging or left-turn decisions in simulation relative to a memory-free baseline.","The framework's bidirectional adaptation implies AV deployment is a co-evolution process: as human drivers learn to exploit or trust AVs, AV policies should update in response, which would make field deployment a continuous calibration loop rather than a one-time release."],"forward_implications":["If the framework's priority ordering is right, near-term SCAV development should concentrate on anticipation capability, the ability to read other road users' intended actions, before investing heavily in other social features.","Long-term research and development should shift toward bidirectional behavioural adaptation and spatial-temporal memory, which experts ranked as the most critical capabilities for the 5-10 year horizon.","Safety should be enforced by a dedicated module that sits outside the socially compliant decision-making layer, meaning social behaviour is constrained but not determined by safety logic.","Social compliance should be treated as a dynamic trade-off between ego-vehicle benefits and network-level outcomes, so optimal individual behaviour is not assumed to be optimal for the road network.","The five methodological families identified in the scoping review, including imitation learning, reinforcement learning with utility models, game-theoretic and field-based models, trajectory prediction, and optimisation-based tuning, are complementary and should be combined in future SCAV systems."],"supporting_citations":[{"why":"Supplies the definition of socially compliant driving and the social value orientation approach that the framework builds on.","marker":"Schwarting et al. (2019)"},{"why":"Provides the scoping-study methodology that structures the literature review.","marker":"Arksey and O'Malley (2005)"},{"why":"Supplies the PRISMA-ScR reporting extension that the five-step review process adapts.","marker":"Tricco et al. (2018)"},{"why":"Earlier review of social interactions in autonomous driving that this scoping review explicitly extends and differentiates itself from.","marker":"Wang et al. (2022)"},{"why":"Representative game-theoretic model for human-like decision making used in the reviewed methodological landscape.","marker":"Hang et al. (2021)"},{"why":"Risk-based driver model cited as evidence that human-like driving behaviour can emerge from model-based approaches.","marker":"Kolekar et al. (2020)"},{"why":"Survey evidence on selfish versus utilitarian collision algorithms used to argue that social perception and trust are part of SCAV development.","marker":"Joo and Kim (2023)"},{"why":"Defines the automation levels used to scope which vehicles count as socially compliant automated vehicles.","marker":"SAE International (2021)"}],"fun_headline_variants":["Five-module framework for socially compliant AVs validated by experts","New conceptual model makes AVs socially compliant in mixed traffic","Survey confirms five-part design for socially aware automated vehicles","AV social compliance: experts endorse five interdependent modules"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's generalisability rests on the assumption that the 90 survey respondents, drawn mostly from Europe and China with only one from the United States, represent the global range of driving cultures and expert priorities; if they do not, the framework's capability set and timing could be regionally biased.","fun_headline_variants_meta":{"raw":{"variants":["Five-module framework for socially compliant AVs validated by experts","New conceptual model makes AVs socially compliant in mixed traffic","Survey confirms five-part design for socially aware automated vehicles","AV social compliance: experts endorse five interdependent modules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000491,"raw_usage":{"total_tokens":2452,"prompt_tokens":1023,"completion_tokens":1429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":1364}},"tokens_in":639,"tokens_out":1429,"duration_ms":13051,"temperature":1.0,"reasoning_tokens":1364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:12.723993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a replication survey with balanced representation from North America, Japan, and India returned an average importance rating at or below the neutral midpoint for any of the nine capabilities, the paper's claim that all framework elements are significant would be contradicted. A sharper test would be a controlled driving experiment in which an AV implementing the full framework is compared with a conventional AV; if human drivers do not predict the framework AV's actions more accurately in merging and left-turn scenarios, the claim that these modules enable social compliance would fail.","supporting_citations":[],"review_version":1}