{"id":"14aa8fbd-1971-4438-94c2-791c9f47f4b6","arxiv_id":"2506.14268","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Dubai residents broadly accept robot avatars for service tasks, preferring physical robots over digital ones, clearly robotic or cartoonish designs over androids and animal-like forms, and informational tasks in commercial and transport settings.","lead":"Researchers surveyed 1,001 Dubai residents about whether they would accept robot and virtual customer-service avatars in an ideal future Dubai. The results map which appearances, settings, and tasks people accept most, and they show that acceptance varies strongly by culture and gender.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Acceptance items ask about an 'ideal society for Dubai,' so the headline 'high acceptance' may measure aspirational endorsement, not actual acceptance; the Discussion's limitations do not address this framing.","rationale":"The reader's condition focuses on sample representativeness, which is a legitimate external-validity concern. But the 'ideal society' framing threatens internal validity of the measured construct: the paper defines acceptance as the avatar being 'willingly incorporated into society,' yet the items ask whether avatars would be permitted in an imagined ideal future. If that wording inflates agreement, no amount of improved sampling fixes the headline. The proposed A/B experiment would settle whether the wording matters. If ratings are insensitive to the framing, the concern is resolved and the reader's conditional verdict can stand. Secondary issues—English-only administration, non-probability panel recruitment, fragile cluster labels, and missing raw data—remain important but are less load-bearing for the central descriptive claim than the possibility that the instrument measures idealized desirability rather than acceptance.","tokens_in":13712,"tokens_out":8225,"duration_ms":95193,"concrete_test":"Run a randomized A/B wording experiment with the same panel and stimuli: Condition A repeats the original 'ideal society / would be permitted' wording; Condition B asks 'Do you personally support deploying robot avatars for customer service in Dubai today?' with identical response scales and item order. Compare the six appearance rates, modality rates, and top/bottom settings between conditions; if any headline rate drops by more than 10 percentage points or the appearance/setting ordering changes, the central claim should be reframed as 'ideal-society endorsement' rather than acceptance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing premise is that the survey items measure acceptance rather than idealized desirability. Throughout Sections 1.1-1.3 and 2.1, participants are asked to 'envision an ideal society for Dubai' and then to say whether robot avatars 'would be permitted' or 'would be used' for each appearance, setting, and task. This wording removes exactly the real-world frictions—cost, trust, safety, job displacement, unfamiliarity—that define whether a technology is accepted. A person can easily agree that robots would exist in an ideal future while opposing their actual deployment today. The Discussion's limitation paragraph acknowledges the survey was 'hypothetical,' but it does not acknowledge the 'ideal society' frame, so the stated limitation does not cover the strongest threat. Because every headline rate (67.3% robot, 56.9% digital, appearance and setting rankings) is produced by this same frame, the central descriptive claim may be an artifact of question wording even within the surveyed sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a large-scale survey (N = 1001) of Dubai residents on acceptance of cybernetic avatars in customer-service roles. Participants rated acceptance of physical robot avatars versus digital avatars, six robot appearance categories, twenty deployment settings, and thirteen tasks. The main descriptive findings are that physical robots are accepted more than digital avatars (67.3% vs 56.9%), highly anthropomorphic robotic and cartoonish appearances are favored over android, hybrid, low-anthropomorphic, and animal-like appearances, and acceptance is higher in commercial, transport, and cultural settings and for informational tasks than in healthcare, education, and emotionally sensitive roles. The paper also reports community-cluster and gender comparisons and a thematic analysis of open-ended reasons for android acceptance and rejection.","tokens_in":13872,"tokens_out":6914,"duration_ms":69977,"significance":"Taken at face value, the study provides one of the largest datasets on cybernetic-avatar acceptance in a multicultural urban setting, with transparent reporting of quota sampling, attention checks, multiple-testing corrections in the main analyses, and supplementary tables. Its direct policy link to Dubai's deployment strategy makes the descriptive rankings practically useful. However, the headline claims are currently stronger than the measurement and sampling design support: the 'ideal society' question frame and the non-probability, English-only sample place important limits on what can be concluded about actual acceptance among Dubai residents. Because these issues bear directly on the central conclusion, the contribution is promising but needs revision.","major_comments":[{"comment":"The acceptance items ask participants to 'envision an ideal society for Dubai' and then to state whether robot avatars 'would be permitted' or 'would be used' in that ideal scenario. This wording elicits idealized desirability rather than acceptance under real-world constraints such as cost, trust, safety, job displacement, and unfamiliarity, and it applies to every headline rate in Sections 3.1–3.4. The Discussion's limitation paragraph acknowledges that the survey was 'hypothetical,' but it does not address the 'ideal society' framing specifically, which is the strongest threat to construct validity. I recommend either reframing all conclusions as 'stated acceptance in an idealized future scenario' or adding items that present realistic trade-offs, and discussing how the framing may inflate agreement rates.","section":"§1.1–1.3, §2.1, §4 (Limitations)"},{"comment":"The sample is a non-probability quota sample recruited through an external panel and administered only in English. The authors explicitly state in §2.3.1 that the goal was balanced representation rather than statistical representativeness, yet the abstract and Discussion generalize to 'public in Dubai' and 'Dubai residents.' English-only administration is likely to exclude a substantial share of Dubai's non-English-speaking residents, and the cluster definitions are heterogeneous (e.g., the 'Western' cluster includes Peru, Cuba, Panama, and the Dominican Republic). The population-level claims should be softened, and the limitations section should explicitly discuss language, panel-recruitment, and cluster-definition biases.","section":"§2.2, §2.3.1, Table 1"},{"comment":"The binomial tests in Section 3.2.1 compare each community cluster's agreement rate to the corresponding agreement rate in the overall sample, but the overall sample includes the cluster being tested. This violates the independence assumption of the exact binomial test and biases the comparison, since the baseline proportion already contains the cluster's own respondents. For example, the Emirati android agreement rate (69.4%) is tested against 50.4%, a proportion that already includes the 170 Emirati participants. The cluster-level claims, including the abstract's statement about Emirati and 'Other Asia' differences, should be reanalyzed with the target cluster excluded from the baseline proportion, or with an omnibus test such as chi-square followed by post-hoc comparisons. The same issue applies to the modality comparisons in Section 3.1.","section":"§3.2.1, Tables 3A–8A; §3.1, Tables 1A–2A"},{"comment":"Bonferroni correction is applied separately within each appearance type (six cluster comparisons per appearance, α = .0083), but not across the six appearance types. With 36 cluster-by-appearance binomial tests, the expected number of false positives under the null is close to 1.8 at the per-test level; the significant findings for hybrid android and cartoonish appearances should be assessed with a family-wise correction across all appearance tests, or the results should be labeled as exploratory.","section":"§3.2.1"}],"minor_comments":[{"comment":"The word 'Cyberne3c' appears in the title and in the abstract heading; this should be corrected to 'Cybernetic.'","section":"Title and abstract heading"},{"comment":"The sentence 'Responses were recoded into a binary outcome (“agree” = 0 vs. “not agree” = 1' appears inconsistent with the later statement that 'Agree' responses were treated as successes; please clarify the dummy coding.","section":"§3.2.1"},{"comment":"The text says 'Table 6 presents the percentage of respondents who agreed...' but the relevant table is numbered Table 3; the cross-reference should be corrected.","section":"§3.4"},{"comment":"Wilcoxon and Mann-Whitney tests are reported with p-values but no effect sizes or confidence intervals; given the large sample, reporting rank-biserial correlation or Cliff's delta would help readers assess the magnitude of the differences.","section":"§3.1, §3.2.1"},{"comment":"The full survey instrument is not included in the paper or supplementary information; providing the complete questionnaire would improve transparency and reproducibility.","section":"§2.1 and Appendix/SI"},{"comment":"The thematic analysis of open-ended responses reports no reliability statistics, coding scheme, or intercoder agreement; the authors should describe how many responses fell into each theme and how the AI-assisted grouping was validated.","section":"§3.2.2"},{"comment":"The Discussion describes the sample as capturing 'residents of over 80 countries,' but the analysis is based on six broad community clusters rather than country-level representation; please rephrase or clarify the level at which countries are represented.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The 'ideal society' concern is well-founded and should be the first issue the authors address, because it affects all headline acceptance rates. The self-inclusive binomial tests are a more straightforward statistical fix and also affect the abstract's cluster claims. I see no reason to suspect bad faith; the issues are internal-validity and generalizability problems that are fixable in revision. The dataset and descriptive rankings would be a useful contribution if these load-bearing concerns are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a decent, useful paper. It reports the largest survey I know of on cybernetic avatar acceptance in a genuinely multicultural city, with over a thousand Dubai residents from six community clusters. The descriptive results are clear and the analysis is appropriate for what it is: Wilcoxon, Mann-Whitney, binomial tests with Bonferroni corrections. The main findings—robot avatars preferred over digital, highly anthropomorphic robotic and cartoonish forms beating androids and animal-like designs, high support in malls and airports, low support in healthcare—are new for this population and presented with enough detail to be actionable.\n\nThe strongest part is the dataset itself. The survey has attention checks, a quota balanced on gender and community, and a transparent breakdown of exclusions. The open-ended analysis on android acceptance is a nice addition, with themes like comfort, alignment with Dubai's innovation agenda, job displacement, and religious concerns. That is real value.\n\nThe soft spots are both about interpretation, not the raw numbers. First, the stress-test is right: participants were asked to envision an 'ideal society for Dubai' before saying whether robot avatars 'would be permitted' or 'would be used.' That wording invites aspirational agreement. The Discussion acknowledges the survey was 'hypothetical' but never mentions the ideal-society frame, so the stated limitation doesn't cover the strongest threat. This doesn't sink the paper—people can still express a preference under that prompt—but it means the headline 'high acceptance' should be read as 'endorsement in an idealized scenario,' not 'acceptance of actual deployment today.'\n\nSecond, the reader's concern about generalizability is legitimate. The sample is quota-based and English-only, and the 'Western' cluster lumps Peru with Canada. The authors are upfront that they didn't aim for a probability sample, and the balanced design is defensible for cross-group comparison. But the title says 'evidence from Dubai' and the Discussion implies city-wide guidance, so the framing overreaches. That's fixable by softening the language and making the limits explicit. Also, no raw data or stimuli are provided, which hampers replication.\n\nThis paper is for applied HRI and policy researchers, not for someone building theory. It deserves a serious referee. I'd accept it conditional on revision—add a paragraph on the ideal-society framing and recalibrate the generalizability claims. Put it on the reading group list.","headline":"A solid, useful descriptive survey of avatar acceptance in Dubai with a clean dataset; the main caveats are the 'ideal society' framing and the non-representative sample, both manageable in revision.","tokens_in":14389,"tokens_out":2245,"would_cite":false,"duration_ms":25833,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dubai residents broadly accept robotic customer-service avatars, with clear preferences for physical robots, clearly robotic or cartoonish designs, and commercial spaces.","keywords":["cybernetic avatars","robot avatars","technology acceptance","multicultural contexts","social robots","human-robot interaction","service sector","Dubai"],"falsifier":"Run the same acceptance questions on a probability sample drawn from Dubai's official resident records, administered in Arabic, Hindi, Urdu, Tagalog, and English, and check whether the headline figures replicate: overall 67.3% agreement with physical robot avatars, 61.1% agreement with highly anthropomorphic robotic appearances, and the Emirati cluster's 69.4% agreement with androids; failure to reproduce these percentages within sampling error would undermine the paper's central claim.","tokens_in":13540,"feed_emoji":"🤖","tokens_out":7614,"duration_ms":75092,"temperature":0.7,"pith_summary":"This paper reports a survey of 1,001 residents of Dubai measuring public acceptance of cybernetic avatars, defined as robots or digital characters that combine autonomous abilities with remote human control. Its core finding is that acceptance is generally high but conditional: physical robots are preferred over digital avatars, distinctly anthropomorphic robotic appearances and cartoonish designs beat androids and animal forms, and people accept avatars for informational and logistical tasks in commercial, transport, and cultural spaces far more than in healthcare, schools, or emotionally sensitive roles. The study also claims that cultural background shapes these preferences, with Emirati respondents unusually accepting of androids and respondents from the Other Asia cluster unusually fond of cartoonish designs. If these results hold, they give service providers and city governments a concrete map of where and in what form avatar customer service will be welcomed, and where it will meet resistance.","feed_headline":"Dubai survey: robot avatars accepted, with look-based limits","feed_subtitle":"Survey of 1,001 residents: physical robots beat digital avatars; malls and airports welcome them, healthcare and schools do not.","key_machinery":"The central instrument is a survey battery built around a six-category appearance typology—ultra-realistic android, hybrid android, highly anthropomorphic robotic-looking, low anthropomorphic robotic-looking, cartoonish, and animal-looking—each illustrated by two visual stimuli and rated on a three-point agree-neutral-disagree scale. The same battery asks about two modalities (physical robot versus digital avatar), twenty service settings, and thirteen tasks, plus attitude and fear scales. The comparisons are carried by a stratified quota sample of 1,001 participants balanced by gender and by six community clusters (Emirati, Middle East, South Asia, Other Asia, Western, Other Africa), which is what allows the paper to attribute differences such as Emirati android acceptance or Other Asia cartoonish acceptance to cultural background rather than to chance sampling.","core_discovery":"On the paper's own terms, the central claim is that public acceptance of cybernetic avatars in Dubai is real but patterned. Overall, 67.3% of the 1,001 respondents agreed that physical robots should serve customers, versus 56.9% for digital avatars. Of six appearance categories, highly anthropomorphic robotic-looking designs drew the most agreement (61.1%), followed by cartoonish (53.3%) and android (50.4%) designs; animal-like forms drew the least (39.0%) and the most disagreement (28.0%). Acceptance was highest in shopping malls (74.5%), airports (69.6%), museums (69.1%), and metro stations (68.9%), and lowest in healthcare settings (31.6 to 37.4%) and schools (39.4%). Task preferences followed the same logic: information, guidance, recycling collection, and carrying items were widely accepted, while handling complaints (43.2%) and companionship (50.2%) were not. The paper further claims that community identity shifts these preferences: Emirati respondents agreed with android avatars at 69.4% versus the overall 50.4%, while the Other Asia cluster accepted cartoonish avatars at 68.4% versus the overall 53.3% and accepted androids at only 33.3%.","pith_inferences":["The paper leaves implicit that stated acceptance of still images may not predict actual behavior with a moving, interacting robot; a field trial in a Dubai shopping mall comparing stated acceptance with observed willingness to approach and use a physical avatar would test that gap.","Because the 'Western' cluster groups Latin American countries such as Peru, Cuba, Panama, and the Dominican Republic with European and North American countries, its 40.4% cartoonish-acceptance figure should not be read as a unified Western cultural response; a country-level re-analysis could change the cluster conclusions.","The same relative ordering—clearly robotic and cartoonish forms over androids and animals, commercial over healthcare settings, informational over emotionally sensitive tasks—may extend from cybernetic avatars to conventional service robots, since the authors note the relevance to social and humanoid robotics.","The open-ended android responses suggest that acceptance is tied to Dubai's innovation-brand identity while rejection is tied to uncanny-valley discomfort, job displacement, and religious or ethical objections; these mechanisms could be measured directly as predictors in a follow-up survey."],"forward_implications":["Robotic customer-service avatars deployed in Dubai's commercial, transport, and cultural venues—shopping malls, airports, museums, and metro stations—are likely to meet broad public approval, while hospitals, clinics, schools, and nursing homes will face majority resistance.","Highly anthropomorphic robotic-looking designs are the safest default appearance, followed by cartoonish designs, while animal-like and hybrid-android forms are more likely to be rejected.","Physical robot avatars should be prioritized over screen-based or VR avatars when the goal is to maximize acceptance.","Appearance preferences vary by community, so a single uniform avatar design may be less accepted than designs tailored to the demographic profile of each deployment site.","Task assignment matters as much as appearance: information, guidance, multilingual support, and object carrying are welcome, whereas complaint handling and companionship roles are not."],"supporting_citations":[{"why":"Defines cybernetic avatars as hybrid autonomous-and-teleoperated systems, fixing the object the survey measures.","marker":"Horikawa et al., 2023"},{"why":"Supplies the definition of acceptance as willing incorporation of a robot into society, which the survey operationalizes.","marker":"Broadbent et al., 2009"},{"why":"Supplies the Erica android image used as the ultra-realistic android stimulus in the appearance ratings.","marker":"Glas et al., 2016"},{"why":"Supplies the Ibuki childlike android image used as the second android stimulus in the appearance ratings.","marker":"Nakata et al., 2022"},{"why":"Provides the lists of social-robot appearances, service settings, and tasks from which the survey response categories were built.","marker":"Aymerich-Franch and Ferrer, 2020, 2023"},{"why":"Supplies the attitude-toward-robots and fear-of-robots scale items used to contextualize the acceptance measures.","marker":"Aymerich-Franch and Gómez, 2024"},{"why":"Gives Dubai's official population structure, used to justify the multi-community quota sampling strategy.","marker":"Dubai Statistics Center, 2024"}],"fun_headline_variants":["Dubai prefers robot customer service, with look and venue limits","Robotic-looking avatars lead acceptance in Dubai; animals lag","Emiratis favor androids, Other Asia prefers cartoonish in Dubai","Dubai survey: 67% accept robot staff, but looks and settings matter","Physical robot avatars beat digital ones in Dubai's multicultural test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the quota sample drawn by an external panel in a single language (English) and grouped into six broad community clusters actually represents the attitudes of Dubai's population; if the sample misses Arabic-only or other non-English-speaking residents, or if the clusters lump unlike communities together, the overall acceptance rates and cluster comparisons do not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Dubai prefers robot customer service, with look and venue limits","Robotic-looking avatars lead acceptance in Dubai; animals lag","Emiratis favor androids, Other Asia prefers cartoonish in Dubai","Dubai survey: 67% accept robot staff, but looks and settings matter","Physical robot avatars beat digital ones in Dubai's multicultural test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000978,"raw_usage":{"total_tokens":4241,"prompt_tokens":1117,"completion_tokens":3124,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":3032}},"tokens_in":733,"tokens_out":3124,"duration_ms":25350,"temperature":1.0,"reasoning_tokens":3032,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:17:36.775925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same acceptance questions on a probability sample drawn from Dubai's official resident records, administered in Arabic, Hindi, Urdu, Tagalog, and English, and check whether the headline figures replicate: overall 67.3% agreement with physical robot avatars, 61.1% agreement with highly anthropomorphic robotic appearances, and the Emirati cluster's 69.4% agreement with androids; failure to reproduce these percentages within sampling error would undermine the paper's central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines cybernetic avatars as hybrid autonomous-and-teleoperated systems, fixing the object the survey measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the lists of social-robot appearances, service settings, and tasks from which the survey response categories were built."}],"review_version":1}