{"id":"af09cc04-714f-4061-a8d6-9cbc6647b3b8","arxiv_id":"2504.18044","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A cross-country survey and interview study finds that ChatGPT users and experts perceive transparency, bias, and data collection as the main ethical concerns, and that perceptions differ across groups for trust, security, toxicity, and social norms.","lead":"This paper surveys 111 ChatGPT users and interviews 38 experts in Germany, Iran, and the US about perceived ethics and social norms, covering bias, trustworthiness, security, toxicity, social norms, and ethical data. It reports that transparency, bias, and data collection are the most frequent ethical concerns, though it measures perceptions rather than actual ChatGPT behavior.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3's Kruskal-Wallis grouping variable 'Ethic' is undefined; without it, the quantitative evidence for significant cross-group differences in ethics perceptions is uninterpretable.","rationale":"The reader's weakest_assumption identifies exactly this gap, and I agree. The paper explicitly presents the Kruskal-Wallis result as the quantitative confirmation that ethical perceptions differ across groups (Section 4.1.2, Discussion, and Conclusion). A test whose grouping variable is never defined is not a minor reporting omission: with df=4 but only three countries, the five 'Ethic' groups cannot be recovered from the demographic table, and if the data are repeated measures, the test is misapplied. Since no raw data are released, there is no way to referee the table. This supports the reader's CONDITIONAL verdict. I would not move to REJECT because the qualitative evidence and item-level percentages may still support a weaker descriptive claim about perceived ethical concerns; the paper's possible contribution is a perception taxonomy, not the specific H statistics. But the quantitative component must be defined and corrected, or explicitly downgraded to descriptive percentages.","tokens_in":30388,"tokens_out":4516,"duration_ms":48961,"concrete_test":"Obtain the anonymized data or codebook and identify the five levels of 'Ethic' used in Table 3. If they are independent participant subgroups, reproduce the Kruskal-Wallis tests and verify the group assignments. If they are the same participants' responses across five ethical categories, rerun the analysis with the Friedman test; if Trustworthiness and Social Norms lose significance, Section 4.1.2's claim of cross-group variation is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative claim of statistically significant perception differences rests entirely on Section 4.1.2 and Table 3. Table 3 reports Kruskal-Wallis H statistics with df=4 and 'Grouping Variable: Ethic', but the manuscript never defines this variable, the five groups, how participants were assigned to them, or which pairwise comparisons were made. The Kruskal-Wallis test requires independent samples; if 'Ethic' refers to the five ethical categories (Bias, Trustworthiness, Security, Toxicology, Social Norms) measured on the same 111 participants, the independence assumption is violated and the correct repeated-measures test is Friedman's. In that case the reported p-values (e.g., Trustworthiness H=31.243, p=.000; Social Norms H=19.037, p=.001) cannot support Section 4.1.2's conclusion that 'participant perceptions vary significantly across different groups.' Since the abstract and Discussion cite these significant differences as evidence supporting the taxonomy, the quantitative pillar of the mixed-method claim is unverifiable until the grouping is specified or the test is corrected. The absence of Ethical Data from Table 3, despite the text saying six categories, compounds the ambiguity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-method study of user and expert perceptions of AI ethics and social norms in ChatGPT. The quantitative component is an online Likert-scale survey of 111 participants in Germany, Iran, and the US, organized around six categories: bias, trustworthiness, security, toxicology, social norms, and ethical data. The qualitative component consists of semi-structured interviews with 38 experts from the same three countries. The authors claim that their quantitative results, analyzed with Kruskal-Wallis tests, show significant differences across ethical categories, and that the interview analysis supports a six-category taxonomy of ethical concerns, with transparency and bias in unsupervised data collection identified as major issues. The paper presents a questionnaire in the appendix, a thematic analysis of expert interviews, and a proposed framework for evaluating LLM ethics.","tokens_in":30599,"tokens_out":6548,"duration_ms":65228,"significance":"If the findings are valid, the study would provide a useful cross-country, perception-based taxonomy of ChatGPT ethics and a set of user- and expert-identified concerns, including transparency, bias, data collection, and social norms. The qualitative corpus—38 experts interviewed in three languages and three countries—is a valuable resource, and the inclusion of the full questionnaire in the appendix is a reproducible feature. The mixed-method design is appropriate in principle. However, the abstract overstates the evidence by claiming the study evaluates whether ChatGPT itself \"operates following ethics,\" when only perceptions were measured. More importantly, the only inferential statistical analysis, the Kruskal-Wallis tests in Table 3, is uninterpretable as reported because the grouping variable is undefined and the independence assumption is not established. These issues are load-bearing because the abstract and Discussion rely on the significant p-values to support the taxonomy claim.","major_comments":[{"comment":"The Kruskal-Wallis analysis is the only quantitative inferential evidence for the claim that perceptions vary across ethical categories, but the grouping variable \"Ethic\" is never defined. The text says the test assessed differences among \"five independently sampled groups,\" yet no group membership, sample sizes, or sampling procedure are reported. If the five groups are the ethical categories measured on the same 111 participants, the independence assumption is violated and the correct test would be Friedman's; if the groups are something else (e.g., countries or user subgroups), the label \"Ethic\" and df=4 do not match the reported design. Consequently, the p-values (e.g., Trustworthiness H=31.243, p=.000; Social Norms H=19.037, p=.001) cannot be checked, and the claims in the abstract and Section 5.1 that significant differences were found are not supported as reported. The authors must specify the grouping, justify independence, report group sizes, and either use an appropriate test or remove the inferential claim. Relatedly, Table 3 omits the Ethical Data category even though the text claims it presents statistics for each of the six categories.","section":"Section 4.1.2, Table 3"},{"comment":"The abstract states that the study \"aims to evaluate whether ChatGPT in an empirical context operates following ethics and social norms,\" and the conclusion describes obstacles \"identified as ChatGPT's ethical concerns.\" The data, however, are self-reported Likert-scale opinions from 111 participants and semi-structured expert interviews; the study does not directly probe ChatGPT's outputs or behavior. The title's \"From What to How\" and the Discussion's claims about ChatGPT's capabilities therefore overreach what the design can show. These findings should be framed as perceptions, concerns, or reported experiences of users and experts, not as direct evidence about whether ChatGPT itself follows ethics. This is a load-bearing distinction because the paper's central contribution is presented as an evaluation of ChatGPT's ethical operation rather than a study of user and expert perceptions.","section":"Abstract and Section 5.1"},{"comment":"The qualitative analysis is described as \"guided by both deductive and inductive reasoning,\" but Figure 2 lists principal themes (Generalization, Challenges, Social Norms, Toxicology, Trustworthiness, Bias, Security) that map closely onto the survey's six categories. It is not explained how the deductive coding into the survey categories interacted with the inductive coding, whether the codebook allowed new categories to emerge, or how disagreements between the two coders were resolved. Without this information, the qualitative results risk simply confirming the authors' own framework rather than providing independent evidence for the taxonomy. The authors should report the coding protocol, the distribution of codes across themes, and any inductive themes that arose outside the original six categories; at minimum, an intercoder reliability statistic or a description of the adjudication process would strengthen the validity claim.","section":"Sections 3.1.2, 3.3, and 4.2"}],"minor_comments":[{"comment":"The participant description is internally inconsistent: the text says \"60% females and 40% males\" and then reports \"56.76% female, 43.24% male,\" and it says 112 individuals completed the survey while the analysis uses 111. These figures should be reconciled.","section":"Section 3.2.1"},{"comment":"The phrase \"quantitative interview study\" appears in both places and appears to be a typo for \"qualitative interview study.\" Similarly, Section 3.3 says \"both the quantitative survey study and the quantitative interview study\" when the latter is qualitative.","section":"Sections 3.1.2 and 3.3"},{"comment":"The reporting of the transparency item is confusing: Q6 states that \"understanding how ChatGPT's results were generated is somewhat challenging,\" but the text says 12.9% agreed with the ease of comprehension and 65.1% \"indicated trust that the results . . . cannot be explained.\" The direction of the percentages appears to be reversed or the coding of the item is unclear; please clarify what the reported percentages represent.","section":"Section 4.1.1, Q6"},{"comment":"Several reference entries are incomplete or contain informal metadata, including [69], [72], and [82]. A full reference cleanup is needed before publication.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a valuable qualitative corpus and a clearly presented questionnaire, but the quantitative pillar is currently unverifiable because the Kruskal-Wallis grouping variable is undefined. The abstract also overstates the design's ability to evaluate ChatGPT's actual ethical behavior. I believe these issues are fixable within the scope of a revision, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is an honest perception study with a real methodological flaw in its inferential statistics. The new thing here is the dataset: 111 survey responses and 38 expert interviews across Iran, Germany, and the US, asking about ChatGPT's bias, trustworthiness, security, toxicology, social norms, and data ethics. That cross-country design is a legitimate extension of prior single-dimension work, and the interview excerpts give you concrete, quotable material on how users in different regulatory environments experience the same tool. If you work on applied AI ethics or CSCW, the descriptive percentages alone are a useful map of concerns, especially the transparency and data-collection worries.\n\nThe paper does not, however, support its own abstract. It claims to evaluate whether ChatGPT 'operates following ethics,' but the design only measures perceptions. That is a re-scoping problem, not fatal, but it needs to be fixed.\n\nThe bigger issue is Table 3. The Kruskal-Wallis tests require independent groups, and the 'Grouping Variable: Ethic' is never defined. Given the survey design, the five rows are almost certainly five ethical categories rated by the same 111 participants, which means the correct test is Friedman's (or a mixed model). If so, those p-values are invalid. The paper also omits Ethical Data from the table despite claiming six categories. This is a load-bearing flaw because the Discussion cites those significant differences as evidence for the taxonomy. Without clarification, the quantitative part is uninterpretable.\n\nThere are smaller problems: no data or codebook, a convenience sample recruited via social media, and some sloppy writing (e.g., the participant demographics seem internally inconsistent at one point). None of these are disqualifying on their own.\n\nNet: the qualitative analysis and descriptive statistics earn the paper a serious look, but the inferential statistics need to be corrected and the claims re-scoped before I'd trust the conclusions. I'd send it to peer review—a good referee can force the authors to fix the test and release the instrument. I wouldn't cite it until then.","headline":"Useful cross-country perception data; the inferential statistics are currently uninterpretable due to an undefined grouping variable.","tokens_in":31114,"tokens_out":2229,"would_cite":false,"duration_ms":21895,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a mixed-method study of ChatGPT users and experts identifies six ethical dimensions of AI, with transparency and bias in data collection as the most salient concerns.","keywords":["AI ethics","social norms","ChatGPT","bias","trustworthiness","security","toxicology","ethical data"],"falsifier":"Re-analyze the survey with the Ethic grouping variable explicitly defined; if the five groups cannot be specified from the questionnaire or the identified p-values for Trustworthiness, Security, Toxicology, and Social Norms do not reproduce, the claim of significant cross-group differences collapses.","tokens_in":1,"feed_emoji":"⚖️","tokens_out":5960,"duration_ms":115358,"temperature":0.7,"pith_summary":"This paper tries to establish that ordinary users and AI experts perceive ChatGPT as ethically uneven, and that the most widely felt problems are a lack of transparency and bias embedded in the way its training data are collected. The evidence is a mixed-method study: a Likert-scale survey of 111 ChatGPT users in Germany, Iran, and the United States, plus semi-structured interviews with 38 experts. The authors organize both strands around six dimensions of AI ethics—bias, trustworthiness, security, toxicology, social norms, and ethical data—and report that most respondents found ChatGPT's output hard to explain and were unsure whether its data collection is ethical. If this picture holds, then making LLMs acceptable for everyday work, healthcare, and collaborative settings requires changing how data are gathered and how model behavior is explained, not just adding safety filters. The quantitative analysis also claims significant differences across participant groups for trustworthiness, security, toxicology, and social norms, while bias perceptions did not differ significantly.","feed_headline":"Survey finds ChatGPT ethics concerns cluster on transparency and bias","feed_subtitle":"A six-dimension ethics taxonomy from 111 users and 38 experts points to transparency and data-collection bias as the main gaps.","key_machinery":"The central object is a taxonomy of six AI-ethics dimensions—bias, trustworthiness, security, toxicology, social norms, and ethical data—with fourteen sub-dimensions and twenty-two Likert-scale questions. The argument is carried by triangulating a quantitative survey with a qualitative interview study, and the quantitative inference that perceptions vary across groups rests on the Kruskal-Wallis $H$ test, a non-parametric rank test for differences among independent groups. The machinery also includes thematic analysis of 38 expert interviews, coded by two researchers, which supplies the reasons behind the survey numbers.","core_discovery":"The paper's central claim is that a six-part taxonomy captures the ethical and social-norm concerns people actually have about ChatGPT, and that within that taxonomy the dominant perceived failures are transparency and bias from unsupervised data collection. On the survey, a majority of participants disagreed that an outside observer can understand how ChatGPT's results are produced, and a large share expressed uncertainty about whether the data behind ChatGPT were collected ethically. The Kruskal-Wallis $H$ test is then used to claim that perceptions differ significantly across groups for trustworthiness, security, toxicology, and social norms, but not for bias, which the authors take as evidence that ethics judgments are context-dependent and require qualitative follow-up. From the expert interviews, the paper further claims that bias is experienced differently by region—users in Iran emphasize data and access limitations, while users in the US and Germany emphasize gender and race—and that trust hinges on transparency, reliability, and open data practices.","pith_inferences":["The same survey instrument could be applied to other LLMs to test whether the transparency-and-bias gap is ChatGPT-specific or common to all large language models, a question the paper raises but does not answer.","The reported country differences imply that ethics benchmarks for chatbots should be validated per culture rather than once globally, for example by building country-specific bias test sets.","A behavioral extension would ask participants to identify biased or non-transparent ChatGPT outputs in a controlled prompt set, connecting self-reported perceptions to measurable system behavior."],"forward_implications":["If the taxonomy and findings are correct, transparency about training-data collection and output generation is the first thing to fix in LLM-based tools.","Because trustworthiness, security, toxicology, and social norms showed significant cross-group differences, ethics guidelines that are uniform across countries and user groups will miss real variation in perception.","The lack of a significant difference for bias means bias may be a constant concern across groups, so it needs different detection methods than the other dimensions.","The regional pattern in the interviews suggests that mitigation should be localized, for example by including non-Western training data and addressing access restrictions.","The six-dimension structure gives subsequent studies a ready-made questionnaire for evaluating ChatGPT and other LLMs."],"supporting_citations":[{"why":"Supplies the six-category ethical-risk taxonomy that the paper adapts into its bias, trustworthiness, security, toxicology, social norms, and ethical data dimensions.","marker":"[101]"},{"why":"Prior diagnostic analysis of ChatGPT's ethical risks from bias, reliability, robustness, and toxicity that the study builds on and extends empirically.","marker":"[109]"},{"why":"Defines the Kruskal-Wallis H test used for the quantitative significance claims in Table 3.","marker":"[62]"},{"why":"Provides the social-norms and fairness framework that shapes the social norms category and interview questions.","marker":"[16]"},{"why":"Source of the survey question about whether ChatGPT is biased against conservatives.","marker":"[61]"},{"why":"Supporting evidence for ChatGPT's poor hate-speech and counter-speech performance, cited in the toxicology discussion.","marker":"[17]"},{"why":"Establishes the globally convergent AI ethics principles of transparency, justice, and fairness that the discussion uses to interpret the findings.","marker":"[48]"}],"fun_headline_variants":["Transparency and bias lead ChatGPT ethics concerns","ChatGPT ethics survey: top gaps are transparency and bias","Users and experts flag transparency and bias in ChatGPT","Six AI ethics dimensions, but transparency and bias dominate","Study: ChatGPT's ethical weak spots are transparency and bias"],"cache_read_input_tokens":33408,"weakest_assumption_plain":"The quantitative finding that perceptions differ across groups depends on the five levels of a grouping variable called Ethic in the Kruskal-Wallis test, but the paper does not say what those five groups are, how participants were sorted into them, or whether the samples are independent.","fun_headline_variants_meta":{"raw":{"variants":["Transparency and bias lead ChatGPT ethics concerns","ChatGPT ethics survey: top gaps are transparency and bias","Users and experts flag transparency and bias in ChatGPT","Six AI ethics dimensions, but transparency and bias dominate","Study: ChatGPT's ethical weak spots are transparency and bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1302,"prompt_tokens":866,"completion_tokens":436,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":361}},"tokens_in":482,"tokens_out":436,"duration_ms":4227,"temperature":1.0,"reasoning_tokens":361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:25:28.414580+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the survey with the Ethic grouping variable explicitly defined; if the five groups cannot be specified from the questionnaire or the identified p-values for Trustworthiness, Security, Toxicology, and Social Norms do not reproduce, the claim of significant cross-group differences collapses.","supporting_citations":[],"review_version":1}