{"id":"3ede30cb-eca3-4b51-9c48-f3a98a1abaf3","arxiv_id":"2502.00632","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 168 AI-use declarations from 2024 Elsevier articles, ChatGPT is the dominant tool and readability improvement is the top stated purpose, with reported differences between author groups.","lead":"This paper analyzed 168 statements in which authors of Elsevier journal articles disclosed their use of AI writing tools. It found ChatGPT dominates, most authors use AI for readability and grammar, and author team composition correlates with how AI is used.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's quantitative core is internally inconsistent: abstract and full text report different sample sizes and contradictory p-values for the same hypotheses, so the central claims are not defined as posted.","rationale":"The reader's weakest_assumption (declarations may be incomplete or inaccurate) is a real threat to construct validity and is conceded in Section 7. My primary concern is more immediate: regardless of declaration accuracy, the paper as posted reports contradictory sample sizes and p-values for the same hypotheses in the abstract and full text. This is a concrete, checkable correctness failure that independently supports rejection. I agree with the reader's overall verdict and recommendation that the data and coding materials must be released and reconciled before the quantitative claims can be considered reliable. I mark agreement as partial because the reader located the weakest link in the reliability of disclosures, whereas I locate it in the unreproducibility and internal inconsistency of the reported statistics; both point to the same conclusion.","tokens_in":12544,"tokens_out":5438,"duration_ms":61111,"concrete_test":"Recompute the Fisher-Freeman-Halton exact test from the 2x6 contingency tables in Table 6, Panel A and Panel B, using R's fisher.test (or equivalent) and compare the computed p-values with the abstract (0.2359, 0.0008) and full text (0.0483, 0.0012). If the recomputed values match neither pair, the association claims fail outright; if they match one pair, the other version is erroneous and must be corrected. Also check whether the declaration counts total 135 or 168 by auditing the extracted declarations once the dataset is released.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims are not stable within the paper. The abstract reports 135 AI declarations from 8,633 articles, 73.3% ChatGPT usage, and Fisher-Freeman-Halton p-values of 0.2359 (native-speaker status) and 0.0008 (team composition). The full text reports 168 declarations from 8,859 articles, 77% ChatGPT usage, and p-values of 0.0483 and 0.0012 for the same hypotheses. Table 6 provides enough information to recompute the exact tests, so this is a checkable internal contradiction, not a matter of interpretation. The data availability section only offers data 'on reasonable request' with no code or release, so a reader cannot tell which analysis generated the abstract and which generated the body. Since hypotheses H1 and H2 drive the paper's conclusions about language background and team composition, the conflicting p-values are load-bearing: one of the two versions is wrong, and without raw data the published result is unverifiable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript analyzes Elsevier journal declarations of generative-AI use in academic writing, combining content analysis, Fisher-Freeman-Halton exact tests, and text mining to describe which AI tools authors declare, for what purposes, and whether tool-purpose patterns differ by native-language status and team composition. The full-text version reports 168 declarations from 8,859 articles, finds ChatGPT dominant (77% of usage), readability and grammar as the top declared purposes, and significant associations for both native-speaker status (p = 0.0483) and team composition (p = 0.0012). The abstract reports different numbers, including 135 declarations, 73.3% ChatGPT usage, a non-significant native-speaker result (p = 0.2359), and a team-composition p-value of 0.0008.","tokens_in":12701,"tokens_out":2841,"duration_ms":28363,"significance":"If the results were internally consistent, the study would provide a useful descriptive snapshot of declared AI use in a large publisher's journals and could inform editorial policy discussions. The research question is timely, and the use of disclosure statements as a data source is a plausible approach, even though it captures declarations rather than actual use. The paper's value is currently undermined because its central descriptive and inferential claims are not stable across the abstract and the full text, and the lack of a public data release prevents independent adjudication.","major_comments":[{"comment":"The sample size is inconsistent: the abstract reports 135 AI declarations from 8,633 articles, while Section 4.1 reports 168 declarations and the body text reports 8,859 articles (8,633 Elsevier plus 226 conference papers). All percentages and statistical tests depend on this sample, so the central quantitative claims are not defined as posted.","section":"Abstract vs. §4.1, §5.1, §5.2"},{"comment":"The native-speaker-status hypothesis (H1) yields contradictory results: the abstract reports no significant association (p = 0.2359), while Section 5.2 and Table 6 report a significant association (p = 0.0483). Since H1 is one of the paper's two main hypotheses, the conclusion about language background is internally inconsistent and cannot be accepted as stated.","section":"Abstract vs. §5.2 and Table 6, Panel A"},{"comment":"The team-composition hypothesis (H2) also differs between abstract (p = 0.0008) and full text (p = 0.0012). Although both values indicate significance at the 0.01 level, the discrepancy shows that the abstract and the full text are based on different computations or datasets, which undermines confidence in the reported exact test results.","section":"Abstract vs. §5.2 and Table 6, Panel B"},{"comment":"The paper states that data were collected from 27 Scopus major categories, but Table 1 lists only 26 rows. Additionally, 'Ultrasonics Sonochemistry' appears twice (under Chemical Engineering and under Physics and Astronomy), and the journal listed for Nursing ('Journal of Functional Foods') is not a nursing journal. These errors cast doubt on the accuracy of the journal-selection and data-collection description in Section 4.1.","section":"Table 1"},{"comment":"The distribution of declared purposes is inconsistent: the abstract reports readability at 57.8% and grammar checking at 19.3%, whereas Section 5.1 reports 51% and 22% for the same categories. Since these percentages are key descriptive results, the manuscript does not provide a single stable account of its own main findings.","section":"Abstract vs. §5.1"},{"comment":"The data availability section states only that datasets are available 'on reasonable request' and does not provide code or a data repository. Given the internal contradictions between the abstract and the full text, the absence of a public, verifiable dataset makes it impossible for readers to determine which analysis generated the reported results.","section":"Data availability statement"}],"minor_comments":[{"comment":"The interpretation of the word-frequency differences would benefit from a clear statement that the reported 'Difference' values are raw per-1,000-word differences without a statistical test, because the text implies a meaningful contrast without providing uncertainty measures.","section":"§5.3, Table 7"},{"comment":"The sentence 'Figure 2 shows that 117 authors use ChatGPT... accounting for 77% of total usage' would be clearer if it stated the denominator (all tool mentions, which includes multiple tools per author) and how the percentage was calculated.","section":"§5.1, Figure 2 caption"},{"comment":"The discussion cites 'the significant influence of team composition (p = 0.0012)' and 'language background (p = 0.0483)' using the full-text values; the abstract uses different values, and this inconsistency should be resolved before the paper can be considered publishable.","section":"§6, Discussion"},{"comment":"There are numerous typographical and stylistic errors, including the misspelling 'World-cloud Statement' in Figure 4, the inconsistent phrase 'bibliometric analysis.' in a reference, and the reference to 'W AME' with irregular spacing.","section":"Throughout"},{"comment":"The sentence 'Future research... focusing on evolution and current landscape' is incomplete and needs to be rephrased to clearly state the planned future work.","section":"§7, Conclusion"}],"recommendation":"reject","confidential_remarks":"The manuscript, as posted, contains irreconcilable differences between the abstract and the body on sample size, key percentages, and both hypothesis-test p-values. These are not presentation issues; they change the conclusions of the paper. The absence of a public dataset or code means the inconsistencies cannot be resolved by a reader. While the topic is relevant and a corrected version might be salvageable, the current manuscript does not meet the standard of a self-consistent, verifiable research report. If the authors resubmit with a reconciled dataset and provide the underlying data, the study could merit further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the underlying dataset is worth having, but the posted version contradicts itself on all the headline numbers. The abstract says 135 declarations from 8,633 articles, with p=0.2359 for native-speaker status and p=0.0008 for team composition. The body says 168 declarations from 8,859 articles, with p=0.0483 and p=0.0012. Those are not cosmetic differences—one version says H1 is significant, the other says it isn't. The paper can't be used as-is for anything quantitative.\n\nWhat's genuinely new: the author collected Elsevier's AI disclosure declarations for a set of 2024 open-access journals and coded them for tool and purpose, then tested associations with author language background and team composition. That's a real, publishable dataset—I haven't seen these declarations analyzed this way. The qualitative findings (ChatGPT dominance, readability and grammar as top purposes, non-native speakers using more grammar tools) align with common sense and prior anecdotal evidence. The Fisher-Freeman-Halton exact test is the right tool for small contingency tables, and Table 6 actually gives enough numbers to recompute the p-values, which is a good sign.\n\nThe soft spots are the internal inconsistencies, and they're load-bearing. Table 1 lists 26 categories, not the 27 claimed, duplicates Ultrasonics Sonochemistry under both Chemical Engineering and Physics, and labels the Nursing row with Journal of Functional Foods. The methodology says 8,633 articles were collected while the body abstract says 8,859. Data availability is only 'on reasonable request,' no code or release, so no one can check which set of numbers is right. And as the paper itself concedes in Section 7, declarations may be incomplete or inaccurate—so all the p-values really describe declaration behavior, not AI use, unless that bias is argued away.\n\nWho should read this: anyone working on journal disclosure policy or studying AI adoption in academic writing. The topic is timely and the descriptive snapshot is useful. But the quantitative claims need to be reconciled and the data released before this is citable.\n\nMy recommendation: not acceptable in current form. I'd want a revision that fixes the numbers, corrects Table 1, and posts the data and code. If the author does that, it's a solid empirical note. Given the originality of the dataset, I'd send it to a referee rather than desk-reject, but with a clear message that the internal contradictions must be resolved first.","headline":"Underlying dataset is new and worth having, but the paper's own numbers don't match across abstract and body, so it can't be used as posted.","tokens_in":13208,"tokens_out":3027,"would_cite":false,"duration_ms":29386,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-use declarations show ChatGPT dominates academic writing, with readability and grammar the leading purposes.","keywords":["academic writing","ChatGPT","AI usage declarations","journal policies","large language models","content analysis","Fisher-Freeman-Halton exact test","Elsevier journals"],"falsifier":"Take a random sample of the same 2024 Elsevier articles, run a validated AI-text classifier and have blinded human experts check for AI-assisted passages, then compare detected AI assistance against declared assistance; if undeclared AI use is common and correlates with team type, tool, or purpose, the reported distributions and p-values would shift.","tokens_in":12320,"feed_emoji":"🤖","tokens_out":8860,"duration_ms":79965,"temperature":0.7,"pith_summary":"This paper claims that the disclosure statements Elsevier journals require from authors provide a usable window onto how AI is actually used in academic writing. Analyzing 168 declarations collected from 8,859 articles and conference papers published in 2024 across 27 Scopus categories, the author reports that ChatGPT (including GPT-3.5 and GPT-4) dominates, accounting for 77% of declared tool mentions, and that the two most common declared purposes are improving readability (51%) and grammar checking (22%). The paper further claims an association between team composition and declared purpose: international teams lean more heavily on grammar assistance than single-country teams, with a Fisher-Freeman-Halton exact test yielding $p = 0.0012$. If these patterns hold, journal policies and AI-literacy programs can be targeted at the tasks authors actually use AI for, rather than treating all AI use as a single undifferentiated category.","feed_headline":"ChatGPT is 77% of declared AI tool use in journal papers","feed_subtitle":"Readability and grammar checking top the list, and international teams use AI differently from single-country teams.","key_machinery":"The machinery is Elsevier's standardized 'Declaration of Generative AI and AI-assisted technologies in the writing process' template, which asks authors to name their AI tool and state its intended purpose, supported by a three-part analytical workflow. First, content analysis uses a coding framework to classify tools and to sort purposes into nine categories such as readability, grammar, proofreading, translation, and content generation. Second, the Fisher-Freeman-Halton exact test, an extension of Fisher's exact test for contingency tables with small expected cell counts, tests whether native-speaker status or team composition is associated with purpose. Third, text mining—word frequencies, bigrams, and a bipartite tool-purpose network—cross-checks the coding and visualizes which tools are paired with which purposes. The load-bearing step is the coding of free-text declarations into the nine purpose categories, because every distribution and test result depends on that classification.","core_discovery":"The central discovery, as the author states it, is that AI tool use in academic writing is both concentrated and purpose-driven: one tool family, ChatGPT, dominates, and the declared purposes skew toward lower-level language tasks rather than higher-level content generation. The quantitative evidence is a set of frequency distributions from coded declarations—77% of tool mentions are ChatGPT, 51% of purposes are readability, and 22% are grammar—plus two Fisher-Freeman-Halton exact tests on small contingency tables. The team-composition test is presented as highly significant ($p = 0.0012$), with international teams showing a higher share of grammar use (30.4% vs 21.3%) and no declared proofreading or analysis uses; the native-speaker test is reported in the body as significant at $p = 0.0483$, with non-native speakers using grammar checking more and translation tools exclusively. The author interprets these patterns as evidence that AI tools help level language barriers in scholarly communication and that policies should distinguish language polishing from deeper content generation.","pith_inferences":["An editorial inference: the paper's aggregate percentages likely combine two selection effects—which journals require or encourage declarations and which authors choose to comply—so the 77% ChatGPT figure should be read as the share among declared users, not among all AI-assisted papers.","The paper's own limitation section notes that declarations may be incomplete or inaccurate; if under-reporting is more common among certain teams or purposes, the reported p-values describe declaration behavior rather than actual AI use.","A testable extension: run the same coding and tests on declarations from non-Elsevier publishers or on later years to see whether ChatGPT dominance and the readability/grammar focus are publisher-specific or a stable feature of AI-assisted academic writing."],"forward_implications":["Journal policies can stop treating AI use as one undifferentiated practice: since declared uses are mostly readability and grammar, tiered policies that permit language polishing while scrutinizing content generation would match observed behavior.","Because ChatGPT accounts for the large majority of declared usage, publisher guidance and detection efforts that focus on ChatGPT (across versions) would cover most of the current disclosure space.","International teams' higher reliance on grammar assistance suggests AI tools are serving as a language-equity mechanism; if that is true, restricting AI editing could disproportionately burden non-native-English-speaking researchers.","The significant team-composition association implies that usage patterns are not uniform across collaboration structures, so AI-literacy training and support should be tailored to team context rather than applied generically."],"supporting_citations":[{"why":"Documents that only 24% of top publishers provide generative-AI guidance, motivating why Elsevier's disclosure template is a rare and valuable data source.","marker":"Ganjavi et al. (2024)"},{"why":"Independent survey of EFL university students finding ChatGPT and Grammarly are the most used AI writing tools, corroborating the paper's dominance finding.","marker":"Selim (2024)"},{"why":"Controlled experiment showing ChatGPT improves the output quality of weaker writers, providing background for why readability and grammar are the top declared purposes.","marker":"Noy & Zhang (2023)"},{"why":"Establishes the linguistic disadvantage of scholars writing in English as an additional language, grounding the native-speaker hypothesis.","marker":"Flowerdew (2019)"},{"why":"Documents academic writing challenges faced by international graduate students, reinforcing the rationale for testing native-speaker status.","marker":"Singh (2015)"},{"why":"Shows international teams differ from domestic teams in media use for complex tasks, supporting the team-composition hypothesis.","marker":"Bjorvatn & Wald (2019)"},{"why":"Links cultural diversity and information and communication technology use in global virtual teams, used to justify why team composition might change AI tool purposes.","marker":"Shachaf (2008)"},{"why":"Discusses the difficulty of enforcing editorial policies on AI-generated papers, supporting the paper's concern that declarations may be incomplete or inaccurate.","marker":"Hu (2024)"}],"fun_headline_variants":["ChatGPT dominates AI writing tools in journals at 77%","Journal AI use: readability and grammar checking lead","International teams differ in AI grammar use in papers","AI in academic writing: 77% ChatGPT, mostly polishing","Cross-journal AI analysis: ChatGPT, readability, grammar"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the AI-use declarations researchers submit to journals are complete and accurate enough that their patterns reflect real AI use rather than only what authors chose to disclose.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT dominates AI writing tools in journals at 77%","Journal AI use: readability and grammar checking lead","International teams differ in AI grammar use in papers","AI in academic writing: 77% ChatGPT, mostly polishing","Cross-journal AI analysis: ChatGPT, readability, grammar"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1499,"prompt_tokens":908,"completion_tokens":591,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":512}},"tokens_in":524,"tokens_out":591,"duration_ms":5703,"temperature":1.0,"reasoning_tokens":512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:14:00.593305+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the same 2024 Elsevier articles, run a validated AI-text classifier and have blinded human experts check for AI-assisted passages, then compare detected AI assistance against declared assistance; if undeclared AI use is common and correlates with team type, tool, or purpose, the reported distributions and p-values would shift.","supporting_citations":[{"cited_title":", Eppler, M.B","cited_arxiv_id":null,"evidence_quote":"Documents that only 24% of top publishers provide generative-AI guidance, motivating why Elsevier's disclosure template is a rare and valuable data source."},{"cited_title":"APACrefauthors \\ 2024","cited_arxiv_id":null,"evidence_quote":"Independent survey of EFL university students finding ChatGPT and Grammarly are the most used AI writing tools, corroborating the paper's dominance finding."},{"cited_title":"APACrefauthors \\ 2019","cited_arxiv_id":null,"evidence_quote":"Establishes the linguistic disadvantage of scholars writing in English as an additional language, grounding the native-speaker hypothesis."},{"cited_title":"APACrefauthors \\ 2015","cited_arxiv_id":null,"evidence_quote":"Documents academic writing challenges faced by international graduate students, reinforcing the rationale for testing native-speaker status."},{"cited_title":"\\ Wald, A.E","cited_arxiv_id":null,"evidence_quote":"Shows international teams differ from domestic teams in media use for complex tasks, supporting the team-composition hypothesis."},{"cited_title":"APACrefauthors \\ 2008","cited_arxiv_id":null,"evidence_quote":"Links cultural diversity and information and communication technology use in global virtual teams, used to justify why team composition might change AI tool purposes."}],"review_version":1}