{"id":"2e23fcf0-789b-4f32-97f8-746d353da971","arxiv_id":"2506.08738","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Computer science-only teams now supply a growing majority of societally-oriented AI research on arXiv, even though interdisciplinary teams remain more likely to produce such work.","lead":"This paper measures whether AI research papers that discuss ethical or societal concerns are written by interdisciplinary teams, using over 100,000 arXiv papers from 2014 to 2024. It finds that computer science-only teams now produce a growing majority of this societally-oriented output, challenging the idea that interdisciplinary collaboration drives societal AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 3's raw sentence-share trend is not normalized for paper counts or length, and the paper's baseline team-composition claim is internally contradictory (75% interdisciplinary vs. CS-only majority), so the 49%→71% shift cannot be interpreted as a decline in interdisciplinarity.","rationale":"The paper asks a timely question and assembles a large corpus with an explicit classifier and several robustness checks; those are real assets. My concern is not with the classifier per se but with the inferential link between the raw sentence-share metric and the headline 'less interdisciplinary' claim. Figure 3's share is a volume-weighted total, so it cannot distinguish 'CS-only teams write more societal prose because there are more of them or their papers are longer' from 'CS-only teams are internally shifting toward societal concerns.' The paper's own Section 3.1 simultaneously reports a stable ~75% interdisciplinary baseline and a Figure 2 caption saying CS-only teams are the stable majority; those cannot both be true. Since the 49→71 trend is only meaningful relative to the team-type paper-share baseline, this contradiction is load-bearing. A reader cannot determine whether the observed shift is a compositional artifact or a genuine intellectual change. The reader's weakest_assumption about measurement invariance is related and reasonable, but the sharpest problem is the missing normalization and the unresolved denominator. A per-paper time trend and a baseline reconciliation would settle it. Because this concern directly undercuts the paper's central empirical assertion, I support the reader's REJECT verdict without adjustment.","tokens_in":14972,"tokens_out":5073,"duration_ms":57197,"concrete_test":"Recompute Figure 3 using per-paper societal orientation (percent of societally-oriented sentences per paper) averaged within each team type by year, and also compute a decomposition of the raw share change into (a) paper-count share, (b) mean section length, and (c) mean per-paper intensity. If the per-paper intensity trend for CS-only teams is flat or negative, or if CS-only paper-count share already exceeded 71% in 2014, the headline claim fails. Separately, report the actual annual shares of CS-only and interdisciplinary teams to resolve the Figure 2/text contradiction; if CS-only teams were a stable majority, the 49→71 shift is a convergence toward paper share, not an interdisciplinary decline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—'the proportion of all sentence-level societal orientation attributable to CS-only teams rose from 49.0% in 2014 to 71.2% in 2024' (Section 3.1)—is computed as a raw share of societally-oriented sentences, not as a per-paper percentage. Such a share can rise simply because CS-only papers became more numerous or longer, even if their per-paper societal orientation was flat or falling. The paper reports no normalized time trend of societal orientation by team type: Figure 1 pools 2014–2024, and Appendix B's regression controls for article length but does not decompose the temporal shift in attribution. The baseline needed to interpret this shift is also contradictory: the text says 'the share of interdisciplinary teams has remained relatively stable at about 75%,' while the Figure 2 caption says 'CS-only teams consistently comprise the majority of research output.' If CS-only teams were already roughly 75% of papers, then a 71.2% societal-output share shows underrepresentation relative to paper share, not dominance; if they were roughly 25%, the shift is dramatic. The paper cannot support its headline conclusion until this denominator is resolved and the raw share is decomposed into paper-count, length, and per-paper intensity components.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 101,919 arXiv papers (2014–2024) from four subfields to examine whether interdisciplinary research teams lead the integration of societal and ethical concerns into AI research. Using a supervised sentence-level classifier and a team-composition typology based on Semantic Scholar author histories, the authors report two main findings: (i) interdisciplinary teams produce societally-oriented research at higher per-paper rates, but (ii) the aggregate share of society-oriented sentences attributable to CS-only teams rose from 49.0% in 2014 to 71.2% in 2024, with similar growth in society-oriented research questions. They conclude that societal AI research has become less interdisciplinary, and discuss possible explanations including evolving CS norms, a shift to applied research, and computational social science.","tokens_in":15098,"tokens_out":3273,"duration_ms":38743,"significance":"If the central empirical claim were established, the paper would provide a large-scale, timely challenge to the widespread policy assumption that interdisciplinary collaboration is the main driver of societal orientation in AI research. The paper's strengths include a large corpus, a clearly defined measurement concept, transparency about the codebook and prompts, and robustness checks on classification thresholds and author-assignment thresholds. However, the headline conclusion rests on a raw sentence-share aggregation that is not normalized for paper counts or length, and the text contains a direct internal contradiction about the baseline team-composition distribution. These issues are load-bearing: until they are resolved, the 49%→71% shift cannot be interpreted as a decline in interdisciplinarity.","major_comments":[{"comment":"The central claim that CS-only teams' share of societal output rose from 49.0% to 71.2% is computed as a raw share of all societally-oriented sentences, not as a per-paper rate. Such a share can rise simply because CS-only papers became more numerous or longer, even if their per-paper societal orientation was flat or falling. The paper does not report a normalized time trend of societal orientation by team type: Figure 1 pools 2014–2024, and Appendix B's regression controls for article length but does not decompose the temporal shift in attribution. The authors must decompose the Figure 3 trend into paper-count, length, and per-paper intensity components, and report per-paper societal orientation over time by team type, before the headline conclusion can be drawn.","section":"Section 3.1 / Figure 3"},{"comment":"There is a direct internal contradiction about the baseline. The text states that 'the share of interdisciplinary teams has remained relatively stable at about 75%,' while the Figure 2 caption states that 'CS-only teams consistently comprise the majority of research output.' These claims cannot both be true: if CS-only teams are a majority of papers, the interdisciplinary share cannot be 75%. Since the interpretation of the 71.2% sentence-level share depends on whether CS-only teams are about 25% or about 75% of all papers, this denominator must be resolved. The authors should report the exact annual shares of all four team types and reconcile the text with Figure 2.","section":"Section 3.1 / Figure 2"},{"comment":"The classifier is trained on 1,002 sentences, with no reported inter-annotator agreement, no validation on full papers, and no uncertainty propagation into the Figure 3 trend. If classifier errors correlate with writing style, section availability, or paper length—which is plausible given that the classifier was trained on sentences sampled from subfields and keyword filters—the temporal trend could be an artifact. The paper should report classifier performance by team type and year, validate on held-out full papers (or at least on the Abstract/Introduction/Conclusion sections used in the corpus), and propagate classification uncertainty into the aggregate shares or provide a sensitivity analysis bounding the trend.","section":"Section 5 / Classifier validation"},{"comment":"The Methods section states that '86% of the papers were missing abstracts, 18% were missing their introduction section, and 35% were missing the conclusion.' The 86% figure is implausible for arXiv papers and is inconsistent with the claim that only 35 papers were missing all three sections; it is likely a typo (perhaps 8.6%). If that many abstracts were genuinely unavailable, the research-question extraction (which relies on title and abstract) and the sentence-level measure would be severely affected, and missingness correlated with year or team type could bias the trend. The authors must correct this figure and clarify how missing sections are handled in the aggregation.","section":"Section 5 / Data preprocessing"},{"comment":"The fixed-effects regression establishes a pooled association between team type and societal orientation, but it does not address the temporal decomposition that is central to the paper's claim. The appendix should include a regression or decomposition that separates the period effect on CS-only output share into within-team intensity changes and compositional changes. Without this, the robustness checks in Appendix A and B do not support the Figure 3 trend, only the cross-sectional team-type differences.","section":"Appendix B"}],"minor_comments":[{"comment":"The F1 score is reported inconsistently: Section 2.2 says the final model achieved F1 = 0.93, while Section 5 says logistic regression achieved 0.94 accuracy and 0.93 F1, and also that the SciBERT model had 0.90 accuracy. The authors should state clearly which model was used for the main results and provide a single consistent performance table.","section":"Section 5 / Classifier performance"},{"comment":"The text says 'we selected three of the most frequent topics' but the Figure 4 caption says 'each row represents one of the four most frequent societal topics'; also the caption lists Gender and Race, Language and Translation, and Medical Imaging, which is three. Please align the text and the figure.","section":"Section 3.2 / Figure 4"},{"comment":"The annotation section states that two research assistants annotated sentences with an author as tie-breaker, but no inter-annotator agreement statistic (e.g., Cohen's kappa) is reported. Please add this for transparency.","section":"Section 5 / Annotation"},{"comment":"The first paragraph contains a typo: 'V oeneky' appears in the references and main text; the correct author name is 'Voeneky' (also on line 'V oeneky et al.'). Please correct throughout.","section":"Section 1"},{"comment":"The codebook example for 'Special Note on Bias' is helpful, but the paper should clarify how the classifier handles technical uses of 'bias' given that the codebook explicitly excludes them; currently the classifier is trained on sentences, and this nuance may not be captured reliably.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"This manuscript addresses a timely and important question, and its data-collection effort is substantial. However, the central empirical claim is not yet supported because of the raw-share aggregation issue and the internal contradiction about the team-composition baseline. Both are fixable with additional analyses, so I do not recommend rejection, but the revision needs to be substantive. Peer reviewers should specifically ask for a decomposition of the Figure 3 trend and a resolution of the 75% vs. majority contradiction. I also noted that the authors state replication materials will be made available 'upon publication,' which is weaker than the journal's typical data-availability expectation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has one genuinely useful result buried under a headline that doesn't survive contact with the data. The cross-sectional pattern—interdisciplinary teams, especially SSH-inclusive ones, write more societally-oriented papers—is credible and well-supported by multiple measures. That part is worth publishing. But the paper's central temporal claim, that CS-only teams now account for 71% of societal output, is built on a raw sentence share that isn't normalized for paper counts or length, and the text contradicts its own denominator.\n\nFirst, the contradiction: Section 3.1 says the share of interdisciplinary teams has stayed stable at about 75%, while Figure 2's caption says CS-only teams are the consistent majority. Both cannot be true. If CS-only teams are ~75% of papers, then their 71% sentence share shows underrepresentation, not dominance. If they are ~25%, the shift is dramatic. The paper never resolves this, and the entire interpretation hinges on it.\n\nSecond, the 49%→71% figure is a share of all societally-oriented sentences. It can rise simply because CS-only papers became more numerous or longer, even if their per-paper societal orientation was flat. The paper reports no per-paper normalized trend by team type. Appendix B's regression controls for article length, but it doesn't decompose the attribution shift. So the claim that CS-only teams are 'increasingly integrating societal concerns' is not actually demonstrated.\n\nThere are also smaller issues: the 86% missing-abstract rate in the PDF pipeline is unexplained, and if abstracts were missing for most papers, it's unclear what the LLM was prompted with for the research-question measure. Classifier uncertainty isn't propagated into the aggregate trends.\n\nTo be fair, the paper does several things well: the corpus is large, the annotation and classifier are described in detail, robustness checks on thresholds are included, and the cross-sectional finding is likely real. The topic is timely and the framing is honest about limitations.\n\nAs it stands, the central claim is not supported. The paper needs a re-analysis with per-paper normalization, a resolved team-composition baseline, and a clearer account of the abstract pipeline. If those are fixed, this could be a solid contribution. I'd send it to a serious referee in its current form only if the editor expects heavy revision; otherwise, it's a reject-and-resubmit.","headline":"Solid cross-sectional result, but the headline temporal claim rests on an unnormalized sentence share and an unresolved denominator contradiction.","tokens_in":15731,"tokens_out":3558,"would_cite":false,"duration_ms":38330,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that societally oriented AI research is becoming less interdisciplinary in aggregate: computer science-only teams' share of all societal-content sentences rose from 49.0% in 2014 to 71.2% in 2024, while teams including…","keywords":["societal orientation","AI research","interdisciplinary collaboration","computer science teams","text classification","AI ethics","research policy","topic modeling"],"falsifier":"Human annotators blind to team type would re-label a stratified sample of full papers from 2014 and 2024 across CS-only and SSH/NSM-inclusive teams; if the human-annotated societal sentence shares do not reproduce the 49% to 71% shift in CS-only attribution, or if classifier error rates differ by team type, the central claim fails.","tokens_in":14654,"feed_emoji":"🤖","tokens_out":13194,"duration_ms":138521,"temperature":0.7,"pith_summary":"What is being established: societally oriented AI research is becoming less interdisciplinary at the level of the field, even though interdisciplinary teams still produce societally oriented papers at a higher rate. On a corpus of over 100,000 AI papers from 2014 to 2024, the share of all societal-content sentences attributable to computer-science-only teams rose from 49.0% to 71.2%, while the share attributable to teams including social scientists or humanists fell from 25.7% to 3.8%. The paper argues that this combination means the societal turn in AI is being driven from within computer science, not by the cross-disciplinary collaboration that policy guidelines usually recommend. A sympathetic reader would care because the result challenges a widely held assumption about how AI research becomes responsive to social concerns, and it reframes the value of social-science and humanities participation in AI.","feed_headline":"Societal AI research is shifting to computer-science-only teams","feed_subtitle":"Computer-science-only teams grew from 49% to 71% of societal AI output in a decade, though mixed teams lead per paper.","key_machinery":"The load-bearing object is a sentence-level classifier that labels each sentence of a paper's abstract, introduction, and conclusion as expressing societal orientation or not; aggregating these labels gives the paper's societal-content share. The classifier is a logistic-regression model trained on 1,002 manually annotated sentences using embeddings from a transformer language model specialized for scientific text, reaching an F1 of 0.93. A second measure feeds a paper's title and abstract to a large language model to generate the main research question, which is then passed through the same classifier to determine whether the paper's central focus is societal. Team disciplinary composition is inferred from each author's prior publication history: authors are labeled CS, natural-science/medicine, or social-science/humanities when at least 90% of their prior publications fall in one category, and papers are then grouped into CS-only, SSH-inclusive, NSM-inclusive, and fully interdisciplinary teams, with 80% and 75% thresholds used as robustness checks.","core_discovery":"The central discovery is that the aggregate locus of societal orientation in AI research has moved away from interdisciplinary teams. Per-paper, the ordering is exactly what policy intuition predicts: teams including social scientists or humanists average 20.9% societally oriented sentences and 35% societally focused research questions, against 7.8% and 8% for CS-only teams, and a fixed-effects regression confirms these gaps after controlling for year, subfield, team size, and length. Yet when the total societal output of the field is summed by year, the CS-only contribution rises from 49.0% in 2014 to 71.2% in 2024 (β=2.05, p=0.001), and SSH-inclusive teams' contribution collapses from 25.7% to 3.8%; the same pattern appears for societally framed research questions. Topic-level analysis shows the CS-only dominance is broad, not niche: they generate the largest number of high-scoring papers on topics such as gender and race, language and translation, and medical imaging. The authors conclude that evolving norms within computer science, a shift toward applied research, and the rise of computational social science are plausible drivers, and they raise the open question of what distinctive contribution social scientists and humanists can make if technical teams are already absorbing societal concerns.","pith_inferences":["Not tested in the paper: a formal decomposition of the 49% to 71% shift into per-paper intensity, team size, and subfield composition would show whether the rise is driven mostly by CS-only authors writing more societally per paper or by shifts in where papers are published.","Because the definition of societal orientation bundles normative values with applied societal topics, re-running the analysis with those two components separated could reveal whether the CS-only rise comes from ethical framing or from mentioning applications such as healthcare and misinformation.","An obvious extension is to compare the preprint trend with peer-reviewed proceedings only, which would indicate whether the shift reflects preprint norms or formal publication norms.","Another unaddressed possibility is that CS-only societal papers cluster in large industrial research groups with dedicated ethics and policy teams; if so, the relevant driver would be organizational resources rather than disciplinary composition."],"forward_implications":["The topic analysis implies that CS-only teams' societal output is broad: they supply the largest volume of top-relevance papers on gender and race, language and translation, and medical imaging, not just a single ethical subfield.","If the trend is real, the field's aggregate societal output is increasingly produced by teams that per paper express societal orientation at roughly a third the rate of SSH-inclusive teams, so overall societal engagement could remain flat or decline even as CS-only volume grows.","The paper's three proposed explanations—internal norm change, a shift from foundational to applied research, and the rise of computational social science—are each compatible with the data, meaning the paper establishes the shift but not which mechanism drives it.","Institutional mechanisms such as broader-impact statements and ethics review become plausible causes rather than failed interventions: the rise in CS-only societal output is consistent with norms diffusing through the technical community rather than through interdisciplinary collaboration."],"supporting_citations":[{"why":"Supplies the preprint-repository paper corpus and download pipeline on which the entire 100k-paper analysis is built.","marker":"Clement et al., 2019"},{"why":"Provides each author's publication history and field-of-study labels used to infer team disciplinary composition.","marker":"Kinney et al., 2023"},{"why":"Provides the scientific-text language model whose embeddings the sentence-level societal-orientation classifier fine-tunes.","marker":"Beltagy, Lo and Cohan, 2019"},{"why":"Grounds the annotation codebook for identifying ethical values in AI research statements.","marker":"Ashurst, Hine, Sedille and Carlier, 2022"},{"why":"Survey of AI researchers' concerns used to define the societal concerns that the codebook looks for.","marker":"Grace et al., 2024"},{"why":"Survey of machine-learning researchers' ethics and governance views used alongside Grace et al. to shape the codebook.","marker":"Zhang et al., 2021"},{"why":"Provides the topic-modeling method used to show that CS-only teams contribute the largest volume of high-scoring papers across societal topics.","marker":"Grootendorst, 2022"}],"fun_headline_variants":["CS-only teams surge to 71% of societal AI research","Societal AI output shifts decisively to CS-only teams","Interdisciplinary AI teams lose ground: CS-only dominates output","Per paper, interdisciplinary wins; overall, CS-only rules societal AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The trend rests on the assumption that the sentence-level classifier's labels measure societal orientation equally well for computer-science-only and interdisciplinary teams in every year, rather than tracking differences in writing style, section availability, or paper length.","fun_headline_variants_meta":{"raw":{"variants":["CS-only teams surge to 71% of societal AI research","Societal AI output shifts decisively to CS-only teams","Interdisciplinary AI teams lose ground: CS-only dominates output","Per paper, interdisciplinary wins; overall, CS-only rules societal AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001089,"raw_usage":{"total_tokens":4603,"prompt_tokens":1054,"completion_tokens":3549,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":3479}},"tokens_in":670,"tokens_out":3549,"duration_ms":31137,"temperature":1.0,"reasoning_tokens":3479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:03:59.490045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Human annotators blind to team type would re-label a stratified sample of full papers from 2014 and 2024 across CS-only and SSH/NSM-inclusive teams; if the human-annotated societal sentence shares do not reproduce the 49% to 71% shift in CS-only attribution, or if classifier error rates differ by team type, the central claim fails.","supporting_citations":[],"review_version":1}