Pith. sign in

REVIEW 5 major objections 5 minor 2 references

Societal AI Research Has Become Less Interdisciplinary

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that societally oriented AI research is becoming less interdisciplinary in aggregate: computer science-only teams' share of all societal-content sentences rose from 49.0% in 2014 to 71.2% in 2024, while teams including…

desk verdict Solid cross-sectional result, but the headline temporal claim rests on an unnormalized sentence share and an unresolved denominator contradiction. read the letter →

arxiv 2506.08738 v2 pith:2TGRQ3WR submitted 2025-06-10 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords societalorientationAIresearchinterdisciplinarycollaborationcomputerscienceteamstextclassificationethicspolicytopicmodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

What is being established: societally oriented AI research is becoming less interdisciplinary at the level of the field, even though interdisciplinary teams still produce societally oriented papers at a higher rate. On a corpus of over 100,000 AI papers from 2014 to 2024, the share of all societal-content sentences attributable to computer-science-only teams rose from 49.0% to 71.2%, while the share attributable to teams including social scientists or humanists fell from 25.7% to 3.8%. The paper argues that this combination means the societal turn in AI is being driven from within computer science, not by the cross-disciplinary collaboration that policy guidelines usually recommend. A sympathetic reader would care because the result challenges a widely held assumption about how AI research becomes responsive to social concerns, and it reframes the value of social-science and humanities participation in AI.

What carries the argument

The load-bearing object is a sentence-level classifier that labels each sentence of a paper's abstract, introduction, and conclusion as expressing societal orientation or not; aggregating these labels gives the paper's societal-content share. The classifier is a logistic-regression model trained on 1,002 manually annotated sentences using embeddings from a transformer language model specialized for scientific text, reaching an F1 of 0.93. A second measure feeds a paper's title and abstract to a large language model to generate the main research question, which is then passed through the same classifier to determine whether the paper's central focus is societal. Team disciplinary composition is inferred from each author's prior publication history: authors are labeled CS, natural-science/medicine, or social-science/humanities when at least 90% of their prior publications fall in one category, and papers are then grouped into CS-only, SSH-inclusive, NSM-inclusive, and fully interdisciplinary teams, with 80% and 75% thresholds used as robustness checks.

What would settle it

Human annotators blind to team type would re-label a stratified sample of full papers from 2014 and 2024 across CS-only and SSH/NSM-inclusive teams; if the human-annotated societal sentence shares do not reproduce the 49% to 71% shift in CS-only attribution, or if classifier error rates differ by team type, the central claim fails.

Watch

Extended reading notes

Core claim

The central discovery is that the aggregate locus of societal orientation in AI research has moved away from interdisciplinary teams. Per-paper, the ordering is exactly what policy intuition predicts: teams including social scientists or humanists average 20.9% societally oriented sentences and 35% societally focused research questions, against 7.8% and 8% for CS-only teams, and a fixed-effects regression confirms these gaps after controlling for year, subfield, team size, and length. Yet when the total societal output of the field is summed by year, the CS-only contribution rises from 49.0% in 2014 to 71.2% in 2024 (β=2.05, p=0.001), and SSH-inclusive teams' contribution collapses from 25.7% to 3.8%; the same pattern appears for societally framed research questions. Topic-level analysis shows the CS-only dominance is broad, not niche: they generate the largest number of high-scoring papers on topics such as gender and race, language and translation, and medical imaging. The authors conclude that evolving norms within computer science, a shift toward applied research, and the rise of computational social science are plausible drivers, and they raise the open question of what distinctive contribution social scientists and humanists can make if technical teams are already absorbing societal concerns.

Load-bearing premise

The trend rests on the assumption that the sentence-level classifier's labels measure societal orientation equally well for computer-science-only and interdisciplinary teams in every year, rather than tracking differences in writing style, section availability, or paper length.

Editorial extensions

If this is right

  • The topic analysis implies that CS-only teams' societal output is broad: they supply the largest volume of top-relevance papers on gender and race, language and translation, and medical imaging, not just a single ethical subfield.
  • If the trend is real, the field's aggregate societal output is increasingly produced by teams that per paper express societal orientation at roughly a third the rate of SSH-inclusive teams, so overall societal engagement could remain flat or decline even as CS-only volume grows.
  • The paper's three proposed explanations—internal norm change, a shift from foundational to applied research, and the rise of computational social science—are each compatible with the data, meaning the paper establishes the shift but not which mechanism drives it.
  • Institutional mechanisms such as broader-impact statements and ethics review become plausible causes rather than failed interventions: the rise in CS-only societal output is consistent with norms diffusing through the technical community rather than through interdisciplinary collaboration.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not tested in the paper: a formal decomposition of the 49% to 71% shift into per-paper intensity, team size, and subfield composition would show whether the rise is driven mostly by CS-only authors writing more societally per paper or by shifts in where papers are published.
  • Because the definition of societal orientation bundles normative values with applied societal topics, re-running the analysis with those two components separated could reveal whether the CS-only rise comes from ethical framing or from mentioning applications such as healthcare and misinformation.
  • An obvious extension is to compare the preprint trend with peer-reviewed proceedings only, which would indicate whether the shift reflects preprint norms or formal publication norms.
  • Another unaddressed possibility is that CS-only societal papers cluster in large industrial research groups with dedicated ethics and policy teams; if so, the relevant driver would be organizational resources rather than disciplinary composition.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper analyzes 101,919 arXiv papers (2014–2024) from four subfields to examine whether interdisciplinary research teams lead the integration of societal and ethical concerns into AI research. Using a supervised sentence-level classifier and a team-composition typology based on Semantic Scholar author histories, the authors report two main findings: (i) interdisciplinary teams produce societally-oriented research at higher per-paper rates, but (ii) the aggregate share of society-oriented sentences attributable to CS-only teams rose from 49.0% in 2014 to 71.2% in 2024, with similar growth in society-oriented research questions. They conclude that societal AI research has become less interdisciplinary, and discuss possible explanations including evolving CS norms, a shift to applied research, and computational social science.

Significance. If the central empirical claim were established, the paper would provide a large-scale, timely challenge to the widespread policy assumption that interdisciplinary collaboration is the main driver of societal orientation in AI research. The paper's strengths include a large corpus, a clearly defined measurement concept, transparency about the codebook and prompts, and robustness checks on classification thresholds and author-assignment thresholds. However, the headline conclusion rests on a raw sentence-share aggregation that is not normalized for paper counts or length, and the text contains a direct internal contradiction about the baseline team-composition distribution. These issues are load-bearing: until they are resolved, the 49%→71% shift cannot be interpreted as a decline in interdisciplinarity.

major comments (5)
  1. [Section 3.1 / Figure 3] The central claim that CS-only teams' share of societal output rose from 49.0% to 71.2% is computed as a raw share of all societally-oriented sentences, not as a per-paper rate. Such a share can rise simply because CS-only papers became more numerous or longer, even if their per-paper societal orientation was flat or falling. The paper does not report a normalized time trend of societal orientation by team type: Figure 1 pools 2014–2024, and Appendix B's regression controls for article length but does not decompose the temporal shift in attribution. The authors must decompose the Figure 3 trend into paper-count, length, and per-paper intensity components, and report per-paper societal orientation over time by team type, before the headline conclusion can be drawn.
  2. [Section 3.1 / Figure 2] There is a direct internal contradiction about the baseline. The text states that 'the share of interdisciplinary teams has remained relatively stable at about 75%,' while the Figure 2 caption states that 'CS-only teams consistently comprise the majority of research output.' These claims cannot both be true: if CS-only teams are a majority of papers, the interdisciplinary share cannot be 75%. Since the interpretation of the 71.2% sentence-level share depends on whether CS-only teams are about 25% or about 75% of all papers, this denominator must be resolved. The authors should report the exact annual shares of all four team types and reconcile the text with Figure 2.
  3. [Section 5 / Classifier validation] The classifier is trained on 1,002 sentences, with no reported inter-annotator agreement, no validation on full papers, and no uncertainty propagation into the Figure 3 trend. If classifier errors correlate with writing style, section availability, or paper length—which is plausible given that the classifier was trained on sentences sampled from subfields and keyword filters—the temporal trend could be an artifact. The paper should report classifier performance by team type and year, validate on held-out full papers (or at least on the Abstract/Introduction/Conclusion sections used in the corpus), and propagate classification uncertainty into the aggregate shares or provide a sensitivity analysis bounding the trend.
  4. [Section 5 / Data preprocessing] The Methods section states that '86% of the papers were missing abstracts, 18% were missing their introduction section, and 35% were missing the conclusion.' The 86% figure is implausible for arXiv papers and is inconsistent with the claim that only 35 papers were missing all three sections; it is likely a typo (perhaps 8.6%). If that many abstracts were genuinely unavailable, the research-question extraction (which relies on title and abstract) and the sentence-level measure would be severely affected, and missingness correlated with year or team type could bias the trend. The authors must correct this figure and clarify how missing sections are handled in the aggregation.
  5. [Appendix B] The fixed-effects regression establishes a pooled association between team type and societal orientation, but it does not address the temporal decomposition that is central to the paper's claim. The appendix should include a regression or decomposition that separates the period effect on CS-only output share into within-team intensity changes and compositional changes. Without this, the robustness checks in Appendix A and B do not support the Figure 3 trend, only the cross-sectional team-type differences.
minor comments (5)
  1. [Section 5 / Classifier performance] The F1 score is reported inconsistently: Section 2.2 says the final model achieved F1 = 0.93, while Section 5 says logistic regression achieved 0.94 accuracy and 0.93 F1, and also that the SciBERT model had 0.90 accuracy. The authors should state clearly which model was used for the main results and provide a single consistent performance table.
  2. [Section 3.2 / Figure 4] The text says 'we selected three of the most frequent topics' but the Figure 4 caption says 'each row represents one of the four most frequent societal topics'; also the caption lists Gender and Race, Language and Translation, and Medical Imaging, which is three. Please align the text and the figure.
  3. [Section 5 / Annotation] The annotation section states that two research assistants annotated sentences with an author as tie-breaker, but no inter-annotator agreement statistic (e.g., Cohen's kappa) is reported. Please add this for transparency.
  4. [Section 1] The first paragraph contains a typo: 'V oeneky' appears in the references and main text; the correct author name is 'Voeneky' (also on line 'V oeneky et al.'). Please correct throughout.
  5. [Appendix C] The codebook example for 'Special Note on Bias' is helpful, but the paper should clarify how the classifier handles technical uses of 'bias' given that the codebook explicitly excludes them; currently the classifier is trained on sentences, and this nuance may not be captured reliably.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the societal-orientation measure is trained on independent human annotations and the headline trend is a descriptive aggregation, not a fitted prediction.

full rationale

The paper's central claim—the rise in the CS-only share of societal output from 49.0% to 71.2%—is a descriptive aggregation of two independently constructed variables: sentence-level societal labels from a classifier trained on 1,002 manually annotated sentences, and team-type labels inferred from Semantic Scholar publication histories. Neither variable is fit to the target trend; no parameter is tuned to reproduce Figure 3. The robustness checks (classifier thresholds 0.6/0.8, field-of-study thresholds 80%/75%) vary definitions without re-estimating the outcome, which is consistent with a measurement-validity exercise rather than circularity. The only self-citations (Gilardi et al. 2024; Kotarcic et al. 2022) appear in the introduction as contextual examples and play no role in the derivation. Appendix B's regression controls for article length and team size and uses the same classifier labels, but it is a descriptive confirmatory analysis, not a prediction derived from fitted inputs. Thus the dominant threats are construct validity and denominator interpretation, not circular reasoning.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on measurement assumptions rather than derivation. The most consequential are the validity of the societal-orientation classifier and the field-of-study assignment. Thresholds are varied in robustness checks, but the aggregate trend is not tested against per-paper rates.

free parameters (3)
  • Sentence classification probability threshold = 0.5 (robustness at 0.6 and 0.8)
    Sentences are labeled societal if predicted probability exceeds 0.5; thresholds are manually set. The paper's main figures use 0.5.
  • Author field-of-study assignment threshold = 90% (robustness at 80% and 75%)
    An author is labeled single-discipline if at least 90% of prior publications fall in one umbrella category; chosen by hand.
  • Topic model parameters = UMAP n_neighbors=20, n_components=5; HDBSCAN min_cluster_size=300
    Parameters for BERTopic are chosen by the authors and affect which societal topics are identified in Section 3.2.
assumptions (4)
  • domain assumption Semantic Scholar field-of-study tags validly proxy author disciplinary identity.
    Used to assign each author to CS, NSM, or SSH in Sections 2.1 and 5; misclassification would change all team-type comparisons.
  • domain assumption Abstract, introduction, and conclusion sections represent each paper's societal orientation.
    Section 2.2 says only these sections were analyzed; the paper does not test whether omitting other sections changes the measure.
  • domain assumption The logistic regression classifier trained on 1,002 sentences generalizes across the full corpus and over time.
    Methods report test F1=0.93 on a small balanced sample; no full-paper validation or precision estimate on the natural corpus is provided.
  • domain assumption The Llama 3.1 model extracts valid main research questions from title and abstract.
    Section 5 says 200 output questions were manually reviewed, but given 86% of papers were missing abstracts in the PDF extraction, the input text for most papers is unclear.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Societal AI Research Has Become Less Interdisciplinary." pith.science (2026). https://pith.science/paper/2TGRQ3WR

@misc{pith2026250608738,
  author       = {Pith},
  title        = {Pith review of: Societal AI Research Has Become Less Interdisciplinary},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2TGRQ3WR}},
  note         = {Machine review of arXiv:2506.08738}
}
read the original abstract

As artificial intelligence (AI) systems become deeply embedded in everyday life, calls to align AI development with ethical and societal values have intensified. Interdisciplinary collaboration is often championed as a key pathway for fostering such engagement. Yet it remains unclear whether interdisciplinary research teams are actually leading this shift in practice. This study analyzes over 100,000 AI-related papers published on ArXiv between 2014 and 2024 to examine how ethical values and societal concerns are integrated into technical AI research. We develop a classifier to identify societal content and measure the extent to which research papers express these considerations. We find a striking shift: while interdisciplinary teams remain more likely to produce societally-oriented research, computer science-only teams now account for a growing share of the field's overall societal output. These teams are increasingly integrating societal concerns into their papers and tackling a wide range of domains - from fairness and safety to healthcare and misinformation. These findings challenge common assumptions about the drivers of societal AI and raise important questions. First, what are the implications for emerging understandings of AI safety and governance if most societally-oriented research is being undertaken by exclusively technical teams? Second, for scholars in the social sciences and humanities: in a technical field increasingly responsive to societal demands, what distinctive perspectives can we still offer to help shape the future of AI?

Figures

Figures reproduced from arXiv: 2506.08738 by the authors.

Figure 1
Figure 1. A. Average proportion of societally-oriented sentences per paper (the Abstract, Introduc￾tion and Conclusion sections), by team type. B. Share of papers whose research question centers on a societal issue, by team type. Error bars represent 95% confidence intervals. However, this pattern tells only part of the story. As shown in [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Proportion of team types in AI research, 2014–2024. Each line represents the annual [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. A. Share of total sentence-level societal orientation attributable to each team type, by year. B. Share of papers with societal research questions attributable to each team type, by year. Both panels use stacked area charts, with the y-axis representing the proportion of all societally-oriented content (Panel A) or papers (Panel B) in each year. Both results show a clear and consistent rise in the contribution of CS… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Topic-specific societal orientation by team type. Each row represents one of the four most [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Classifier Threshold Sensitivity. Average societal orientation (%) by team type across three sentence-level classification thresholds: 0.5 (original), 0.6, and 0.8. Results show consistent ordering across team types and highlight the robustness of interdisciplinary eff…
Figure 6
Figure 6. Figure 6: Field-of-Study Assignment Thresholds. Average societal orientation (%) by team type across three thresholds (90%, 80%, 75%) used to classify authors’ primary field based on their prior publications. Lowering the threshold makes it easier for authors with mixed publicat…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 linked inside Pith

  1. [2]

    Visions, Values, V oices: A Survey of Artificial Intelligence Researchers

    Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), ed. Kevin Duh, Helena Gomez and Steven Bethard. Association for Computational Linguistics...

  2. [2024]

    Ethical concern identification in NLP: A corpus of ACL Anthology ethics statements

    “Ethical concern identification in NLP: A corpus of ACL Anthology ethics statements.” arXiv. Kasirzadeh, Atoosa. 2025. “Two Types of AI Existential Risk: Decisive and Accumulative.” Philosophical Studies. Kinney, Rodney Michael, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexan- dra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.