Pith. sign in

REVIEW 24 cited by

The Semantic Scholar Open Data Platform

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.10140 v2 pith:BBQ3P6FU submitted 2023-01-24 cs.DL cs.CL

classification cs.DLcs.CL
keywords datagraphsemanticopenplatformscholarscientificedges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The volume of scientific output is creating an urgent need for automated tools to help scientists keep up with developments in their field. Semantic Scholar (S2) is an open data platform and website aimed at accelerating science by helping scholars discover and understand scientific literature. We combine public and proprietary data sources using state-of-the-art techniques for scholarly PDF content extraction and automatic knowledge graph construction to build the Semantic Scholar Academic Graph, the largest open scientific literature graph to-date, with 200M+ papers, 80M+ authors, 550M+ paper-authorship edges, and 2.4B+ citation edges. The graph includes advanced semantic features such as structurally parsed text, natural language summaries, and vector embeddings. In this paper, we describe the components of the S2 data processing pipeline and the associated APIs offered by the platform. We will update this living document to reflect changes as we add new data offerings and improve existing services.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 57 citations worldwide. Full citation record

  1. Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

    cs.DL 2026-07 conditional novelty 8.0 of 10

    At ICLR, equally scored borderline papers from outside top-25 institutions are accepted less often, a gap concentrated in preprint-identifiable submissions; outcome tests find no evidence of a higher bar.

  2. The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing

    cs.CL 2026-07 unverdicted novelty 7.0 of 10

    NLP authors show migration from *ACL flagship tracks (–19.2pp) to Findings (+14.8pp) and ML venues (+8.6pp), with new authors increasing ML share from 5% to 21% and causal inference indicating a citation premium drive...

  3. The Reciprocal Impact of Science and Software: A Cross-Corpus Analysis of How Research Shapes Software and Software Enables Research

    cs.DL 2026-06 unverdicted novelty 7.0 of 10

    Science and software impact each other through complementary strata, but sparse paper–repo linkage makes reuse–citation coupling gap-sensitive and prevents strong decoupling claims.

  4. Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs

    cs.AI 2026-08 conditional novelty 6.0 of 10

    An LLM-assisted assessment tool, a provenance ontology, and an open knowledge graph holding 140 research integrity assessments of 95 randomised clinical trial publications.

  5. Works on My QPU: Reproducibility in Quantum Computing Research

    quant-ph 2026-07 conditional novelty 6.0 of 10

    Manual review of 127 NISQ papers plus automated scan of ~5000 QC papers finds ~25% code availability and ~65% execution failure among those with code, with concrete recommendations.

  6. Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing

    cs.DL 2026-07 conditional novelty 6.0 of 10

    An editor-native platform unifies research, writing, and publishing in one LaTeX environment with compile-verified agents and patent-to-paper impact signals.

  7. Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Xcientist is a research harness that externalizes an AI scientist's literature grounding, idea evolution, experiments, and repairs into auditable artifacts, demonstrated on memory, traffic forecasting, and PDE-solving tasks.

  8. Crystal: Characterizing Relative Impact of Scholarly Publications

    cs.DL 2026-03 unverdicted novelty 6.0 of 10

    Joint LLM ranking of all citations within a paper identifies impactful references more accurately than isolated classification, gaining +9.5% accuracy and +8.3% F1.

  9. BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A continued-pretrained ModernBERT encoder for biomedical and clinical text claims SOTA on several clinical NLP tasks, with caveats about data overlap between pretraining and evaluation.

  10. Societal AI Research Has Become Less Interdisciplinary

    cs.CL 2025-06 reject novelty 6.0 of 10

    Computer science-only teams now supply a growing majority of societally-oriented AI research on arXiv, even though interdisciplinary teams remain more likely to produce such work.

  11. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.

  12. Toward Living Narrative Reviews: An Empirical Study of the Processes and Challenges in Updating Survey Articles in Computing Research

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Interviews with 11 computing survey authors show that keeping narrative surveys up to date is valued but unrewarded, and that updates fall into empirical, structural, and interpretive types.

  13. Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

    cs.CL 2025-01 conditional novelty 6.0 of 10

    In human evaluations by three professional editors, GPT-4V captions for scientific figures were preferred over author-written captions and over captions from challenge-winning models.

  14. "Dialogue" vs "Dialog" in NLP and AI research: Statistics from a Confused Discourse

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Analysis of tens of thousands of papers shows NLP/AI research mixes 'dialogue' and 'dialog' with no clear trend, author, or context explanation.

  15. Quantifying the Dynamics of Harm Caused by Retracted Research

    cs.DL 2024-12 reject novelty 6.0 of 10

    Citing a retracted paper is associated with a growing citation deficit that is larger for indirect citations and in lower-impact journals, a pattern the authors call 'attention escape'.

  16. How Do We Engage with Other Disciplines? A Framework to Study Meaningful Interdisciplinary Discourse in Scholarly Publications

    cs.DL 2026-01 conditional novelty 5.0 of 10

    A new citation-purpose taxonomy applied to NLP+CSS papers finds that most out-of-discipline citations are shallow and that automated classification of citation purpose is not yet reliable.

  17. MedSEBA: Synthesizing Evidence-Based Answers Grounded in Evolving Medical Literature

    cs.CL 2025-08 conditional novelty 5.0 of 10

    MedSEBA is a RAG-based medical question-answering system that provides cited key arguments, per-study stance labels, and temporal consensus visualization, evaluated by a 10-person user study.

  18. Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ClimateEval unifies 25 climate-related NLP tasks into one benchmark and shows that open-source LLMs gain from few-shot examples but lag on misinformation and fine-grained entity recognition.

  19. Who Gets Recommended? Investigating Gender, Race, and Country Disparities in Paper Recommendations from Large Language Models

    cs.IR 2024-12 reject novelty 5.0 of 10

    LLM recommendations of important AI research favor recent, well-cited, team-authored papers, but do not measurably over-represent male, white, or developed-country scholars relative to a human-curated benchmark.

  20. Hallucination Detector: A hybrid LLM and Semantic Scholar tool calling for detecting hallucination in scientific literature on AtomGPT.org

    cs.DL 2026-07 conditional novelty 4.0 of 10

    AtomGPT's hybrid LLM+Semantic Scholar checker flags 94 of 100 confirmed hallucinated NeurIPS 2025 citations, driven mainly by author mismatch rather than title similarity.

  21. Position: Olfaction Standardization is Essential for the Advancement of Embodied Artificial Intelligence

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A call to add olfaction, with standardized data and benchmarks, to the list of core modalities that embodied AI systems should sense and reason about.

  22. On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts

    cs.CL 2025-02 conditional novelty 4.0 of 10

    With few-shot prompting, Llama 3.1 classifies paper titles and abstracts into five ORKG top-level fields at 0.82 accuracy, about 0.08 above a BERT baseline.

  23. Demo: Interactive Visualization of Semantic Relationships in a Biomedical Project's Talent Knowledge Graph

    cs.SI 2025-01 unverdicted novelty 4.0 of 10

    A web demo maps about 28,000 biomedical researchers and 1,179 datasets into a searchable 2D space and uses GPT-4o to explain collaborator and dataset recommendations.

  24. Charting the Future of Scholarly Knowledge with AI: A Community Perspective

    cs.DL 2025-08 unverdicted novelty 2.0 of 10

    A community perspective on how AI can support scholarly knowledge extraction, organization, and communication, with a proposed classification and ethical considerations.

Pith tools