Pith. sign in

REVIEW 1 cited by

Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.06578 v3 pith:S4R4LEBD submitted 2023-09-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords evidencehypothesisscientificdatasetdiscernhypotheseslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies in the social sciences. We compare the performance of LLMs to several state-of-the-art benchmarks and highlight opportunities for future research in this area. The dataset is available at https://github.com/Sai90000/ScientificHypothesisEvidencing.git

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Using Large Language Models for Automated Grading of Student Writing about Science

    cs.CL 2024-12 conditional novelty 4.0 of 10

    GPT-4 grading of short astronomy essays matched instructor grades when given a model answer and rubric, and was closer to the instructor than peer grading.

Pith tools