Pith. sign in

REVIEW 4 cited by

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09939 v2 pith:L2RY3QMM submitted 2024-05-16 cs.CL cs.AI

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

classification cs.CL cs.AI
keywords sciencescientificsciqagansweringdatasetframeworkllmsquestion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA evaluator, which work together to extract diverse and research-level questions and answers from scientific papers. Utilizing this framework, we construct a large-scale, high-quality, open-ended science QA dataset containing 188,042 QA pairs extracted from 22,743 scientific papers across 24 scientific domains. We also introduce SciQAG-24D, a new benchmark task designed to evaluate the science question-answering ability of LLMs. Extensive experiments demonstrate that fine-tuning LLMs on the SciQAG dataset significantly improves their performance on both open-ended question answering and scientific tasks. To foster research and collaboration, we make the datasets, models, and evaluation codes publicly available, contributing to the advancement of science question answering and developing more interpretable and reasoning-capable AI systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

    cs.CL 2026-02 reject novelty 6.0

    DBench-Bio builds a dynamic biology benchmark from post-release abstracts, but LLM-generated gold answers and unverified per-model temporal separation undermine its claim to measure knowledge discovery.

  2. Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models

    cs.AI 2025-06 conditional novelty 6.0

    Active Indexing with synthetic data augmentation for bidirectional fact-source binding during pretraining yields up to 30.2% higher citation precision than passive identifier appending on CitePretrainBench for Qwen models.

  3. ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment

    cs.AI 2026-05 unverdicted novelty 5.0

    ForeSci is a temporally controlled benchmark with 500 tasks for assessing LLM agents on forward-looking AI research judgments in four domains using cutoff-aligned knowledge bases.

  4. ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge

    cs.CE 2025-07 unverdicted novelty 5.0

    ChemDFM-R is a chemical reasoning LLM trained via a four-stage pipeline on the ChemFG dataset of functional-group annotations for molecules and reactions, reaching performance comparable to or better than commercial m...