Pith. sign in

REVIEW 3 cited by

Synthetic Context Generation for Question Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13188 v1 pith:6U7OHFNQ submitted 2024-06-19 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelssyntheticcontextcontextsgenerationlanguagequestionadvancements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite rapid advancements in large language models (LLMs), QG remains a challenging problem due to its complicated process, open-ended nature, and the diverse settings in which question generation occurs. A common approach to address these challenges involves fine-tuning smaller, custom models using datasets containing background context, question, and answer. However, obtaining suitable domain-specific datasets with appropriate context is often more difficult than acquiring question-answer pairs. In this paper, we investigate training QG models using synthetic contexts generated by LLMs from readily available question-answer pairs. We conduct a comprehensive study to answer critical research questions related to the performance of models trained on synthetic contexts and their potential impact on QG research and applications. Our empirical results reveal: 1) contexts are essential for QG tasks, even if they are synthetic; 2) fine-tuning smaller language models has the capability of achieving better performances as compared to prompting larger language models; and 3) synthetic context and real context could achieve comparable performances. These findings highlight the effectiveness of synthetic contexts in QG and paves the way for future advancements in the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

    cs.CL 2025-06 conditional novelty 6.0 of 10

    BioMol-MQA is a new multimodal QA dataset for polypharmacy in which LLMs perform poorly zero-shot but much better when given gold context.

  2. Multi-Hop Question Generation via Dual-Perspective Keyword Guidance

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A dual-perspective keyword guidance framework with two decoders improves multi-hop question generation over prior keyword-based methods.

  3. LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements

    cs.CL 2025-05 conditional novelty 5.0 of 10

    HSE-Bench is a new 1,020-question LLM benchmark for HSE compliance reasoning, and the paper claims LLMs rely on semantic matching rather than structured legal reasoning.

Pith tools