Pith. sign in

REVIEW 1 cited by

Reasoning Circuits: Few-shot Multihop Question Generation with Structured Rationales

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.08466 v1 pith:SBK4M4OD submitted 2022-11-15 cs.CL

classification cs.CL
keywords generationrationalereasoningmodelperformancequestionrationalesbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-hop Question Generation is the task of generating questions which require the reader to reason over and combine information spread across multiple passages using several reasoning steps. Chain-of-thought rationale generation has been shown to improve performance on multi-step reasoning tasks and make model predictions more interpretable. However, few-shot performance gains from including rationales have been largely observed only in +100B language models, and otherwise require large scale manual rationale annotation. In this work, we introduce a new framework for applying chain-of-thought inspired structured rationale generation to multi-hop question generation under a very low supervision regime (8- to 128-shot). We propose to annotate a small number of examples following our proposed multi-step rationale schema, treating each reasoning step as a separate task to be performed by a generative language model. We show that our framework leads to improved control over the difficulty of the generated questions and better performance compared to baselines trained without rationales, both on automatic evaluation metrics and in human evaluation. Importantly, we show that this is achievable with a modest model size.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automatic Generation of Inference Making Questions for Reading Comprehension Assessments

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-4o generates high-quality reading comprehension questions, but only 42.6% correctly target the specified bridging inference type.

Pith tools