Pith. sign in

REVIEW 11 cited by

Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.02235 v1 pith:3ELDSIUX submitted 2021-01-06 cs.CL

classification cs.CL
keywords questionreasoningansweringstepsstrategiesstrategyqabenchmarkdecomposition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of $\sim$66%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit

    cs.MA 2026-06 unverdicted novelty 7.0 of 10

    ECHO is a clipped policy-gradient method that uses posterior-sensitive rewards to give turn-level epistemic credit in multi-turn information-seeking tasks, outperforming trajectory-level GRPO on a new Clue Selector Ga...

  2. Beyond Prediction: Tail-Aware Scheduling for LLM Inference

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Presents a distribution-aware scheduling framework for LLM inference that reduces P99 TTLT by 35-50% and TTFT by 34-47% versus SRPT with perfect length knowledge using statistical signals instead of predictions.

  3. Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery

    cs.AI 2026-06 conditional novelty 7.0 of 10

    2-bit quantized reasoning models exhibit process failures like loops and delayed commitment that degrade end-to-end performance, but FP16 planning and loop rescue recover accuracy on MATH-500 from 17.2% to 74.2% for Q...

  4. Proper Scoring Rules for Agentic Uncertainty Quantification

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    Introduces Trajectory Proper Score (TPS) as a strictly proper family of trajectory-level scoring rules that elicits the complete prefix-conditioned success probability process.

  5. Argumentative Large Language Models for Explainable and Contestable Claim Verification

    cs.CL 2024-05 unverdicted novelty 6.0 of 10

    ArgLLMs build argumentation frameworks from LLMs to support explainable and contestable formal reasoning for claim verification.

  6. Efficient Reasoning on the Edge

    cs.LG 2026-03 accept novelty 5.5 of 10

    LoRA adapters, budget-forced GRPO, dynamic switching, parallel verification and FPTQuant enable practical chain-of-thought reasoning on quantized Qwen2.5-7B for edge devices.

  7. When LLM Rationales Become User-Facing: Effects on Trust Perception, Decision-Making, and Gaze Behaviors

    cs.HC 2026-06 unverdicted novelty 5.0 of 10

    Two linked user studies find that LLM rationale correctness and certainty framing affect trust and decision confidence while presentation format does not, and incorrect rationales increase gaze attention and pupil size.

  8. JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    JT-Safe-V2 is a safety-by-design LLM that reports SOTA scores on both capability and safety benchmarks while Safe-MoMA cuts inference cost over 30 percent.

  9. Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Error Reflection Prompting, a chain-of-thought variant that includes an incorrect answer and error recognition, is claimed to improve LLM reasoning performance and interpretability.

  10. Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering

    cs.AI 2026-07 unverdicted novelty 3.0 of 10

    The study applies Bayesian uncertainty propagation to agentic RAG pipelines on StrategyQA and HotpotQA, reporting better discrimination on HotpotQA than on StrategyQA using standard calibration and selective-predictio...

  11. The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey

    cs.AI 2024-04 unverdicted novelty 3.0 of 10

    A survey of emerging AI agent architectures that organizes single and multi-agent designs around reasoning, planning, tool use, communication, and reflection phases.

Pith tools