Pith. sign in

REVIEW 4 cited by

SEMQA: Semi-Extractive Multi-Source Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04886 v2 pith:6XNTJ37I submitted 2023-11-08 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords semi-extractiveansweringanswerscapabilitieslanguagemodelstaskabstractive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently proposed long-form question answering (QA) systems, supported by large language models (LLMs), have shown promising capabilities. Yet, attributing and verifying their generated abstractive answers can be difficult, and automatically evaluating their accuracy remains an ongoing challenge. In this work, we introduce a new QA task for answering multi-answer questions by summarizing multiple diverse sources in a semi-extractive fashion. Specifically, Semi-extractive Multi-source QA (SEMQA) requires models to output a comprehensive answer, while mixing factual quoted spans -- copied verbatim from given input sources -- and non-factual free-text connectors that glue these spans together into a single cohesive passage. This setting bridges the gap between the outputs of well-grounded but constrained extractive QA systems and more fluent but harder to attribute fully abstractive answers. Particularly, it enables a new mode for language models that leverages their advanced language generation capabilities, while also producing fine in-line attributions by-design that are easy to verify, interpret, and evaluate. To study this task, we create the first dataset of this kind, QuoteSum, with human-written semi-extractive answers to natural and generated questions, and define text-based evaluation metrics. Experimenting with several LLMs in various settings, we find this task to be surprisingly challenging, demonstrating the importance of QuoteSum for developing and studying such consolidation capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RCStat: A Statistical Framework for using Relative Contextualization in Transformers

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RCStat uses pre-softmax attention logits to define a Relative Contextualization score that improves adaptive KV-cache eviction and attention-head selection for attribution on LLaMA models.

  2. TokenShapley: Token Level Context Attribution with Shapley Value

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TokenShapley computes token-level Shapley attributions from context to response by treating context tokens as (prefix, token) data points in a KNN datastore.

  3. DBRouting: Routing End User Queries to Databases for Answerability

    cs.CL 2025-01 conditional novelty 5.0 of 10

    The paper introduces DB routing, a task of ranking databases for answerability, with synthesized benchmarks from Spider and BIRD showing that LLMs outperform embeddings but both struggle with many or similar databases.

  4. PLD+: Accelerating LLM inference by leveraging Language Model Artifacts

    cs.CL 2024-12 conditional novelty 5.0 of 10

    PLD+ accelerates LLM inference on input-guided tasks by ranking prompt-derived draft spans with hidden states or attention heads, beating tuning-free baselines and often surpassing the tuned EAGLE method.

Pith tools