Pith. sign in

REVIEW 4 cited by

Investigating Answerability of LLMs for Long-Form Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08210 v1 pith:44WV3LD2 submitted 2023-09-15 cs.CL

classification cs.CL
keywords llmsquestionssummarieschallengingopen-sourceabstractiveansweringcapabilities
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As we embark on a new era of LLMs, it becomes increasingly crucial to understand their capabilities, limitations, and differences. Toward making further progress in this direction, we strive to build a deeper understanding of the gaps between massive LLMs (e.g., ChatGPT) and smaller yet effective open-source LLMs and their distilled counterparts. To this end, we specifically focus on long-form question answering (LFQA) because it has several practical and impactful applications (e.g., troubleshooting, customer service, etc.) yet is still understudied and challenging for LLMs. We propose a question-generation method from abstractive summaries and show that generating follow-up questions from summaries of long documents can create a challenging setting for LLMs to reason and infer from long contexts. Our experimental results confirm that: (1) our proposed method of generating questions from abstractive summaries pose a challenging setup for LLMs and shows performance gaps between LLMs like ChatGPT and open-source LLMs (Alpaca, Llama) (2) open-source LLMs exhibit decreased reliance on context for generated questions from the original document, but their generation capabilities drop significantly on generated questions from summaries -- especially for longer contexts (>1024 tokens)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AIVA: An AI-based Virtual Companion for Emotion-aware Interaction

    cs.CV 2025-09 conditional novelty 5.0 of 10

    AIVA combines a multimodal sentiment network with LLM prompt engineering, text-to-speech, and an animated avatar to produce emotion-aware companion responses, with MSPN reported to outperform prior multimodal sentimen...

  2. An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring

    cs.MA 2025-05 conditional novelty 5.0 of 10

    A credibility-scoring framework for multi-agent LLM systems, learning agent trustworthiness on the fly and weighting outputs accordingly, improves accuracy under adversarial conditions in some benchmarks.

  3. Inferring Questions from Programming Screenshots

    cs.SE 2025-04 conditional novelty 5.0 of 10

    Multimodal LLMs, especially GPT-4o and Gemini, can infer plausible Stack Overflow questions from code and IDE screenshots with moderate similarity to the original posts, but performance drops on complex screenshots.

  4. An Empirical Study of Evaluating Long-form Question Answering

    cs.IR 2025-04 conditional novelty 5.0 of 10

    In long-form QA, LLM-based evaluators correlate with human judgments better than ROUGE or BERTScore, but they are biased by answer length, question type, self-reinforcement, and rare-word usage, and fine-grained promp...

Pith tools