Pith. sign in

REVIEW 3 cited by

Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03025 v1 pith:2CRPD4MT submitted 2022-04-06 cs.CL

classification cs.CL
keywords feedbacksystemexplanationsmodelquestionsystemsaccuracyanswer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most research on question answering focuses on the pre-deployment stage; i.e., building an accurate model for deployment. In this paper, we ask the question: Can we improve QA systems further \emph{post-}deployment based on user interactions? We focus on two kinds of improvements: 1) improving the QA system's performance itself, and 2) providing the model with the ability to explain the correctness or incorrectness of an answer. We collect a retrieval-based QA dataset, FeedbackQA, which contains interactive feedback from users. We collect this dataset by deploying a base QA system to crowdworkers who then engage with the system and provide feedback on the quality of its answers. The feedback contains both structured ratings and unstructured natural language explanations. We train a neural model with this feedback data that can generate explanations and re-score answer candidates. We show that feedback data not only improves the accuracy of the deployed QA system but also other stronger non-deployed systems. The generated explanations also help users make informed decisions about the correctness of answers. Project page: https://mcgill-nlp.github.io/feedbackqa/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation

    cs.CL 2026-08 accept novelty 7.0 of 10

    Making a persisted evidence record the exclusive input to a later LLM verdict reduces preference agreement and order robustness relative to one-call structured judging.

  2. Aligning Black-box Language Models with Human Judgments

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A linear correction learned from 100 labeled examples per task substantially improves LLM-human agreement on evaluation tasks, including for smaller models.

  3. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools