Pith. sign in

REVIEW 4 cited by

Distilling Knowledge from Reader to Retriever for Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.04584 v2 pith:C5I6S47P submitted 2020-12-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords retrieveransweringdocumentsquestionknowledgemethodsmodelobtain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The task of information retrieval is an important component of many natural language processing systems, such as open domain question answering. While traditional methods were based on hand-crafted features, continuous representations based on neural networks recently obtained competitive results. A challenge of using such methods is to obtain supervised data to train the retriever model, corresponding to pairs of query and support documents. In this paper, we propose a technique to learn retriever models for downstream tasks, inspired by knowledge distillation, and which does not require annotated pairs of query and documents. Our approach leverages attention scores of a reader model, used to solve the task based on retrieved documents, to obtain synthetic labels for the retriever. We evaluate our method on question answering, obtaining state-of-the-art results.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploiting Structural Properties for Efficient Constraint-Aware HNSW Hyperparameter Tuning

    cs.DB 2026-07 conditional novelty 6.0 of 10

    CHAT uses HNSW-specific monotonic and unimodal structure plus resource surrogates to tune M, efc, and efs under constraints, beating black-box tuners by up to 45% throughput or 11% recall and up to 44× faster convergence.

  2. Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Transformers systematically deviate from the Bayes-optimal predictor under high-ambiguity contexts on a new HMM benchmark, and a Monte Carlo predictor that decouples task inference from token prediction partly closes ...

  3. PROBE: Benchmarking Code Generation in Large Language Models

    cs.SE 2026-07 conditional novelty 5.0 of 10

    A multi-language evaluation framework measuring correctness, solution proximity, and code quality finds current LLMs pass at most ~0.70 per language and worsen sharply with problem difficulty.

  4. RAG for Geoscience: What We Expect, Gaps and Opportunities

    cs.ET 2025-08 conditional novelty 5.0 of 10

    Geoscience AI needs a new RAG paradigm that retrieves multimodal Earth data, reasons under physical constraints, generates machine-readable artifacts, and verifies outputs against models and sensors.

Pith tools