Pith. sign in

REVIEW

Pruning the Index Contents for Memory Efficient Open-Domain QA

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.10697 v2 pith:ZJVMYSWQ submitted 2021-02-21 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords indexcontentsworkmassivenovelonlyopen-domainpipeline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work presents a novel pipeline that demonstrates what is achievable with a combined effort of state-of-the-art approaches. Specifically, it proposes the novel R2-D2 (Rank twice, reaD twice) pipeline composed of retriever, passage reranker, extractive reader, generative reader and a simple way to combine them. Furthermore, previous work often comes with a massive index of external documents that scales in the order of tens of GiB. This work presents a simple approach for pruning the contents of a massive index such that the open-domain QA system altogether with index, OS, and library components fits into 6GiB docker image while retaining only 8% of original index contents and losing only 3% EM accuracy.

Discussion (0). Continue with ORCID to comment.

Pith tools