Pith. sign in

REVIEW 1 cited by

Delaying Interaction Layers in Transformer-based Encoders for Efficient Open Domain Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.08422 v1 pith:EI2JRWK5 submitted 2020-10-16 cs.CL

classification cs.CL
keywords modelstransformer-basedallowansweringcorpusdomainefficientodqa
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open Domain Question Answering (ODQA) on a large-scale corpus of documents (e.g. Wikipedia) is a key challenge in computer science. Although transformer-based language models such as Bert have shown on SQuAD the ability to surpass humans for extracting answers in small passages of text, they suffer from their high complexity when faced to a much larger search space. The most common way to tackle this problem is to add a preliminary Information Retrieval step to heavily filter the corpus and only keep the relevant passages. In this paper, we propose a more direct and complementary solution which consists in applying a generic change in the architecture of transformer-based models to delay the attention between subparts of the input and allow a more efficient management of computations. The resulting variants are competitive with the original models on the extractive task and allow, on the ODQA setting, a significant speedup and even a performance improvement in many cases.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Sample Anti-Aliasing and Constrained Optimization for 3D Gaussian Splatting

    cs.CV 2025-08 conditional novelty 3.0 of 10

    Combining 4x multisample anti-aliasing with adaptive error weighting and gradient-difference loss produces modest SSIM/LPIPS gains over vanilla 3DGS on Mip-NeRF360 and Tanks&Temples.

Pith tools