A language-model-only, two-phase retrieval pipeline with hard-negative training and a grid-searched ensemble is reported to outperform sparse, dense, and generative baselines on a Japanese legal retrieval test set and a small MS MARCO subset.
PrimeQA: The Prime Repository for State-of-the-Art Multilingual Question Answering Research and Development
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The field of Question Answering (QA) has made remarkable progress in recent years, thanks to the advent of large pre-trained language models, newer realistic benchmark datasets with leaderboards, and novel algorithms for key components such as retrievers and readers. In this paper, we introduce PRIMEQA: a one-stop and open-source QA repository with an aim to democratize QA re-search and facilitate easy replication of state-of-the-art (SOTA) QA methods. PRIMEQA supports core QA functionalities like retrieval and reading comprehension as well as auxiliary capabilities such as question generation.It has been designed as an end-to-end toolkit for various use cases: building front-end applications, replicating SOTA methods on pub-lic benchmarks, and expanding pre-existing methods. PRIMEQA is available at : https://github.com/primeqa.
fields
cs.IR 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Optimizing Multi-Stage Language Models for Effective Text Retrieval
A language-model-only, two-phase retrieval pipeline with hard-negative training and a grid-searched ensemble is reported to outperform sparse, dense, and generative baselines on a Japanese legal retrieval test set and a small MS MARCO subset.