Pith. sign in

REVIEW 2 cited by

QAEncoder: Towards Aligned Representation Learning in Question Answering Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20434 v3 pith:IKJ5JNQ2 submitted 2024-09-30 cs.CL

classification cs.CL
keywords qaencoderembeddingdocumentqueriessystemsaccurateacrossadditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses. However, the inherent gap between user queries and relevant documents hinders precise matching. We introduce QAEncoder, a training-free approach to bridge this gap. Specifically, QAEncoder estimates the expectation of potential queries in the embedding space as a robust surrogate for the document embedding, and attaches document fingerprints to effectively distinguish these embeddings. Extensive experiments across diverse datasets, languages, and embedding models confirmed QAEncoder's alignment capability, which offers a simple-yet-effective solution with zero additional index storage, retrieval latency, training costs, or catastrophic forgetting and hallucination issues. The repository is publicly available at https://github.com/IAAR-Shanghai/QAEncoder.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A human-annotated benchmark with 4,800 QA pairs evaluates how well AI models can retrieve and generate interleaved text-and-image answers.

  2. A Comprehensive Survey on Imbalanced Data Learning

    cs.LG 2025-02 conditional novelty 3.0 of 10

    A structured survey and benchmark that groups imbalanced data learning methods into data re-balancing, feature representation, training strategy, and ensemble learning.

Pith tools