REVIEW 3 cited by
MCRanker: Generating Diverse Criteria On-the-Fly to Improve Point-wise LLM Rankers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The most recent pointwise Large Language Model (LLM) rankers have achieved remarkable ranking results. However, these rankers are hindered by two major drawbacks: (1) they fail to follow a standardized comparison guidance during the ranking process, and (2) they struggle with comprehensive considerations when dealing with complicated passages. To address these shortcomings, we propose to build a ranker that generates ranking scores based on a set of criteria from various perspectives. These criteria are intended to direct each perspective in providing a distinct yet synergistic evaluation. Our research, which examines eight datasets from the BEIR benchmark demonstrates that incorporating this multi-perspective criteria ensemble approach markedly enhanced the performance of pointwise LLM rankers.
Forward citations
Cited by 3 Pith papers
-
Likert or Not: LLM Absolute Relevance Judgments on Fine-Grained Ordinal Scales
Pointwise LLM scoring with an 11-point ordinal scale is statistically competitive with listwise ranking for 31 of 40 model-dataset combinations on NDCG@10.
-
Precise Zero-Shot Pointwise Ranking with LLMs through Post-Aggregated Global Context Information
A summary-based anchor document enables contrastive pointwise scoring that, when averaged with ordinary pointwise scores, improves zero-shot LLM reranking.
-
LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews
LGAR combines zero-shot LLM graded relevance scoring with monoT5 re-ranking to rank abstracts for systematic reviews, outperforming QA-based baselines by 5-10 pp MAP on two benchmarks.
Discussion (0). Sign in to comment.