REVIEW 3 cited by
On the Difficulty of Evaluating Baselines: A Study on Recommender Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Numerical evaluations with comparisons to baselines play a central role when judging research in recommender systems. In this paper, we show that running baselines properly is difficult. We demonstrate this issue on two extensively studied datasets. First, we show that results for baselines that have been used in numerous publications over the past five years for the Movielens 10M benchmark are suboptimal. With a careful setup of a vanilla matrix factorization baseline, we are not only able to improve upon the reported results for this baseline but even outperform the reported results of any newly proposed method. Secondly, we recap the tremendous effort that was required by the community to obtain high quality results for simple methods on the Netflix Prize. Our results indicate that empirical findings in research papers are questionable unless they were obtained on standardized benchmarks where baselines have been tuned extensively by the research community.
Forward citations
Cited by 3 Pith papers
-
How Reliable Are Semantic-ID Tokenizer Comparisons in Generative Recommendation?
Semantic-ID tokenizers produce collisions affecting up to 30.5% of items across four datasets, inflating Hit@10 by up to 103.36% and making prior tokenizer comparisons unreliable.
-
DABstep: Data Agent Benchmark for Multi-step Reasoning
DABstep releases 450+ real-world financial analytics tasks requiring iterative code and documentation reasoning; the best baseline solves 14.55% of hard tasks.
-
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
Across 17 argument-mining datasets, BERT, RoBERTa, DistilBERT, and WRAP show strong in-benchmark performance but poor cross-dataset generalization, with evidence of shortcut learning tied to content words.
Discussion (0). Continue with ORCID to comment.