Pith. sign in

REVIEW 1 cited by

Exploring Data Splitting Strategies for the Evaluation of Recommendation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.13237 v1 pith:JUYUS725 submitted 2020-07-26 cs.IR

classification cs.IR
keywords splittingstrategysystemsrecommenderdataevaluationimpactmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective methodologies for evaluating recommender systems are critical, so that such systems can be compared in a sound manner. A commonly overlooked aspect of recommender system evaluation is the selection of the data splitting strategy. In this paper, we both show that there is no standard splitting strategy and that the selection of splitting strategy can have a strong impact on the ranking of recommender systems. In particular, we perform experiments comparing three common splitting strategies, examining their impact over seven state-of-the-art recommendation models for two datasets. Our results demonstrate that the splitting strategy employed is an important confounding variable that can markedly alter the ranking of state-of-the-art systems, making much of the currently published literature non-comparable, even when the same dataset and metrics are used.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Double Machine Learning for Adaptive Causal Representation in High-Dimensional Data

    stat.ML 2024-11 reject novelty 4.0 of 10

    With support-points sample splitting, a deep-learning and super-learner hybrid gives lower MSE and deep learning alone gives faster computation than support vector machines in the authors' simulations.

Pith tools