Pith. sign in

REVIEW 3 major objections 4 minor 49 references

On the discriminative power of Hyper-parameters in Cross-Validation and how to choose them

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read nDCG and Precision are the best hyperparameter tuners

desk verdict Useful metric comparison for tuning, but the dominant-hyperparameter claim doesn't follow from the analysis. read the letter →

arxiv 1909.02523 v1 pith:6OJL3XIL submitted 2019-09-05 cs.IR

classification cs.IR
keywords hyperparametertuningcross-validationdiscriminativepowerrecommendersystemsnDCGprecisionBPR-MFnoveltymetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks which evaluation metric best distinguishes between good and bad hyperparameter settings when tuning recommender systems by cross-validation. Using 10-fold cross-validation on MovieLens-1M and Amazon Movies, it compares three models—User-kNN, Item-kNN, and BPR-MF—across grids of hyperparameters and measures accuracy and novelty metrics. The authors find that nDCG@N and Precision@N are the most discriminative accuracy metrics, meaning they most reliably detect significant performance differences between configurations. They also find that for BPR-MF the number of latent factors dominates the other hyperparameters, so tuning effort is best spent there. The paper proposes a general procedure for isolating whether a single hyperparameter matters for a given model.

What carries the argument

Discriminative Power (DP): for a given metric, the sum of p-values over paired statistical tests (Student's t-test) between hyperparameter configurations, averaged across cross-validation folds; a lower DP value means the metric is more sensitive to hyperparameter changes. The paper's per-parameter variant fixes one hyperparameter (e.g., latent factors) and randomly pairs the remaining dimensions, producing a p-value curve whose DP reveals whether that fixed parameter's values drive accuracy differences.

What would settle it

Recompute DP for BPR-MF on MovieLens using all hyperparameter pairs from the full 3150-configuration grid (rather than 25 random pairs) and on the full Amazon grid (rather than the nDCG-selected sub-grid); if the metric ranking no longer places nDCG/Precision first, the paper's conclusion is an artifact of the sampling.

Watch

Extended reading notes

Core claim

The central discovery is that Discriminative Power (DP), computed from p-values of paired statistical tests over hyperparameter pairs, ranks nDCG@N and Precision@N as the strongest accuracy metrics for model selection among the six studied. These metrics show lower DP (sharper separation) than Recall and MRR across nearly all model-dataset combinations, and this ordering survives adding a standard-deviation band across folds. For BPR-MF, fixing each hyperparameter in turn and randomizing the others shows that the number of latent factors produces the smallest DP, indicating that variations in latent-factor count cause the most significant accuracy differences; the paper concludes that this parameter dominates the number of iterations and learning rate.

Load-bearing premise

The conclusions rest on the representativeness of 25 randomly chosen hyperparameter pairs per model and, for Amazon, on a grid sub-selected using nDCG@10; if those choices are unrepresentative, the ranking of metrics could change.

Editorial extensions

If this is right

  • Practitioners can use nDCG@N or Precision@N as the primary selection metric when tuning neighborhood-based models or BPR-MF, instead of averaging many metrics.
  • When tuning BPR-MF, prioritizing the number of latent factors in the search (e.g., more granular values) over iterations and learning rate should yield most of the accuracy gain.
  • Novelty metrics (EFD, EPC) are also highly sensitive to hyperparameter changes, so tuning for accuracy alone may unintentionally drive novelty in ways a combined objective would avoid.
  • The per-parameter DP procedure offers a cheap pre-screening step: before a full grid search, identify which hyperparameters actually matter for the model and dataset.
  • The metric-DP comparisons across folds show that the best accuracy metrics also have acceptable fold-to-fold stability, supporting their use in k-fold cross-validation pipelines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the same DP analysis were run on non-movie domains (news, e-commerce, social), the dominance of nDCG/Precision might not hold; the paper's conclusion is demonstrated only on two movie datasets.
  • The Amazon sub-grid was pre-selected using nDCG@10, which could bias the Amazon comparison in nDCG's favor; a neutral test would select the sub-grid by a different metric or repeat the selection for each metric.
  • The 25 random pairs per model introduce sampling variance; constructing bootstrap confidence intervals around DP values would test whether the metric ordering is statistically stable.
  • The latent-factor dominance result suggests a two-stage tuning strategy: coarse search over latent factors first, then fine-tuning of learning rate and iterations, which the paper does not explicitly recommend but follows directly from its findings.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies how evaluation metrics behave when used to select hyperparameters in cross-validation, and whether some hyperparameters are more important than others. The authors compute the Discriminative Power (DP) of nDCG, Precision, Recall, MRR, EFD, and EPC over grid-search configurations of User-kNN, Item-kNN, and BPR-MF on MovieLens-1M and Amazon Movies. They report that nDCG@N and Precision@N are the most discriminative accuracy metrics, that novelty metrics are also sensitive to hyperparameter changes, and that for BPR-MF the number of latent factors dominates the number of iterations and the learning rate. The paper proposes a procedure for isolating a single hyperparameter in models with several parameters and reports DP tables per hyperparameter value.

Significance. If the results were supported, the paper would make a useful practical contribution to recommender-system evaluation: it would guide practitioners in choosing a metric for cross-validation-based tuning and in prioritizing which BPR-MF hyperparameters to spend compute on. The work builds on an established DP framework from prior work, uses public datasets and standard libraries, and asks a well-motivated question. The empirical claims are specific and falsifiable. However, the significance is moderate because the findings rest on two datasets and on sampling choices that are not fully controlled, and because the central claim about latent-factor dominance appears not to follow from the reported analysis. No code or experimental seeds are provided, which limits reproducibility of the randomness-dependent steps.

major comments (3)
  1. [Section 5 (Table 5)] The conclusion that the number of latent factors is the dominant hyperparameter for BPR-MF does not follow from the reported procedure. The text states that for the latent-factor analysis, each pair of systems is selected to "share the number of latent factors, and differ in the number of iterations and learning rate." Thus a low DP value for a fixed latent-factor row means that variations in iterations and learning rate produce statistically significant metric differences while the latent-factor count is held constant; it is not a measurement of the effect of changing the latent-factor count. To establish dominance one would need to hold iterations and learning rate fixed, vary only the latent factors, and compare that conditional DP with the analogous conditional DPs for the other two parameters. The manuscript itself warns that "the results between the two tables are not comparable," yet the conclusion selects the best DP values across the three hyperparameter analyses. Since this dominance claim appears in the abstract, introduction, and conclusions, the paper's stated contribution (iii) is unsupported by the analysis as written.
  2. [Section 3 (Table 3)] The DP values in Table 3 are computed from a single random draw of 25 pairs of hyperparameter configurations, with no seed, no repeated sampling, and no reported variance. For BPR-MF the MovieLens grid has 15×15×14 = 3,150 configurations, so the pairing universe is enormous, and 25 pairs is a very small sample. Several adjacent DP values in Table 3 are close (e.g., Item-kNN MovieLens nDCG@N 1.710 vs MRR@N 1.766; BPR-MF MovieLens nDCG@N 0.594 vs Precision@N 0.585), so the ranking could plausibly change under another draw. The central claim that nDCG@N and Precision@N are the most discriminative metrics rests on this table. The authors should either use all pairs, or report confidence intervals or repeated random draws, and must at minimum fix and state the random seed.
  3. [Section 2 (Amazon sub-grid)] The Amazon experiments are run on a sub-grid selected by the nDCG@10 value of each hyperparameter plus two neighboring grid values. This makes the comparison across metrics unfair: the configurations are chosen because they are promising according to nDCG@10, so nDCG is evaluated on a grid tailored to its own preferences while the other metrics are evaluated on a grid selected by nDCG. This selection bias affects the Amazon portion of Table 3 and the Amazon portion of Table 5, and it also makes the Section 5 statement that the DP values show the sub-grid choice was "a reasonable choice" circular. A metric-independent subsampling rule, or a separate sub-grid selection for each metric with a corresponding correction, is needed before the Amazon metric ranking can be accepted.
minor comments (4)
  1. [Section 2] There are small text errors: "Intel Xenon" should be "Intel Xeon," and "0,2000038948" uses a comma as a decimal separator, which is inconsistent with the rest of the paper.
  2. [Section 3] The sentence "we generate all possible combinations of pairs of hyperparameters and we randomly take 25 combinations" is ambiguous: it is unclear whether the 25 pairs are drawn once per dataset/model/metric, and whether the same pairs are used for all metrics. Please clarify the sampling protocol.
  3. [Section 4 (Table 4)] The column header "Best + Std Dev" is not self-explanatory. It should state explicitly whether the standard deviation is added to each ordered p-value before computing the DP sum, or to the final DP value; the current text describes the former but the table caption does not.
  4. [Table 5] Table 5 is difficult to read because the MovieLens blocks contain 15, 14, and 15 values while the Amazon blocks contain only 5, 3, and 3 values, with no visual separator between datasets. A panel split or separate rows for the two datasets would greatly improve clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the paper reports direct empirical measurements and does not fit parameters and rename them as predictions.

full rationale

The paper's central claims are empirical measurements: discriminative power (DP) is computed from p-values of paired statistical tests over hyperparameter configurations, and metric rankings are direct observations rather than outputs of a fitted model. The DP methodology is explicitly attributed to prior work by other authors (Valcarce et al., [44]) and is not a self-citation or a uniqueness theorem imported from the authors' own prior work. There is no step where a quantity is defined in terms of the conclusion, and no equation reduces to its own input by construction. The Amazon sub-grid is selected using nDCG@10 before comparing metrics; this is a potential selection-bias concern for the metric-ranking claim, but it is not circularity because all metrics are then independently evaluated on the same selected configurations and no metric is defined in terms of another. The Section 5 'dominant hyperparameter' conclusion is weakened by the fact that the latent-factor analysis holds latent factors fixed and varies the other two parameters, so low DP for latent-factor rows does not directly measure sensitivity to latent factors; however, this is an inferential or validity problem, not a circularity, since the DP values are not constructed from the dominance conclusion. No load-bearing self-citation is present, and no fitted input is presented as a prediction. The honest finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper has no fitted model parameters; the free parameters are experimental choices such as the relevance threshold tau=4 and the grid ranges, which are not fitted to optimize the claim. However, the analysis leans on statistical assumptions (t-test validity) and on the representativeness of random pair sampling and the nDCG-selected sub-grid.

assumptions (3)
  • domain assumption Student's t-test is valid for per-user metric distributions that are typically non-normal.
    The paper uses paired Student's t-tests to compare metrics across user-level values without checking normality or using a nonparametric alternative. This appears in Section 2 under Evaluation protocol.
  • domain assumption The Discriminative Power computed on 25 randomly chosen pairs approximates the true DP over all pairs.
    Section 3 states 'we randomly take 25 combinations' and then computes DP from these pairs. The paper provides no stability analysis, seed, or confidence interval for this estimate.
  • ad hoc to paper The nDCG-based sub-grid selection on Amazon is representative for comparing the discriminative power of all metrics.
    In Section 2, the Amazon grid is restricted to the best nDCG@10 value plus or minus two neighbors. This choice was made for computational feasibility but can bias the relative DP of metrics, and no correction is applied.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the discriminative power of Hyper-parameters in Cross-Validation and how to choose them." pith.science (2026). https://pith.science/paper/6OJL3XIL

@misc{pith2026190902523,
  author       = {Pith},
  title        = {Pith review of: On the discriminative power of Hyper-parameters in Cross-Validation and how to choose them},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6OJL3XIL}},
  note         = {Machine review of arXiv:1909.02523}
}
read the original abstract

Hyper-parameters tuning is a crucial task to make a model perform at its best. However, despite the well-established methodologies, some aspects of the tuning remain unexplored. As an example, it may affect not just accuracy but also novelty as well as it may depend on the adopted dataset. Moreover, sometimes it could be sufficient to concentrate on a single parameter only (or a few of them) instead of their overall set. In this paper we report on our investigation on hyper-parameters tuning by performing an extensive 10-Folds Cross-Validation on MovieLens and Amazon Movies for three well-known baselines: User-kNN, Item-kNN, BPR-MF. We adopted a grid search strategy considering approximately 15 values for each parameter, and we then evaluated each combination of parameters in terms of accuracy and novelty. We investigated the discriminative power of nDCG, Precision, Recall, MRR, EFD, EPC, and, finally, we analyzed the role of parameters on model evaluation for Cross-Validation.

Figures

Figures reproduced from arXiv: 1909.02523 by the authors.

Figure 1
Figure 1. Discriminative Power of Accuracy and Novelty metrics [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 45 canonical work pages

  1. [1]

    Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio, Azzurra Ragone, and Joseph Trotta. 2019. The importance of being dissimilar in Recommendation. to appear

  2. [2]

    Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio, Azzurra Ragone, and Joseph Trotta. 2019. Local Popularity and Time in top-N Recommendation. In Advances in Information Retrieval - 41st European Conf. on IR Research, ECIR 2019, Cologne, Germany, April 14-18, 2019, Proc., Part I . 861–868

  3. [3]

    Alejandro Bellogín and Pablo Sánchez. 2017. Revisiting Neighbourhood-Based Recommenders For Temporal Scenarios. In Proc. of the 1st Workshop on Temporal Reasoning in Recommender Systems co-located with 11th Int. Conf. on Recommender Systems (RecSys 2017), Como, Italy, August 27-31, 2017. 40–44

  4. [4]

    James Bennett, Stan Lanning, et al. 2007. The netflix prize. In Proc. of KDD cup and workshop, Vol. 2007. New York, NY, USA., 35

  5. [5]

    James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011. Algorithms for Hyper-Parameter Optimization. In Advances in Neural Information Processing Systems 24: 25th Annual Conf. on Neural Information Processing Systems 2011. Proc. of a meeting held 12-14 December 2011, Granada, Spain. 2546–2554

  6. [6]

    James Bergstra and Yoshua Bengio. 2012. Random Search for Hyper-Parameter Optimization. Journal of Machine Learning Research 13 (2012), 281–305

  7. [7]

    James Bergstra, Brent Komer, Chris Eliasmith, Dan Yamins, and David D Cox. 2015. Hyperopt: a python library for model selection and hyperparameter optimization. Computational Science & Discovery 8, 1 (2015), 014008

  8. [8]

    James Bergstra, Dan Yamins, and David D Cox. 2013. Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms. In Proc. of the 12th Python in science conf. Citeseer, 13–20

Show all 49 references
  1. [9]

    An- derson

    Alex Beutel, Ed Huai-hsin Chi, Zhiyuan Cheng, Hubert Pham, and John R. An- derson. 2017. Beyond Globally Optimal: Focused Learning for Improved Recom- mendations. In Proc. of the 26th Int. Conf. on World Wide Web, WWW 2017, Perth, Australia, April 3-7, 2017. 203–212

  2. [10]

    Cora, and Nando de Freitas

    Eric Brochu, Vlad M. Cora, and Nando de Freitas. 2010. A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. CoRR abs/1012.2599 (2010)

  3. [11]

    Campos, Fernando Díez, and Iván Cantador

    Pedro G. Campos, Fernando Díez, and Iván Cantador. 2014. Time-aware recom- mender systems: a comprehensive survey and analysis of existing evaluation protocols. User Model. User-Adapt. Interact. 24, 1-2 (2014), 67–119

  4. [12]

    Hurley, and Saul Vargas

    Pablo Castells, Neil J. Hurley, and Saul Vargas. 2015. Novelty and Diversity in Recommender Systems. See [35], 881–918

  5. [13]

    Hurley, and Saul Vargas

    Pablo Castells, Neil J. Hurley, and Saul Vargas. 2015. Novelty and Diversity in Recommender Systems. Springer US, Boston, MA, 881–918

  6. [14]

    Treleaven, and Licia Capra

    Simon Chan, Philip C. Treleaven, and Licia Capra. 2013. Continuous hyperpa- rameter optimization for large-scale recommender systems. In Proc. of the 2013 IEEE Int. Conf. on Big Data, 6-9 October 2013, Santa Clara, CA, USA . 350–358

  7. [15]

    Paolo Cremonesi, Franca Garzotto, Sara Negro, Alessandro Vittorio Papadopoulos, and Roberto Turrin. 2011. Looking for "Good" Recommendations: A Comparative Evaluation of Recommender Systems. In Human-Computer Interaction - INTER- ACT 2011 - 13th IFIP TC 13 Int. Conf., Lisbon, ...

  8. [16]

    Paolo Cremonesi, Yehuda Koren, and Roberto Turrin. 2010. Performance of recommender algorithms on top-n recommendation tasks. In Proc. of the 2010 ACM Conf. on Recommender Systems, RecSys 2010, Barcelona, Spain, September 26-30, 2010. 39–46

  9. [17]

    Paolo Cremonesi, Roberto Turrin, Eugenio Lentini, and Matteo Matteucci. 2008. An evaluation methodology for collaborative recommender systems. In 2008 Int. Conf. on Automated Solutions for Cross Media Content and Multi-Channel Distribution. IEEE, 224–231

  10. [18]

    Bruno Feres de Souza, André Carlos Ponce de Leon Ferreira de Carvalho, Rodrigo Calvo, and Renato Porfirio Ishii. 2006. Multiclass SVM Model Selection Using Particle Swarm Optimization. In 6th Int. Conf. on Hybrid Intelligent Systems (HIS 2006), 13-15 December 2006, Auckland, N...

  11. [19]

    Frauke Friedrichs and Christian Igel. 2005. Evolutionary tuning of multiple SVM parameters. Neurocomputing 64 (2005), 107–117

  12. [20]

    Herlocker, Joseph A

    Jonathan L. Herlocker, Joseph A. Konstan, Loren G. Terveen, and John Riedl

  13. [21]

    Neil Hurley and Mi Zhang. 2011. Novelty and Diversity in Top-N Recommenda- tion – Analysis and Evaluation. ACM Trans. Internet Technol. 10 (March 2011), 14:1–14:30. Issue 4

  14. [22]

    Hoos, and Kevin Leyton-Brown

    Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown. 2014. An Efficient Approach for Assessing Hyperparameter Importance. In Proc. of the 31th Int. Conf. on Machine Learning, ICML 2014, Beijing, China, 21-26 June 2014 (JMLR Workshop and Conf. Proc.), Vol. 32. JMLR.org, 754–762

  15. [23]

    Frank Hutter, Jörg Lücke, and Lars Schmidt-Thieme. 2015. Beyond Manual Tuning of Hyperparameters. KI 29, 4 (2015), 329–337

  16. [24]

    Legenstein

    Michael Jahrer, Andreas Töscher, and Robert A. Legenstein. 2010. Combining predictions for accurate recommender systems. In Proc. of the 16th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining, Washington, DC, USA, July 25-28, 2010, Bharat Rao, Balaji Krishnapuram, A...

  17. [25]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. 20, 4 (2002), 422–446

  18. [26]

    Ron Kohavi. 1995. A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection. In Proc. of the Fourteenth Int. Joint Conf. on Artificial Intelligence, IJCAI 95, Montréal Québec, Canada, August 20-25 1995, 2 Volumes. Morgan Kaufmann, 1137–1145

  19. [27]

    Nathan Nan Liu and Qiang Yang. 2008. EigenRank: a ranking-oriented approach to collaborative filtering. In Proc. of the 31st Annual Int. ACM SIGIR Conf. on Research and Development in Information Retrieval, SIGIR 2008, Singapore, July 20-24, 2008, Sung-Hyon Myaeng, Douglas W. ...

  20. [28]

    Shane Culpepper

    Xiaolu Lu, Alistair Moffat, and J. Shane Culpepper. 2016. The effect of pooling and evaluation depth on IR metrics. Inf. Retr. Journal 19, 4 (2016), 416–445

  21. [29]

    Xin Luo, MengChu Zhou, Shuai Li, Zhu-Hong You, Yunni Xia, and Qingsheng Zhu. 2016. A Nonnegative Latent Factor Model for Large-Scale Sparse Matrices in Recommender Systems via Alternating Direction Method. IEEE Trans. Neural Netw. Learning Syst. 27, 3 (2016), 579–592

  22. [30]

    Gurmeet Singh Manku, Sridhar Rajagopalan, and Bruce G. Lindsay. 1999. Random Sampling Techniques for Space Efficient Online Computation of Order Statistics of Large Datasets. In SIGMOD 1999, Proc. ACM SIGMOD Int. Conf. on Management of Data, June 1-3, 1999, Philadelphia, Penns...

  23. [31]

    Julian John McAuley and Jure Leskovec. 2013. From amateurs to connoisseurs: modeling the evolution of user expertise through online reviews. In 22nd Int. World Wide Web Conf., WWW ’13, Rio de Janeiro, Brazil, May 13-17, 2013. 897–908

  24. [32]

    McNee, John Riedl, and Joseph A

    Sean M. McNee, John Riedl, and Joseph A. Konstan. 2006. Being Accurate is Not Enough: How Accuracy Metrics Have Hurt Recommender Systems. In CHI ’06 Extended Abstracts on Human Factors in Computing Systems (CHI EA ’06) . 1097–1101

  25. [33]

    DiCarlo, and David D

    Nicolas Pinto, David Doukhan, James J. DiCarlo, and David D. Cox. 2009. A High- Throughput Screening Approach to Discovering Good Forms of Biologically Inspired Visual Representation. PLoS Computational Biology 5, 11 (2009)

  26. [34]

    Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme

  27. [35]

    Francesco Ricci, Lior Rokach, and Bracha Shapira (Eds.). 2015. Recommender Systems Handbook. Springer

  28. [36]

    Tetsuya Sakai. 2006. Evaluating evaluation metrics based on the bootstrap. In SIGIR 2006: Proc. of the 29th Annual Int. ACM SIGIR Conf. on Research and Development in Information Retrieval, Seattle, Washington, USA, August 6-11, 2006 . 525–532

  29. [37]

    Konstan, and John Riedl

    Badrul Munir Sarwar, George Karypis, Joseph A. Konstan, and John Riedl. 2000. Analysis of recommendation algorithms for e-commerce. In EC. 158–167

  30. [38]

    Guy Shani and Asela Gunawardana. 2011. Evaluating Recommendation Systems. In Recommender Systems Handbook. 257–297

  31. [39]

    Smith, Logan Mitchell, Christophe G

    Michael R. Smith, Logan Mitchell, Christophe G. Giraud-Carrier, and Tony R. Martinez. 2014. Recommending Learning Algorithms and Their Associated Hy- perparameters. In Proc. of the Int. Workshop on Meta-learning and Algorithm Selec- tion co-located with 21st European Conf. on ...

  32. [40]

    Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. 2012. Practical Bayesian Optimization of Machine Learning Algorithms. InAdvances in Neural Information Processing Systems 25: 26th Annual Conf. on Neural Information Processing Systems

  33. [41]

    Harald Steck. 2013. Evaluation of recommendations: rating-prediction and rank- ing. In Seventh ACM Conf. on Recommender Systems, RecSys ’13, Hong Kong, China, October 12-16, 2013, Qiang Yang, Irwin King, Qing Li, Pearl Pu, and George Karypis (Eds.). ACM, 213–220

  34. [42]

    Gábor Takács, István Pilászy, Bottyán Németh, and Domonkos Tikk. 2008. Matrix factorization and neighbor based algorithms for the netflix prize problem. In Proc. of the 2008 ACM Conf. on Recommender Systems, RecSys 2008, Lausanne, Switzerland, October 23-25, 2008 . 267–274

  35. [43]

    Gábor Takács, István Pilászy, Bottyán Németh, and Domonkos Tikk. 2009. Scal- able Collaborative Filtering Approaches for Large Recommender Systems.Journal of Machine Learning Research 10 (2009), 623–656

  36. [44]

    Daniel Valcarce, Alejandro Bellogín, Javier Parapar, and Pablo Castells. 2018. On the Robustness and Discriminative Power of Information Retrieval Metrics for top-N Recommendation. In Proc. of the 12th ACM Conf. on Recommender Systems (RecSys ’18). ACM, New York, NY, USA, 260–268

  37. [45]

    Saul Vargas and Pablo Castells. 2011. Rank and relevance in novelty and diversity metrics for recommender systems. InProc. of the 2011 ACM Conf. on Recommender Systems, RecSys 2011, Chicago, IL, USA, October 23-27, 2011 . 109–116

  38. [46]

    Wilkinson, Robert Schreiber, and Rong Pan

    Yunhong Zhou, Dennis M. Wilkinson, Robert Schreiber, and Rong Pan. 2008. Large-Scale Parallel Collaborative Filtering for the Netflix Prize. In Algorithmic Aspects in Information and Management, 4th Int. Conf., AAIM 2008, Shanghai, China, June 23-25, 2008. Proc. 337–348

  39. [2004]

    ACM Trans

    Evaluating collaborative filtering recommender systems. ACM Trans. Inf. Syst. 22, 1 (2004), 5–53

  40. [2009]

    In UAI 2009, Proc

    BPR: Bayesian Personalized Ranking from Implicit Feedback. In UAI 2009, Proc. of the Twenty-Fifth Conf. on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, June 18-21, 2009 . 452–461

  41. [2012]

    2960–2968

    December 3-6, 2012, Lake Tahoe, Nevada, United States. 2960–2968

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.