REVIEW 5 major objections 5 minor 58 references
Score Combining for Contrastive OOD Detection
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A generalized likelihood-ratio test derived from a negative-means model of z-scored inlier scores outperforms heuristic score-combining rules for contrastive out-of-distribution detection.
desk verdict Clean GLRT-based score combining for contrastive OOD detection with an honest writeup, but the claim of beating Fisher/Stouffer rests on sub-0.01 AUROC gaps with no error bars. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the negative-means (NM) hypothesis-testing model and its generalized likelihood ratio test statistic. The model (equation 15) says that inlier z-values are independent standard normals while OOD z-values are normals with every coordinate mean at most $-\epsilon$; the GLRT statistic $t_{\mathrm{GLRT}}(z) = (\tfrac{1}{2} z^{-} - z)^T z^{-}$ with $z^{-}_l = \min\{z_l, -\epsilon\}$ is the log-likelihood ratio after replacing the unknown alternative mean by its constrained maximizer. The machinery also includes the empirical CDF transform $\hat{z}_l(x) = \Phi^{-1}(\hat{F}_l(s_l(x)))$ that converts arbitrary score distributions to z-values, and a conformal p-value threshold (Algorithm 3) that certifies a user-specified false-alarm rate up to a chosen failure probability using a validation set.
What would settle it
Measure the off-diagonal correlations of the empirical z-value vectors on a held-out inlier set; if a GLRT built with the estimated inlier covariance (Appendix A.2) does not beat or match the identity-covariance GLRT on the same novelty sets, the independence assumption is doing essential work rather than being a harmless device.
Extended reading notes
Core claim
The paper's central claim is that score combining for contrastive OOD detection should be posed as a negative-means hypothesis test on z-values rather than as a heuristic sum, max, or weighted average. Under the null hypothesis $H_0 : Z \sim \mathcal{N}(0, I)$, each empirical z-value is standard normal; under the alternative $H_1 : Z \sim \mathcal{N}(\mu, I)$ with $\mu_l \le -\epsilon$ for every $l$, an OOD sample is, on average, worse on every score. Solving this composite testing problem with a GLRT yields the statistic $t_{\mathrm{GLRT}}(z) = (\tfrac{1}{2} z^- - z)^T z^-$, where $z^-_l = \min\{z_l, -\epsilon\}$. The paper claims that thresholding this statistic outperforms the heuristic combining rules of CSI and SupCSI on the tested image benchmarks, and also outperforms Fisher, Bonferroni, Simes/BH, Stouffer, and ALR p-value combiners.
Load-bearing premise
The derivation assumes the z-values of an inlier are independent standard normals; the paper says this independence does not hold, so if the scores are strongly correlated, the GLRT is no longer a likelihood-ratio test for the true distribution.
Editorial extensions
If this is right
- New inlier scores can be added to a contrastive detector by appending their empirical z-values, since the GLRT is agnostic to how the scores were produced; glrt-SupCSI+ combines 24 scores and improves dataset-vs-dataset AUROC over glrt-SupCSI.
- The GLRT removes the need to hand-tune CSI's weighting coefficients $\lambda^{\mathrm{con}}_j$ and $\lambda^{\mathrm{shift}}_j$, because every score enters through its own z-value.
- A validation set of size $v$ gives a threshold such that, with probability at least $1-\delta$, the false-alarm rate is at most $\alpha$, via the beta-distributed conformal p-value described in Section 3.3.
- A positive offset $\epsilon$ (the paper uses 0.25) improves AUROC over the $\epsilon = 0$ cone problem.
- The benefit of adding extra base scores is not guaranteed: glrt-SupCSI+ underperforms glrt-SupCSI on leave-one-class-out CIFAR-10, so base-score selection remains part of the design.
Reading between the lines
- Editorial inference: the average AUROC gap between the GLRT and the best classical combiners (Fisher, Stouffer) is well under a percentage point, while Bonferroni, Simes/BH, and ALR are much further behind, which suggests the GLRT's practical value is stability across datasets rather than a large leap over a well-chosen classical merger.
- Editorial inference: because the independence assumption in (15) is false, a natural stress test is to compare the GLRT against a covariance-aware GLRT on benchmarks with strongly correlated scores; the paper's appendix finds the estimated-covariance variant underperforms, hinting that the identity covariance acts as a helpful inductive bias.
- Editorial inference: the conformal false-alarm machinery of Section 3.3 does not depend on the GLRT statistic itself, so any score combiner could inherit the same finite-sample threshold guarantee by reusing Algorithm 3.
- Editorial inference: since the pipeline only sees marginal z-values, it could fuse heterogeneous evidence (classifier logits, nearest-neighbor distances, density estimates) in other one-class and anomaly-detection tasks, not just contrastive image representations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes to replace the heuristic score-combining rules used by the CSI and SupCSI contrastive OOD detectors (Eqs. (6) and (8)) with a generalized likelihood ratio test (GLRT) derived from a negative-means model on z-values obtained from an empirical CDF transform of the base scores (Eqs. (15) and (20)). The authors also give a conformal-prediction procedure that, using a validation set, yields a false-alarm-rate guarantee (Algorithms 2 and 3). Experiments on dataset-vs-dataset tasks with CIFAR-10, SVHN, LSUN, ImageNet, and CIFAR-100, and on leave-one-class-out tasks with CIFAR-10, claim that the GLRT combination outperforms the original CSI/SupCSI heuristics and the classical Fisher, Bonferroni, Simes/BH, Stouffer, and ALR combiners.
Significance. The statistical derivation of the GLRT statistic is sound and the proposed framework is a principled alternative to ad-hoc ensembling of contrastive scores. The false-alarm-rate guarantee in Section 3.3 is a useful contribution, as is the appendix's explicit examination of the independence assumption through a general-covariance variant. However, the central empirical claim of outperformance is currently not statistically established: the reported gains over Fisher and Stouffer are very small in AUROC and detection rate, the experiments appear to be single-run, and the single tuned parameter epsilon is selected using one of the evaluation tasks. If the empirical claim could be supported by repeated-seed results with proper selection of epsilon, the paper would be a genuinely useful contribution to OOD detection methodology.
major comments (5)
- [Section 4.2, Figure 4] The free parameter epsilon in Eq. (15) is set to 0.25 by tuning on the CIFAR-10-inliers vs CIFAR-100-novelties task, which is one of the evaluation tasks reported in Tables 2, 4, 6, and 8. This leaks OOD information into the method and makes the reported improvements over eps=0 and over the classical combiners hard to interpret. The authors should either choose epsilon on a held-out validation split that does not use any evaluation OOD dataset, or report the main comparisons for a range of epsilon values, including epsilon=0, to show that the conclusions are not artifacts of this choice.
- [Tables 6-9 and 13-14] No error bars, confidence intervals, number of seeds, or paired significance tests are reported. The average AUROC differences between GLRT and Fisher/Stouffer are as small as 0.0002-0.009 (e.g., Table 6 avg 0.9677 vs 0.9670 for Fisher and Stouffer; Table 7 avg 0.9928 vs 0.9926/0.9916), and the DR differences are 0.1-1.0 percentage points. Given the stochasticity of contrastive training, these differences are within plausible run-to-run noise. The central claim that the GLRT outperforms Fisher and Stouffer therefore needs repeated-seed experiments with variance reporting or paired tests before it can be accepted.
- [Section 3.1, Eq. (15)] The authors explicitly state that the independence assumption in Eq. (15) is not expected to hold in practice. The GLRT derivation and its Neyman-Pearson/minimax optimality rely on this assumption, while the actual z-values are derived from correlated scores. Appendix A.2 partially addresses this by substituting a sample covariance, but it only compares identity versus sample covariance on the dataset-vs-dataset AUROC tables and finds that identity performs better. This does not resolve the misspecification concern for correlated non-Gaussian z-values. The authors should either temper the optimality language to describe the GLRT as a heuristic score combination that is principled only under the NM model, or provide additional diagnostics on the actual distribution of the z-vectors.
- [Section 4.3, Tables 8-9] The claim that the GLRT outperforms all classical combiners is not uniformly supported by the tables. In Table 8, Stouffer achieves an average DR of 86.3 compared to 86.2 for GLRT, and in Tables 8-9 GLRT and Fisher tie on average. The text in Section 4.3 acknowledges this, but the abstract and conclusion state the outperformance claim without this qualification. The claims should be restated as 'on average over AUROC' and should acknowledge that the advantage over Fisher/Stouffer is not consistent across all metrics and tasks.
- [Section 3.3, Eq. (24)] The false-alarm-rate guarantee is stated as holding 'with probability at least 1-delta' over the random validation set, and the procedure in Algorithm 2 searches for a threshold a that controls the beta upper tail. This is a valid conformal-style argument, but the guarantee is average over the validation set under the assumption that Xval is drawn i.i.d. under H0. In the experiments, no results using this finite-sample guarantee are reported; all DR-versus-FAR results instead use a threshold computed directly on the test set. The paper should either include experiments demonstrating the conformal procedure, or explicitly state that the guarantee is a methodological contribution not evaluated empirically.
minor comments (5)
- [Abstract and Section 1] The name 'Benjamini-Hochwald' appears in the abstract and contribution list; this should be 'Benjamini-Hochberg'.
- [Algorithm 3] In Algorithm 3, the line 'return “OOD” if bq(x) ≤ a' is preceded by 'white guaranteeing a false-alarm rate'; this should read 'while guaranteeing'.
- [Section 5] The conclusion refers to 'the non-negative means problem (15)', but Eq. (15) is the negative-means problem with means ≤ -ϵ; the wording should be corrected for consistency with Section 3.1.
- [Figure 2] Figure 2 plots the proposed test statistic for m=1, but the paper does not mention in the caption or text that this is the univariate case; a brief note would avoid confusion.
- [Table 1] The row for glrt-SupCSI+ marks both 'Heuristic' and 'GLRT' columns as applicable; it would be clearer to state explicitly that the base scores themselves are computed with the SupCSI/OC-SVM/Mahalanobis methods while the combining is GLRT-based.
Circularity Check
No circularity found: GLRT statistic is a standard derivation from an explicit model, comparisons are external benchmarks, and the disclosed ε selection is a hyper-parameter choice, not a fitted prediction.
full rationale
The derivation is self-contained and the empirical claims are not equivalent to their inputs. The empirical z-values in (21) are probability-integral transforms of the inlier score CDF, so their marginal null distribution is standard normal by construction; the paper explicitly acknowledges that the independence assumption in (15) is false and probes the misspecification in Appendix A.2 by substituting the sample covariance. The GLRT statistic (20) is obtained by a standard closed-form likelihood-ratio computation from the stated negative-means model (15), with no benchmark result or target AUROC used as an input. The comparisons against CSI/SupCSI and against Fisher, Bonferroni, Simes/BH, Stouffer, and ALR use identical base scores and are external empirical benchmarks, so outperformance is not implied by the model construction. The paper contains no load-bearing self-citations: the only optimality citation (Wei et al., 2019) is external and covers the ε=0 special case, while the ε=0.25 choice is a disclosed hyper-parameter selection made on one evaluation pair; that choice could affect one table row but is not a fitted outcome renamed as a prediction. The acknowledged limitations (which base scores to combine, and dependence on transformation sets) are scope restrictions rather than hidden circularity.
Assumptions & free parameters
free parameters (1)
- epsilon (NM model parameter) =
0.25
assumptions (5)
- domain assumption Under H0, the z-values Z_l = Phi^{-1}(F_l(s_l(X))) are i.i.d. standard normal.
- domain assumption Under H1, each z-value has mean at most -epsilon, i.e., all component means are shifted negative.
- domain assumption The empirical CDF bF_l provides a consistent estimate of the true null CDF, so bz_l is approximately standard normal under H0.
- standard math The conformal p-value calibration in Algorithms 2-3 provides the stated false-alarm-rate bound.
- ad hoc to paper epsilon=0.25 is an appropriate fixed value for all tasks.
Cite this review
Pith. "Pith review of Score Combining for Contrastive OOD Detection." pith.science (2026). https://pith.science/paper/24DSEF2T
@misc{pith2026250112204,
author = {Pith},
title = {Pith review of: Score Combining for Contrastive OOD Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/24DSEF2T}},
note = {Machine review of arXiv:2501.12204}
}
read the original abstract
In out-of-distribution (OOD) detection, one is asked to classify whether a test sample comes from a known inlier distribution or not. We focus on the case where the inlier distribution is defined by a training dataset and there exists no additional knowledge about the novelties that one is likely to encounter. This problem is also referred to as novelty detection, one-class classification, and unsupervised anomaly detection. The current literature suggests that contrastive learning techniques are state-of-the-art for OOD detection. We aim to improve on those techniques by combining/ensembling their scores using the framework of null hypothesis testing and, in particular, a novel generalized likelihood ratio test (GLRT). We demonstrate that our proposed GLRT-based technique outperforms the state-of-the-art CSI and SupCSI techniques from Tack et al. 2020 in dataset-vs-dataset experiments with CIFAR-10, SVHN, LSUN, ImageNet, and CIFAR-100, as well as leave-one-class-out experiments with CIFAR-10. We also demonstrate that our GLRT outperforms the score-combining methods of Fisher, Bonferroni, Simes, Benjamini-Hochwald, and Stouffer in our application.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
F. Ahmed and A. Courville. Detecting semantic anomalies. In Proc. AAAI Conf. Artificial Intell., volume 34, pages 3154--3162, 2020
work page 2020
-
[2]
M. S. Anderson, J. Dahl, and L. Vandenberghe. CVXOPT : A python package for convex optimization, 2012. URL abel.ee.ucla.edu/cvxopt
work page 2012
-
[3]
Y. Benjamini and Y. Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. Roy. Statist. Soc. B, 57 0 (1): 0 289--300, 1995
work page 1995
-
[4]
F. Bergamin, P.-A. Mattei, J. D. Havtorn, H. Senetaire, H. Schmutz, L. Maal e, S. Hauberg, and J. Frellsen. Model-agnostic out-of-distribution detection using combined statistical tests. In Proc. Intl. Conf. Artificial Intell. Statist., pages 10753--10776, 2022
work page 2022
-
[5]
L. Bergman and Y. Hoshen. Classification-based anomaly detection for general data. In Proc. Intl. Conf. Learn. Rep., 2020
work page 2020
-
[6]
R. H. Berk and D. H. Jones. Goodness-of-fit test statistics that dominate the K olmogorov statistics. Zeitschrift f \"u r Wahrscheinlichkeitstheorie und verwandte Gebiete , 47 0 (1): 0 47--59, 1979
work page 1979
-
[7]
C. M. Bishop. Novelty detection and neural network validation. IEE Proc. V: Vision, Image & Signal Process., 141 0 (4): 0 217--222, 1994
work page 1994
-
[8]
M. M. Breunig, H.-P. Kriegel, R. T. Ng, and J. Sander. LOF: I dentifying density-based local outliers. In Proc. ACM SIGMOD Intl. Conf. Manag. Data, pages 93--104, 2000
work page 2000
Show all 58 references
-
[9]
Chalapathy and S
R. Chalapathy and S. Chawla. Deep learning for anomaly detection: A survey. arXiv:1901.03407, 2019
1901 arXiv
-
[10]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In Proc. Intl. Conf. Mach. Learn., pages 1597--1607, 2020
2020
-
[11]
R. D. Cousins. Annotated bibliography of some papers on combining significances or p-values. arXiv:0705.2209, 2007
2007 arXiv
-
[12]
Deecke, L
L. Deecke, L. Ruff, R. Vandermeulen, and H. Bilen. Transfer-based semantic anomaly detection. In Proc. Intl. Conf. Mach. Learn., pages 2546--2558, 2021
2021
-
[13]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. IEEE Conf. Comp. Vision Pattern Recog., pages 248--255, 2009
2009
-
[14]
Di Mattia, P
F. Di Mattia, P. Galeone, M. De Simoni, and E. Ghelfi. A survey on GANs for anomaly detection. arXiv:1906.11632, 2019
1906 arXiv
-
[15]
Donoho and J
D. Donoho and J. Jin. Higher criticism for detecting sparse heterogeneous mixtures. Ann. Statist., 32 0 (3): 0 962--994, 2004
2004
-
[16]
Ericsson, H
L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales. Self-supervised representation learning: I ntroduction, advances, and challenges. IEEE Signal Process. Mag., 39 0 (3): 0 42--62, 2022
2022
-
[17]
R. A. Fisher. Statistical methods for research workers. Springer, 1992
1992
-
[18]
Georgescu, A
M.-I. Georgescu, A. Barbalau, R. T. Ionescu, F. S. Khan, M. Popescu, and M. Shah. Anomaly detection in video via self-supervised and multi-task learning. In Proc. IEEE Conf. Comp. Vision Pattern Recog., pages 12742--12752, 2021
2021
-
[19]
Gidaris, P
S. Gidaris, P. Singh, and N. Komodakis. Unsupervised representation learning by predicting image rotations. In Proc. Intl. Conf. Learn. Rep., Apr. 2018
2018
-
[20]
Golan and R
I. Golan and R. El-Yaniv. Deep anomaly detection using geometric transformations. In Proc. Neural Info. Process. Syst. Conf., volume 31, 2018
2018
-
[21]
Haroush, T
M. Haroush, T. Frostig, R. Heller, and D. Soudry. A statistical framework for efficient out of distribution detection in deep neural networks. In Proc. Intl. Conf. Learn. Rep., 2022
2022
-
[22]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proc. IEEE Conf. Comp. Vision Pattern Recog., pages 770--778, 2016
2016
-
[23]
Hendrycks and K
D. Hendrycks and K. Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In Proc. Intl. Conf. Learn. Rep., 2017
2017
-
[24]
Hendrycks, M
D. Hendrycks, M. Mazeika, and T. Dietterich. Deep anomaly detection with outlier exposure. In Proc. Intl. Conf. Learn. Rep., 2019 a
2019
-
[25]
Hendrycks, M
D. Hendrycks, M. Mazeika, S. Kadavath, and D. Song. Using self-supervised learning can improve robustness and uncertainty. In Proc. Neural Info. Process. Syst. Conf., 2019 b
2019
-
[26]
Hoffmann
H. Hoffmann. Kernel PCA for novelty detection. Pattern Recognition, 40 0 (3): 0 863--874, 2007
2007
-
[27]
R. Kaur, S. Jha, A. Roy, S. Park, E. Dobriban, O. Sokolsky, and I. Lee. iDECODe: I n-distribution equivariance for conformal out-of-distribution detection. In Proc. AAAI Conf. Artificial Intell., volume 36, pages 7104--7114, 2022
2022
-
[28]
Khalid, A
U. Khalid, A. Esmaeili, N. Karim, and N. Rahnavard. RODD: A self-supervised approach for robust out-of-distribution detection. In Proc. IEEE Conf. Comp. Vision Pattern Recog. Workshop, pages 163--170, 2022
2022
-
[29]
Khosla, P
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan. Supervised contrastive learning. In Proc. Neural Info. Process. Syst. Conf., volume 33, pages 18661--18673, 2020
2020
-
[30]
Krizhevsky
A. Krizhevsky. Learning multiple layers of features from tiny images. Technical Report, 2009. URL https://www.cs.toronto.edu/ kriz/learning-features-2009-TR.pdf
2009
-
[31]
L. J. Latecki, A. Lazarevic, and D. Pokrajac. Outlier detection with kernel density functions. In Proc. Mach. Learn. Data Mining Pattern Recog., volume 7, pages 61--75, 2007
2007
-
[32]
Laurikkala, M
J. Laurikkala, M. Juhola, E. Kentala, N. Lavrac, S. Miksch, and B. Kavsek. Informal identification of outliers in medical data. In Intl. Workshop Intell. Data Anal. Med. Pharma., pages 20--24, 2000
2000
-
[33]
K. Lee, K. Lee, H. Lee, and J. Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Proc. Neural Info. Process. Syst. Conf., 2018
2018
-
[34]
E. L. Lehmann, J. P. Romano, and G. Casella. Testing Statistical Hypotheses, volume 3. Springer, 2005
2005
-
[35]
Magesh, V
A. Magesh, V. V. Veeravalli, A. Roy, and S. Jha. Multiple testing framework for out-of-distribution detection. arXiv:2206.09522, 2022
2022 arXiv
-
[36]
Netzer, T
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. In Proc. Neural Info. Process. Syst. Conf., 2011
2011
-
[37]
G. Pang, C. Shen, L. Cao, and A. V. D. Hengel. Deep learning for anomaly detection: A review. ACM Comput. Surveys, 54 0 (2): 0 1--38, 2021
2021
-
[38]
Ramaswamy, R
S. Ramaswamy, R. Rastogi, and K. Shim. Efficient algorithms for mining outliers from large data sets. In Proc. ACM SIGMOD Intl. Conf. Manag. Data, pages 427--438, 2000
2000
-
[39]
Reiss and Y
T. Reiss and Y. Hoshen. Mean-shifted contrastive loss for anomaly detection. In Proc. AAAI Conf. Artificial Intell., Washington, DC, Feb. 2023
2023
-
[40]
Reiss, N
T. Reiss, N. Cohen, L. Bergman, and Y. Hoshen. PANDA: A dapting pretrained features for anomaly detection and segmentation. In Proc. IEEE Conf. Comp. Vision Pattern Recog., page 2806–2814, virtual, June 2021
2021
-
[41]
o rnitz, A. Binder, E. M\
L. Ruff, R. Vandermeulen, N. G\" o rnitz, A. Binder, E. M\" u ller, K. M\" u ller, and M. Kloft. Transfer-based semantic anomaly detection. In Proc. Intl. Conf. Learn. Rep., 2020
2020
-
[42]
L. Ruff, J. R. Kauffmann, and R. A. Vandermeulen. A unifying review of deep and shallow anomaly detection. Proc. IEEE, 109 0 (5): 0 756--795, 2021
2021
-
[43]
Sch\" o lkopf, J
B. Sch\" o lkopf, J. C. Platt, J. C. Shawe-Taylor, A. J. Smola, and R. C. Williamson. Estimating the support of a high-dimensional distribution. Neural Comput., 13 0 (7): 0 1443--1471, July 2001
2001
-
[44]
R. J. Simes. An improved B onferroni procedure for multiple tests of significance. Biometrika, 73 0 (3): 0 751--754, 1986
1986
-
[45]
L. N. Smith and N. Topin. Super-convergence: V ery fast training of neural networks using large learning rate. In T. Pham, editor, Artificial Intell. Machine Learning for Multi-Domain Operations Applications, volume 11006, page 11006 12. SPIE, 2019
2019
-
[46]
Sohn, C.-L
K. Sohn, C.-L. Li, J. Yoon, M. Jin, and T. Pfister. Learning and evaluating representations for deep one-class classification. In Proc. Intl. Conf. Learn. Rep., 2021
2021
-
[47]
Szegedy, V
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In Proc. IEEE Conf. Comp. Vision Pattern Recog., 2016
2016
-
[48]
J. Tack, S. Mo, J. Jeong, and J. Shin. CSI: N ovelty detection via contrastive learning on distributionally shifted instances. In Proc. Neural Info. Process. Syst. Conf., pages 11839--11852, 2020
2020
-
[49]
Tax and R
D. Tax and R. Duin. Support vector data description. Mach. Learn., 54: 0 45--66, 2004
2004
-
[50]
V. Vovk. Conditional validity of inductive conformal predictors. In Asian Conf. Mach. Learn., pages 475--490, 2012
2012
-
[51]
G. Walther. The average likelihood ratio for large-scale multiple testing and detecting sparse mixtures. Inst. Math. Stat.(IMS) Collect, 9: 0 317--326, 2013
2013
-
[52]
Wang and P
T. Wang and P. Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proc. Intl. Conf. Mach. Learn., pages 9929--9939, 2020
2020
-
[53]
Wasserman
L. Wasserman. All of Statistics: A Concise Course in Statistical Inference. Springer, 2004
2004
-
[54]
Y. Wei, M. J. Wainwright, and A. Guntuboyina. The geometry of hypothesis testing over convex cones: G eneralized likelihood ratio tests and minimax radii. Ann. Statist., 47 0 (2): 0 994--1024, 2019
2019
-
[55]
Winkens, R
J. Winkens, R. Bunel, A. G. Roy, R. Stanforth, V. Natarajan, J. R. Ledsam, P. MacWilliams, P. Kohli, A. Karthikesalingam, S. Kohl, T. Cemgil, S. M. A. Eslami, and O. Ronneberger. Contrastive training for improved out-of-distribution detection. arXiv:2007.05566v1, July 2020
2007 arXiv
-
[56]
Y. You, B. Ginsburg, and I. Gitman. Large batch training of convolutional networks. arXiv:1708:03888v3, 2017
2017
-
[57]
F. Yu, A. Seff, Y. Zhang, S. Song, T. Funkhouser, and J. Xiao. LSUN : C onstruction of a large-scale image dataset using deep learning with humans in the loop. arXiv:1506.03365, 2015
2015 arXiv
-
[58]
Zimek, E
A. Zimek, E. Schubert, and H.-P. Kriegel. A survey on unsupervised outlier detection in high-dimensional numerical data. Stat. Anal. and Data Mining, 5 0 (5): 0 363--387, 2012
2012
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.