{"id":"9c8985a6-12b2-4626-8921-ec355ec22af1","arxiv_id":"1908.09625","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Posterior-based extreme value outlier rejection outperforms predictive entropy for out-of-distribution detection, while the additional benefit of generative classifiers is partial and benchmark-dependent.","lead":"This paper tests whether generative classifiers detect unseen data better than ordinary discriminative classifiers when paired with extreme value theory. It finds that uncertainty alone is not enough, and that latent-space extreme value rejection helps, with mixed evidence that generative models help further.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's MNIST-trained rows contradict the paper's blanket claims that EVT beats predictive entropy in all cases and that the generative decoder improves latent EVT.","rationale":"Reader's weakest_assumption concerns the Weibull-tail premise inherited from reference [16]. That is a legitimate external validity issue, but the more immediate problem is that the paper's own Table 1 contradicts the universal wording of the central claim. The reader's rationale already notes explicit counterexamples in the MNIST-trained rows, so my read partially agrees; however, I would elevate this from a caveat to the primary load-bearing concern because it targets the headline question 'does OOD detection require generative classifiers?' rather than an auxiliary distributional assumption. The generative model's advantage is strong for FashionMNIST- and SVHN-trained models but reverses for MNIST-trained models; a claim of 'might require generative models' cannot rest on a result that is negative in one of three training distributions. The recommended verdict remains conditional: the empirical study is useful, but the paper must either restrict its claims to the supported regimes, report seeds and error bars, add standard baselines such as ODIN and Mahalanobis, or provide a mechanistic explanation for the dataset-dependence. No verdict change from the reader is needed because the reader already chose CONDITIONAL; this stress-test sharpens the justification.","tokens_in":7675,"tokens_out":13051,"duration_ms":121080,"concrete_test":"Independently rerun the MNIST-trained experiments (variational discriminative and variational generative, same architecture and same beta, at least 10 seeds) and report mean plus/minus standard deviation OOD detection for entropy and latent EVT on all six OOD sets. If the generative latent mean is not above the discriminative latent mean on at least 5 of 6 OOD sets, or if latent is not above entropy on SVHN, the universal claims in Section 3.1 and the abstract fail and must be restricted to FashionMNIST- and SVHN-trained settings. A simpler immediate check is to verify these counterexamples directly from Table 1; if confirmed, the 'all cases' wording should be removed even before re-running.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two universal parts: EVT-based latent rejection outperforms predictive entropy in all cases, and adding a decoder to form a joint generative model further improves latent EVT. Section 3.1 states both, and the abstract generalizes them. Table 1 already contains counterexamples. For the MNIST-trained variational generative model, the SVHN column shows entropy detection at 96.53% versus latent EVT at 96.29%, so latent does not beat entropy. More importantly for the 'generative helps' claim, on MNIST-trained models the variational discriminative latent detector beats the generative latent detector on 5 of 6 OOD sets: FashionMNIST 99.86 vs 96.60, KMNIST 99.53 vs 98.97, CIFAR10 99.98 vs 99.81, CIFAR100 99.97 vs 99.65, SVHN 97.70 vs 96.29; only AudioMNIST favors the generative model (99.65 vs 99.98). The 'all cases' sentence is also false for the standard MNIST-trained classifier, for example CIFAR10 entropy 91.06 vs latent 87.62. These are not external-baseline quibbles; the paper's own table contradicts its headline generalization. Because the title question asks whether generative classifiers are required, a trained-distribution where they clearly hurt is directly load-bearing. The absence of error bars and seeds makes it impossible to tell whether the MNIST reversal is a stable effect or noise, so the blanket wording cannot be defended.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an empirical comparison of three classifier families for out-of-distribution detection: a standard discriminative network, a variational discriminative classifier, and a variational joint generative classifier. For each model, the authors compare outlier rejection based on predictive entropy with rejection based on extreme value theory applied to distances from class-conditional latent means, following their earlier open-set recognition procedure. Experiments are run on three training distributions (FashionMNIST, MNIST, SVHN) and evaluated on seven datasets, with and without Monte Carlo dropout. The central claims are that EVT-based latent rejection outperforms predictive-entropy rejection in all cases and that the joint generative model further improves latent EVT, leading the authors to ask whether classifiers need to be generative in order to recognize what they have not seen.","tokens_in":7952,"tokens_out":4621,"duration_ms":46658,"significance":"If the claims were fully supported, the paper would be a valuable empirical contribution: it would show that latent-space Weibull rejection on a variational classifier can be more reliable than predictive entropy, and that adding a decoder can improve open-set recognition. The study is reasonably broad in its evaluation, including non-image AudioMNIST, and it systematically varies the training distribution. The paper also has strengths in transparency: the evaluation threshold is explicitly defined (95% of in-distribution validation data), and the comparison includes Monte Carlo dropout as an additional epistemic-uncertainty mechanism. However, the headline generalizations are contradicted by the paper's own Table 1 in important settings, and the absence of repeated-seed statistics makes it impossible to separate genuine effects from noise. The topic is timely and the central question is interesting, but the evidence in the current manuscript is not yet sufficient to support the stated conclusions.","major_comments":[{"comment":"The blanket statement in Section 3.1 that 'the EVT approach ... outperforms OOD detection with prediction uncertainty in all cases' is contradicted by Table 1. For the MNIST-trained variational generative model on SVHN, entropy detection is 96.53% while latent EVT detection is 96.29%; for the MNIST-trained standard discriminative classifier on CIFAR10, entropy is 91.06% while latent EVT is 87.62%. The related claim that the joint generative model 'further improves' latent EVT is also contradicted: on MNIST-trained models, the variational discriminative latent detector outperforms the variational generative latent detector on 5 of 6 OOD sets (e.g., FashionMNIST 99.86 vs 96.60, CIFAR10 99.98 vs 99.81, SVHN 97.70 vs 96.29). These are not external-baseline quibbles; they are internal counterexamples to the paper's universal claims and should be analyzed explicitly.","section":"Section 3.1 and Table 1"},{"comment":"All results appear to come from a single run, and no error bars, standard deviations, or repeated-seed experiments are reported. Several of the comparisons that support or contradict the central claims differ by less than one percentage point (e.g., 96.53 vs 96.29 for the MNIST-trained variational generative model on SVHN), so the rankings may be within run-to-run noise. Since the paper's main conclusions depend on these small differences, the absence of variance estimates is load-bearing. Please report means and standard deviations over at least three to five seeds and re-evaluate the universal claims in light of the resulting confidence intervals.","section":"Section 3, Experiments"},{"comment":"Equation (1) as printed is L = E_q[log p_phi(x|z) + log p_xi(y|z)] - KL(q_theta(z|x) || p(z)), yet the text states that 'beta is an additional parameter that weighs the contribution of the Kullback-Leibler divergence' and cites the beta-VAE. The displayed objective does not contain beta. Either the beta is missing from the equation, or the model is not the beta-VAE-style objective described in the text. Please correct the equation and state the actual beta value used in the experiments.","section":"Equation (1)"},{"comment":"The comparison is limited to predictive entropy versus the authors' own latent EVT procedure. Standard out-of-distribution baselines available at the time of the paper (e.g., softmax confidence thresholding, ODIN, or Mahalanobis distance) are not evaluated, and results are reported only at a single 95%-inlier operating point rather than as full ROC/AUPR curves. To support the title-level conclusion about what OOD detection 'requires,' the paper should at least include a standard softmax-confidence baseline and report AUROC or a similar threshold-free metric.","section":"Section 3.1 and Table 1"}],"minor_comments":[{"comment":"There is a typo in the first paragraph: 'necesssarily' should be 'necessarily.'","section":"Introduction"},{"comment":"The text says 'additionally' in 'additionally the KMNIST'; this should be 'additionally, the KMNIST.'","section":"Section 3.1"},{"comment":"The tail-size parameter is denoted eta in the algorithm and described as 5% in the main text, but the paper does not state whether this value was chosen a priori or tuned. Please clarify.","section":"Algorithm 1"},{"comment":"The table is difficult to read because the training dataset appears in the first column but is repeated for each model variant, and the column header 'Latent' is not defined in the caption. Please define 'Latent' as the EVT-based rejection rate and restructure the table to make the model variant and training distribution unambiguous.","section":"Table 1"},{"comment":"The exclusion of likelihood-only VAEs is justified only by citing Nalisnick et al. [17] after the fact. Since the paper's central question concerns generative classifiers, a direct small experiment on a likelihood-only VAE would strengthen the argument.","section":"Section 2"},{"comment":"No code or trained model checkpoints are provided, which makes it difficult to reproduce the EVT calibration procedures in Algorithms 1 and 2, especially because the method is carried over from the authors' prior work [16].","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a workshop-style empirical study that has been submitted to a journal venue. The central empirical claims are internally inconsistent with Table 1, and the lack of repeated-seed statistics is a serious concern. The authors should be asked to either substantially narrow their claims or provide a more careful analysis, including error bars and standard baselines, before the manuscript can be considered. The reliance on the authors' own prior work [16] for the EVT procedure, without an independent implementation or code release, is also worth flagging to the editor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper runs an empirical comparison that I haven't seen before: predictive entropy versus EVT-based latent-space rejection, across discriminative and generative variational classifiers, with and without Monte Carlo dropout. That gap is real and worth testing. The EVT-over-entropy pattern mostly holds, and the paper gives credit to the prior work it builds on. But the paper's own Table 1 contradicts its stronger claims. On MNIST-trained models, the discriminative latent detector beats the generative one on five of six OOD sets, so the statement that the generative decoder further improves latent EVT is not supported. There are also examples where latent EVT is worse than entropy, e.g., the standard MNIST classifier on CIFAR10. The abstract and Section 3.1 state these as blanket claims, so they need qualification. Other soft spots: no error bars or seeds, the beta hyperparameter in Eq. 1 is never defined, and no standard baselines like ODIN or Mahalanobis. The EVT-on-posterior procedure comes from the authors' prior work without independent code, so the reader has to trust that implementation. This is a reasonable workshop-level study, not a finished benchmark paper. It could be conditionally accepted if the authors fix the overgeneralizations, add seeds and error bars, and include at least one standard baseline. The comparison is worth running and the question is relevant, so I would send it to peer review rather than desk reject. To summarize: it's a solid empirical starting point with honest limitations, but the conclusions outrun the data. The right reader is someone deciding whether generative decoders buy OOD robustness; this gives a useful hint, not an answer.","headline":"Useful empirical comparison, but the paper's own Table 1 contradicts its claims that EVT always beats predictive entropy and that generative decoders always help.","tokens_in":520,"tokens_out":1433,"would_cite":false,"duration_ms":38798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Latent distances beat predictive uncertainty for out-of-distribution detection","keywords":["open set recognition","out-of-distribution detection","predictive uncertainty","extreme value theory","Weibull distribution","variational autoencoder","generative classifier","epistemic uncertainty"],"falsifier":"A reader could falsify the mechanism by taking a trained model from the paper's setup, computing latent codes for an out-of-distribution dataset, and checking whether the empirical tail of the cosine distances matches the fitted Weibull distribution: if many out-of-distribution points fall inside the fitted per-class high-density region, the open-space bound is not doing the work the paper assigns it.","tokens_in":7454,"feed_emoji":"🎯","tokens_out":5367,"duration_ms":51703,"temperature":0.7,"pith_summary":"The paper asks whether a classifier must be generative to know what it has not seen. It compares three ways to reject out-of-distribution inputs on image classifiers: reading prediction entropy, fitting extreme-value tails to latent representations, and a combination of both in variational models. Its central finding is that entropy-based rejection alone cannot separate seen from unseen data, while latent-space extreme value theory (EVT) rejection does much better, and a variational classifier paired with a generative decoder does best of all. The conclusion, if true, means that modeling the input distribution $p(x)$ as well as the label distribution $p(y)$ gives the latent codes the structure needed for reliable open-set rejection.","feed_headline":"Latent distances beat predictive uncertainty for outlier rejection","feed_subtitle":"Weibull tails on generative latent codes catch unseen images that entropy alone misses.","key_machinery":"The load-bearing mechanism is per-class Weibull tail fitting on the latent approximate posterior. For each training class, after sampling $z\\sim q_\\theta(z|x)$ for correctly classified training inputs, the model computes the class's mean latent vector $\\bar{S}_c$, fits a Weibull distribution to the cosine distances $\\|S_c - \\bar{S}_c\\|$ with a tail size of 5 percent of training examples per class, and rejects a new input when the Weibull CDF value at its distance to any class mean exceeds a task prior $\\Omega_t$. The joint generative variant adds a decoder $p_\\varphi(x|z)$ to the variational classifier, trained with the $\\beta$-VAE-style ELBO in equation (1), so the latent space is shaped by both label and data reconstruction. This mechanism converts epistemic uncertainty from a soft signal into a hard boundary on where the model can be trusted.","core_discovery":"On FashionMNIST-, MNIST-, and SVHN-trained 14-layer wide residual networks, the authors find that predictive entropy, including entropy from variational inference and Monte Carlo dropout, leaves out-of-distribution datasets heavily overlapping with in-distribution data. Latent EVT meta-recognition—fitting a Weibull distribution to the cosine distances between each correctly classified training example's approximate-posterior sample and its class's latent mean, then rejecting any input whose Weibull CDF exceeds a threshold—removes most of this overlap. Adding a probabilistic decoder to learn the joint model $p(x,y,z)=p(y|z)p(x|z)p(z)$ improves the EVT rejection further, reaching near-perfect outlier detection on most cross-dataset pairs while preserving accuracy; the decoder is what the authors point to as the reason the latent space carries information about the data distribution.","pith_inferences":["A direct comparison the paper leaves untested is a likelihood-only variational autoencoder with the same per-class Weibull calibration; the paper excludes such models by citing earlier failures, so the question of whether the joint training is essential remains open.","The Weibull tail assumption is testable per class on any trained model by checking quantile-quantile plots of the empirical distance tail against the fitted distribution; the paper does not report such a diagnostic.","Because the paper evaluates on 32x32 resized images and relatively small datasets, the next test is whether the same latent-EVT gap persists on larger, natural-image benchmarks at native resolution, where latent structure is less separable."],"forward_implications":["Entropy of the predictive distribution, even averaged over 100 posterior samples or 50 Monte Carlo dropout passes, is not a reliable enough signal to reject unseen datasets on these tasks.","Latent-space Weibull rejection raises outlier detection rates substantially over entropy thresholds for all three model families.","A variational classifier that also models the input distribution with a decoder outperforms the discriminative variational classifier under latent EVT rejection on most tested dataset pairs.","With Monte Carlo dropout added to the generative model, out-of-distribution rejection becomes near-perfect for several cross-dataset pairs in the paper's experiments.","The rejection threshold $\\Omega_t$ is easier to set for the generative model because its rejection rate stays more stable across a wide range of priors."],"supporting_citations":[{"why":"Supplies the baseline EVT meta-recognition approach on penultimate-layer features, which the paper adapts to the latent space.","marker":"[2]"},{"why":"Provides dropout as approximate Bayesian inference, the basis for the Monte Carlo dropout estimates of epistemic uncertainty.","marker":"[5]"},{"why":"Provides the variational autoencoder formulation and ELBO that underpin the latent variable models in equation (1).","marker":"[12]"},{"why":"The authors' prior work that introduces fitting EVT to the approximate posterior in a latent variable model; Algorithms 1 and 2 come from this line.","marker":"[16]"},{"why":"Shows that likelihood-only deep generative models fail to separate seen from unseen data, which is why the paper excludes likelihood-only VAEs.","marker":"[17]"},{"why":"Large empirical study showing predictive uncertainty under dataset shift is not enough, the motivation for contrasting it with EVT rejection.","marker":"[19]"}],"fun_headline_variants":["Latent Weibull distances beat predictive entropy for OOD rejection","Generative classifiers know unseen images; discriminative ones don't","Uncertainty alone fails OOD; generative latent codes succeed","For open-set detection, make your classifier generative","EVT on latent codes outperforms entropy for outlier detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central bet is that, for each class, the distances of correctly classified training examples to their class's average latent representation follow a Weibull tail, and this tail marks the boundary beyond which inputs should be rejected; the paper itself notes that the supporting experiments are small-scale and larger evaluation is still needed.","fun_headline_variants_meta":{"raw":{"variants":["Latent Weibull distances beat predictive entropy for OOD rejection","Generative classifiers know unseen images; discriminative ones don't","Uncertainty alone fails OOD; generative latent codes succeed","For open-set detection, make your classifier generative","EVT on latent codes outperforms entropy for outlier detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2898,"prompt_tokens":803,"completion_tokens":2095,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":419,"completion_tokens_details":{"reasoning_tokens":2012}},"tokens_in":419,"tokens_out":2095,"duration_ms":14021,"temperature":1.0,"reasoning_tokens":2012,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:06:16.634275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could falsify the mechanism by taking a trained model from the paper's setup, computing latent codes for an out-of-distribution dataset, and checking whether the empirical tail of the cosine distances matches the fitted Weibull distribution: if many out-of-distribution points fall inside the fitted per-class high-density region, the open-space bound is not doing the work the paper assigns it.","supporting_citations":[{"cited_title":"Bendale and T","cited_arxiv_id":null,"evidence_quote":"Supplies the baseline EVT meta-recognition approach on penultimate-layer features, which the paper adapts to the latent space."},{"cited_title":"Gal and Z","cited_arxiv_id":null,"evidence_quote":"Provides dropout as approximate Bayesian inference, the basis for the Monte Carlo dropout estimates of epistemic uncertainty."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the variational autoencoder formulation and ELBO that underpin the latent variable models in equation (1)."},{"cited_title":"Unified Probabilistic Deep Continual Learning through Generative Replay and Open Set Recognition","cited_arxiv_id":"1905.12019","evidence_quote":"The authors' prior work that introduces fitting EVT to the approximate posterior in a latent variable model; Algorithms 1 and 2 come from this line."},{"cited_title":"Nalisnick, A","cited_arxiv_id":null,"evidence_quote":"Shows that likelihood-only deep generative models fail to separate seen from unseen data, which is why the paper excludes likelihood-only VAEs."}],"review_version":1}