{"id":"8555a550-c3ac-46ca-9122-72b15099b598","arxiv_id":"2608.05025","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"In canonical joint energy-based models on CIFAR-10, a fixed-noise Predictor-Corrector adaptation shows no detectable advantage over SGLD across training, generation, and OOD detection, and is detectably worse at cold-start generation.","lead":"A replication study compared the standard Langevin sampler with a fixed-noise Predictor-Corrector adaptation in a joint energy-based classifier on CIFAR-10, and found no consistent advantage of the new sampler across training, generation, and out-of-distribution detection. The one reliable difference was in cold-start image generation, where the standard sampler was about five FID points better.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training-protocol no-advantage result depends on the post-hoc margin-10 checkpoint selection; the paper's own fixed-epoch analysis at epoch 105 gives a hierarchical 95% CI excluding zero in PC's favor.","rationale":"I agree with the reader's conditional assessment, but I locate the most load-bearing concern differently. The reader's weakest assumption is the untuned PC parameterisation; I find that concern real yet explicitly scoped in §6.3 ('the selected PC parameterisation'), so it does not directly contradict the paper's stated claim. The margin-10 checkpoint alignment, by contrast, is a load-bearing internal choice: the paper itself reports that a reasonable prospective alternative (fixed epoch 105) yields a hierarchical 95% CI that excludes zero in PC's favor, while fixed epoch 90 excludes zero in SGLD's favor. The central null claim on the training protocol is therefore not stable under an analysis the authors themselves consider worth reporting. The paper is honest about this sensitivity and labels the fixed-epoch intervals exploratory, which is why I do not recommend rejection; but the null conclusion should not be considered established until the checkpoint-selection sensitivity is mapped across all common checkpoints and divergence-distance choices. This is a concrete, feasible check using the already-released per-image scores, and it would either confirm the margin-10 result as robust or force the training-protocol conclusion to be rephrased as epoch-dependent.","tokens_in":16148,"tokens_out":8030,"duration_ms":71983,"concrete_test":"Using the published per-image scores for all four runs, compute the macro-5 AUROC difference (PC−SGLD) and hierarchical seed×image 95% CI at every saved checkpoint common to all runs (epochs 85, 90, 95, 100, 105, 110), and repeat the margin-10 analysis with divergence-distance d ∈ {5, 10, 15}. If any fixed-epoch CI excludes zero or the margin-d results change sign with d, the training-protocol null claim is alignment-dependent and must be rephrased; if all intervals contain zero or cover both signs, the margin-10 conclusion is robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim's training-protocol leg rests on the margin-10 checkpoint alignment in §5.2: each run is scored at its own pre-divergence checkpoint, a choice that requires knowing the future divergence epoch. The paper's own fixed-epoch analysis shows the conclusion is not stable under this choice. At epoch 90, the hierarchical seed×image bootstrap gives Δmacro-5 = −0.009, 95% CI [−0.016,−0.002] (SGLD ahead); at epoch 105 it gives +0.018, CI [+0.003,+0.033] (PC ahead, zero excluded). The authors prefer margin-10 because fixed epochs conflate trajectory phases, but fixed epochs are the prospective, practically relevant comparison, and the sign flip means the 'no consistent method-level advantage' claim is an artifact of which alignment is primary rather than a stable feature of the data. This is compounded by the acknowledged ±1–2 epoch sensitivity of divergence-point identification (§6.3). Since the central conclusion is a null result, the burden is to show the null survives a reasonable range of checkpoint rules.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reproduces canonical JEM (Grathwohl et al.) on WideResNet-28-10 without normalization on two independent runs and compares SGLD with a fixed-noise Predictor-Corrector (PC) adaptation across three protocols: full training-trajectory replacement, cold-start FID generation, and refinement-style multi-OOD AUROC. The reconstruction reaches 92.88% test accuracy versus the canonical 92.90%, with a buffer-FID of 44.46 versus 38.40. The paper documents two failure modes: catastrophic late-training divergence in all four runs and run-dependent SVHN OOD-discrimination dynamics. The main empirical claim is that no consistent method-level advantage of the PC adaptation is detected; cold-start generation favors SGLD by about 5 FID points; and on the training protocol the hierarchical seed-by-image bootstrap interval at the margin-10 checkpoints contains zero, while a seed-level TOST with two runs per method is underpowered. The paper explicitly labels exploratory intervals, reports a sign-flipping fixed-epoch analysis, and identifies the restricted PC parameterization as a limitation.","tokens_in":16452,"tokens_out":6380,"duration_ms":54482,"significance":"If the result holds, it is a useful negative result: sampler-level interventions within the SGLD family are unlikely to resolve canonical JEM's known instability, and the annealed-noise guarantees of the Predictor-Corrector framework do not transfer to the fixed-noise setting. The manuscript is unusually honest about its own limitations: it reports the fixed-epoch sign flip, labels the hierarchical bootstrap as exploratory, interprets the TOST as underpowered, and provides reproducible code and per-image data. The two-seed documentation of run-dependent OOD dynamics is a substantive contribution to the replication literature on JEM. The theoretical explanation in §6.1 is coherent, though post-hoc. The main weaknesses are the sensitivity of the training-protocol conclusion to checkpoint alignment and a potential deviation from the canonical training schedule in the reconstruction.","major_comments":[{"comment":"The training-protocol null is tied to the margin-10 checkpoint rule, and the paper's own fixed-epoch analysis does not support the same conclusion. At epoch 90 the exploratory hierarchical 95% CI on Δmacro-5 is [−0.016,−0.002] (SGLD ahead), and at epoch 105 it is [+0.003,+0.033] (PC ahead, zero excluded); only the margin-10 alignment gives a CI containing zero. Because the margin-10 rule requires post-hoc knowledge of each run's future divergence epoch and is sensitive to the ±1–2 epoch identification error acknowledged in §6.3, the abstract's unqualified statement that a hierarchical bootstrap gives a confidence interval that contains zero is misleading. The main claim should state explicitly that the zero-containing interval depends on the margin-10 selection and that a prospective fixed-epoch comparison gives sign-flipping results.","section":"§5.2, Table 3, and Abstract"},{"comment":"The reconstruction applies K=40 SGLD steps 'from the first epoch onward' and describes this as 'a pure canonical reconstruction without the two-phase 20→40 switch.' If the canonical JEM procedure includes the two-phase switch, then this is a deliberate departure from canonical training, and Table 1's comparison against the canonical reference values conflates sampler-configuration differences with reconstruction fidelity. The manuscript should either verify that the official repository uses K=40 from epoch 1 or re-scope the 'high-fidelity canonical reconstruction' claim and discuss the potential impact on the FID gap and divergence epochs.","section":"§3, Algorithm 1"},{"comment":"The cold-start FID conclusion (ΔFID = +4.87±0.31, SGLD better) is based on a single training checkpoint from one SGLD run; the dispersion covers generation seeds only and does not sample between-run variability, as the paper itself notes in §6.3. The abstract nevertheless presents 'seeded cold-start generation favours SGLD' as a method-level result. This should be qualified as 'on the single evaluated checkpoint' or supported with additional training runs, since the central claim of a detectable difference rests on one model.","section":"§5.3 and Abstract"}],"minor_comments":[{"comment":"The sentence 'the canonical reading of JEM as a model that consistently assigns higher likelihoods ... turns out to be a seed-dependent property' is based on two runs; suggest rewording to 'is seed-dependent in the two runs observed here' to avoid overgeneralization.","section":"§4.2"},{"comment":"The statement that the seed-level percentile interval 'coincides with the full support of the nine possible cross-method pairings' is clear but deserves one additional sentence noting that this support-size argument formally precludes coverage calibration, not just in practice but by construction.","section":"§5.5"},{"comment":"The prospective power calculation assumes a pooled per-run standard deviation of approximately 0.008 macro-5 AUROC; please state explicitly which checkpoints and epochs this estimate is pooled over, since the fixed-epoch analysis suggests the dispersion depends on the alignment rule.","section":"§6.3"},{"comment":"The text contains minor encoding and formatting issues, including 'Fr´echet' in the abstract and 'T able 1' in the full text; these should be cleaned up.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest, carefully scoped negative replication with transparent statistics and reproducible artifacts. The main issues are the fragility of the training-protocol conclusion under the checkpoint-alignment choice and the need to clarify the K=40 deviation from the canonical schedule. Both are fixable within the manuscript's scope. No concerns about author conduct or citation integrity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: read this, and it deserves refereeing, but the headline training-protocol conclusion is less robust than the abstract suggests. The paper does three things well. It reproduces canonical JEM accurately (92.88% vs 92.9%), documents a catastrophic divergence mode in all four runs, and reports run-dependent SVHN OOD dynamics that most of the literature misses because it runs one seed. The hierarchical seed-by-image bootstrap and the honest TOST interpretation are a good statistical template. Code and data are out. That is real credit.\n\nThe soft spot is the margin-10 checkpoint alignment. Each run is scored at its own pre-divergence checkpoint, which requires knowing the future. The paper discloses this and calls it post-hoc, but then leans on it. The fixed-epoch comparison, which is the prospective one, shows the training-protocol effect flipping sign: at epoch 90 SGLD is ahead with CI excluding zero; at epoch 105 PC is ahead with CI excluding zero. That is not 'reinforcing' the null, as the paper says; it is evidence that the null is an artifact of alignment choice rather than a stable feature. The hierarchical CI at margin-10 contains zero, but that CI is conditional on the alignment.\n\nOther caveats are smaller. n=2 per method underpowers everything; the authors say so. The PC variant is untuned and fixed-noise, so the negative result is properly scoped to that parameterization, and they say that too. Protocols 2 and 3 use one training run. None of this is hidden.\n\nWho is it for? People working on JEM/EBM samplers and anyone running replication studies. The failure-mode documentation and the seed-dependence finding will be cited. The central sampler comparison should be treated as provisional until the checkpoint-rule sensitivity is resolved.\n\nRecommendation: send to peer review, ask for a sensitivity analysis over checkpoint rules, or a reframing of the training claim as 'no consistent advantage across protocols, with direction depending on alignment.' With that, it is a solid paper.","headline":"Honest, well-scoped replication with a fragile central null: the training-protocol 'no advantage' result flips sign under the paper's own fixed-epoch analysis.","tokens_in":16876,"tokens_out":1940,"would_cite":true,"duration_ms":16994,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fixed-noise Predictor-Corrector sampler gives canonical JEM no detectable edge over SGLD on CIFAR-10.","keywords":["energy-based models","Joint Energy-Based Model","SGLD","Predictor-Corrector sampler","OOD detection","CIFAR-10","replication study","hierarchical bootstrap"],"falsifier":"Run Protocol 1 with a differently configured PC variant—for instance annealed noise $\\sigma_t$ or a training-time $K_c=3$—and observe a hierarchical 95% CI on the macro-5 AUROC shift whose lower bound exceeds $+0.016$; alternatively, find a paired cold-start FID difference near $-5$ in PC's favour across three seeds. Either result would overturn the paper's no-advantage claim on the training or generation protocol.","tokens_in":15907,"feed_emoji":"⚖️","tokens_out":6156,"duration_ms":46418,"temperature":0.7,"pith_summary":"This paper asks whether a theoretically motivated alternative to stochastic gradient Langevin dynamics (SGLD) — the Predictor-Corrector (PC) sampler from score-based generative modelling — delivers any practical gain when dropped into the canonical Joint Energy-Based Model (JEM) on CIFAR-10. The author reproduces canonical JEM on WideResNet-28-10 (92.88% test accuracy vs the canonical 92.90%) and then tests a fixed-noise PC adaptation across three protocols: full training-trajectory replacement, cold-start generation, and refinement-style OOD detection. The result is a consistent negative: no method-level advantage of PC over SGLD appears on any protocol, the training-protocol 95% confidence interval on the macro-averaged AUROC difference contains zero, and seeded cold-start generation favours SGLD by about five FID points. This is theoretically consistent with the fact that the PC predictor step degenerates under the fixed noise $\\sigma=0.01$ of canonical JEM, so the annealed-noise guarantees do not transfer. A sympathetic reader would take the paper as establishing that sampler-level swaps within the SGLD family are unlikely to fix JEM's instability, and that structural changes are the more promising route.","feed_headline":"PC sampler fails to beat SGLD in canonical JEM","feed_subtitle":"Three protocols show no consistent gain; cold-start FID is about five points worse under the adaptation.","key_machinery":"The load-bearing object is the fixed-noise PC adaptation (Algorithm 2): a deterministic gradient step replaces the degenerate annealed-noise predictor, followed by $K_c$ Langevin corrector steps with constant $\\sigma=0.01$ and step size $\\alpha=1.0$ inherited from canonical SGLD. The argument turns on the degeneration of the predictor step under constant noise: in the variance-exploding SDE, the predictor step is $x' \\leftarrow x + (\\sigma_{i+1}^2 - \\sigma_i^2) s_\\theta(x, \\sigma_{i+1})$, so with $\\sigma_i = \\sigma_{i+1}$ the step vanishes and the theoretical coupling between predictor and corrector is broken. The statistical machinery that carries the comparison is the hierarchical seed-by-image bootstrap, which resamples runs at the seed level before paired-resampling images, and a seed-level Welch two one-sided tests procedure; the paper uses these to separate per-image scoring noise from the between-run variability that dominates the training comparison.","core_discovery":"The central claim is that, within the unchanged canonical JEM setting, a fixed-noise PC adaptation shows no detectable advantage over SGLD on any of three protocols, and on cold-start generation it is detectably worse. Across ten checkpoint–OOD pairs, refinement AUROC differences stay below 0.007; over three paired seeds, cold-start FID is $57.76 \\pm 0.26$ for PC versus $52.89 \\pm 0.12$ for SGLD ($\\Delta = +4.87 \\pm 0.31$); on the training protocol, a hierarchical seed-by-image bootstrap gives a 95% confidence interval on the macro-5 AUROC shift of $[-0.005, +0.016]$, which contains zero, while a seed-level equivalence test with two runs per method cannot establish formal equivalence. The paper also documents two failure modes of canonical JEM: catastrophic late-training divergence with the signature of the canonical outlier-buffer mechanism in all four runs, and run-dependent SVHN OOD-discrimination dynamics. The explanation offered is that the theoretical guarantees of the annealed-noise PC framework do not transfer: with fixed noise, the VE predictor step length is proportional to the difference between adjacent noise levels, which is zero, so the adaptation reduces to an SGLD-like recurrence with half the stochastic noise budget.","pith_inferences":["If the predictor-degeneration argument generalises, the no-advantage result should extend to any constant-noise JEM variant, not just CIFAR-10; a cheap check is to repeat Protocol 1 on a different dataset with the same canonical settings.","The seed-dependent SVHN dynamics suggest that published single-seed OOD evaluations of JEM may be reporting seed artefacts; the natural extension is a multi-seed benchmark of OOD detectors in which stability across seeds is a first-class metric.","The untested alternative—an annealed-noise PC with a native noise schedule—would sit outside canonical JEM and inside the EBM–diffusion hybrid space, where the PC guarantees could reappear; the present result should not be read as evidence against that family.","A tuned PC parameterisation (e.g., larger $K_c$ at training or annealed $\\sigma$) has not been ruled out by these data, so the practical takeaway for practitioners is that the specific fixed-noise adaptation is not worth adopting, not that the predictor-corrector idea is exhausted."],"forward_implications":["Replacing SGLD with the fixed-noise PC adaptation in canonical JEM training yields no detectable method-level change in OOD discrimination: the macro-5 AUROC shift is $+0.006$ with an exploratory 95% CI $[-0.005,+0.016]$.","At an equal gradient budget, cold-start generation is worse with PC: FID rises by about five points across three paired seeds, and the effect is consistent in sign.","At inference, refinement-style OOD detection does not benefit from the PC sampler: all ten checkpoint–dataset AUROC differences stay below 0.007 AUROC, and the static energy score outperforms every refinement variant.","Catastrophic late-training divergence occurs in all four runs regardless of sampler, so the divergence mode is structural to the canonical configuration rather than caused by the choice between SGLD and PC.","Any practical improvement in stability or generation for this model class is more likely to come from structural changes—bounded deterministic samplers, diffusion-like noise schedules, or cooperative training—than from swapping samplers within the static-noise SGLD family."],"supporting_citations":[{"why":"Defines canonical JEM, its training procedure, the SGLD sampler, and the Appendix H.3 instability that motivates the replication.","marker":"[1]"},{"why":"Supplies the Predictor-Corrector sampler framework whose annealed-noise guarantees the paper tests and finds not to transfer.","marker":"[4]"},{"why":"Introduces SGLD, the baseline sampler that the fixed-noise PC adaptation is compared against.","marker":"[2]"},{"why":"Provides the replay-buffer contrastive-divergence mechanism with 5% reinitialisation used in canonical JEM training.","marker":"[5]"},{"why":"Shows a bounded deterministic sampler reaching FID≈9 on CIFAR-10, structural support for the claim that the SGLD sampler family is the limiting factor.","marker":"[9]"},{"why":"Warns that plain Langevin dynamics without annealing 'won't work in practice', supporting the paper's reading of the degeneration mechanism.","marker":"[6]"}],"fun_headline_variants":["No edge for fixed-noise PC over SGLD in JEM","Canonical JEM: PC sampler offers no upside","SGLD still beats PC adaptation in JEM","JEM test: PC lags SGLD by about 5 FID"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fixed-noise PC adaptation, configured with untuned hyperparameters inherited from canonical SGLD ($\\alpha=1.0$, $\\sigma=0.01$, $K_c=1$), fairly represents the PC sampler family for this comparison.","fun_headline_variants_meta":{"raw":{"variants":["No edge for fixed-noise PC over SGLD in JEM","Canonical JEM: PC sampler offers no upside","SGLD still beats PC adaptation in JEM","JEM test: PC lags SGLD by about 5 FID"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000321,"raw_usage":{"total_tokens":1922,"prompt_tokens":1172,"completion_tokens":750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":788,"completion_tokens_details":{"reasoning_tokens":677}},"tokens_in":788,"tokens_out":750,"duration_ms":6420,"temperature":1.0,"reasoning_tokens":677,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:36:40.365893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Protocol 1 with a differently configured PC variant—for instance annealed noise $\\sigma_t$ or a training-time $K_c=3$—and observe a hierarchical 95% CI on the macro-5 AUROC shift whose lower bound exceeds $+0.016$; alternatively, find a paired cold-start FID difference near $-5$ in PC's favour across three seeds. Either result would overturn the paper's no-advantage claim on the training or generation protocol.","supporting_citations":[{"cited_title":"Your classifier is secretly an energy based model and you should treat it like one","cited_arxiv_id":null,"evidence_quote":"Defines canonical JEM, its training procedure, the SGLD sampler, and the Appendix H.3 instability that motivates the replication."},{"cited_title":"Score-Based Generative Modeling through Stochastic Differential Equations","cited_arxiv_id":null,"evidence_quote":"Supplies the Predictor-Corrector sampler framework whose annealed-noise guarantees the paper tests and finds not to transfer."},{"cited_title":"Bayesian learning via stochastic gradient Langevin dynam- ics","cited_arxiv_id":null,"evidence_quote":"Introduces SGLD, the baseline sampler that the fixed-noise PC adaptation is compared against."},{"cited_title":"Implicit Generation and Modeling with Energy Based Models","cited_arxiv_id":null,"evidence_quote":"Provides the replay-buffer contrastive-divergence mechanism with 5% reinitialisation used in canonical JEM training."},{"cited_title":"Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation","cited_arxiv_id":null,"evidence_quote":"Shows a bounded deterministic sampler reaching FID≈9 on CIFAR-10, structural support for the claim that the SGLD sampler family is the limiting factor."}],"review_version":3}