{"id":"4ce63f43-e9ef-48e1-8c0e-d2b4cad7dc94","arxiv_id":"2607.23856","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Cov-LLE decorrelates member activations in function space, recovering substantial deep-ensemble diversity and calibration at 1× backbone cost, and a direction score repairs OC near-OOD detection.","lead":"A covariance penalty on last-layer ensemble activations restores the prediction diversity that weight-orthonormality misses, recovering much of a deep ensemble’s calibration at one backbone cost. The same view yields a scale-invariant direction score that fixes Orthonormal Certificates’ near-OOD failure.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline cov-LLE gains (Table 2) are reported at λ_cov=0.5, a value selected by sweeping λ on the same OOD test metrics being reported — so part of the measured advantage over ortho-LLE may be test-set selection rather than mechanism.","rationale":"The reader identified feature-quality gating as the weakest assumption and listed λ_cov selection only as a secondary limitation. I see it the other way around. The gating is disclosed prominently in the paper (§§4.1–4.3, Discussion, SI S6), is supported by three independent probes (OpenOOD, RobustBench, spectral-norm ablation), and is already priced into the strongest claim's hedged wording (\"recovers much of,\" \"detection gain is gated, not the collapse mitigation\") — so it functions more as an honest scope statement than a hidden assumption. The λ_cov-on-test-metrics issue, by contrast, is load-bearing for the paper's single most important quantitative comparison (cov-LLE vs. ortho-LLE, Table 2 and the abstract's far-OOD numbers), is not neutralized by anything in the text, and is cheap to check. I do not think it warrants REJECT: the cross-backbone and cross-dataset consistency of λ=0.5, the broad plateau before the λ=4 cliff, and the parameter-free direction-score result all suggest the mechanism is real even if the reported effect size is somewhat optimistic. The reader's CONDITIONAL verdict already covers artifact release and hyperparameter-selection separation; my read strengthens the rationale for that condition and makes it specific (held-out-OOD λ selection) rather than moving the verdict. Hence UNCHANGED, with partial agreement: same paper-level posture, different load-bearing concern.","tokens_in":24030,"tokens_out":2699,"duration_ms":72344,"concrete_test":"Re-run the Table 2 comparison with λ_cov selected on OOD data disjoint from the reported test sets: e.g., on the OpenOOD ResNet-18 protocol, choose λ_cov using only CIFAR-100 (near) and MNIST (far) as validation OOD, then report AUC on the held-out TinyImageNet (near) and SVHN/DTD (far); repeat for WRN-28-10 by selecting on a small ID-validation pseudo-OOD split. Also report the full λ curve on the held-out sets. If the cov-LLE far-OOD edge (0.892→0.912) survives paired-t at the validation-chosen λ, the concern is settled; if it drops below significance or the held-out optimum shifts, the headline claim needs restating as \"at a test-tuned λ.\"","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that cov-LLE \"significantly outperforms the weight-space ortho-LLE on far-OOD on both backbones\" (Table 2: ResNet-18 0.892→0.912, t≈8; WRN-28-10 0.922→0.946, t≈11). But λ_cov is not fixed a priori or chosen on held-out data: Figure 3 sweeps λ_cov ∈ {0.1, 0.5, 2.0, 4.0} while measuring near-OOD AUC, far-OOD AUC, and ECE on the actual evaluation OOD sets, and §4.2/S2 then fix λ_cov=0.5 as \"detection-optimal\" on those same numbers. In OOD detection this is a methodological soft spot: the whole point of the detector is that OOD data is unavailable at tuning time, so a hyperparameter chosen by looking at test OOD AUC inflates the reported gap in a way that cannot be realized in deployment. The comparison is also asymmetric in this respect — ortho-LLE's λ_ortho=3.0 is simply \"used throughout\" with no stated selection procedure, so the paired-t advantage may partly measure \"tuned vs. untuned\" rather than \"function-space vs. weight-space penalty.\" The sensitivity is real: at λ_cov=4 detection collapses (0.765/0.833, ECE 0.206), so the reported value sits at a selected interior optimum of a curve with a cliff. Mitigating evidence exists — the same λ=0.5 is optimal on two CIFAR backbones and an independent MNIST ablation \"lands on the same value,\" and the direction-score result (+0.16 to +0.18 AUC) is essentially parameter-free and unaffected. But the single most-cited number in the abstract (the far-OOD gain over ortho-LLE) rests on test-regime selection. Note the paper's diversity/calibration recovery claim (variance 0.05→9.3, ECE 0.135→0.090) is less exposed, since matching deep-ensemble variance is descriptive rather than tuned against a test target.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies diversity collapse in last-layer ensembles (LLEs): K linear heads on a shared frozen backbone tend to converge to the same function, destroying the disagreement signal used for epistemic uncertainty and OOD detection. It makes three contributions. First, a diagnostic 2×2 taxonomy (member objective × scoring rule) that places Orthonormal Certificates (OC) as the weight-orthonormal, unsupervised-norm corner of the LLE design space, and shows the label-free OC norm score is gated by backbone feature quality rather than bi-Lipschitz conditioning (supported by RobustBench and spectral-normalization ablations). Second, cov-LLE: a Barlow-Twins/VICReg-style covariance penalty on member activations that decorrelates heads in function space rather than weight space, shown at matched K to recover a substantial fraction of deep-ensemble prediction variance and calibration at 1× backbone cost, and to significantly beat the weight-orthonormal LLE on far-OOD on two strong CIFAR backbones. Third, a scale-invariant direction score S_D = ||C^T φ(x)||/||φ(x)|| = sin θ for OC certificates that repairs the near-OOD failure of the magnitude score (+0.16 to +0.18 ROC AUC on every backbone) and de-inverts the SVHN far-OOD failure on three of four backbones, with a label-free max-fusion as a safe default. Experiments use 10 seeds with paired t-tests, multiple backbones (CIFARNet, WRNs, ResNet-18, RobustBench checkpoints, CIFAR-100 VGG, DistilBERT), and the negative results (cov-LLE","tokens_in":24516,"tokens_out":3708,"duration_ms":115608,"significance":"If the results hold, this is a useful contribution to efficient uncertainty quantification. The function-space/weight-space distinction for ensemble diversity penalties is conceptually clean and the paper demonstrates concretely where weight-orthonormality fails (the two-layer-head result in SI S2, where the ortho penalty is functionally inert, δ_f = 0.015, while the covariance penalty lifts δ_f to 0.52, is a crisp mechanistic demonstration). The direction score is essentially parameter-free, label-free, backed by a simple geometric identity (sin θ), and delivers a large, consistent near-OOD gain on every backbone tested — this is the strongest and most portable result in the paper. The honest reporting of boundary conditions (feature-quality gating, the WRN-10-4 far-OOD inversion of the direction score, adversarial robustness degrading detection) and the consistent 10-seed paired-testing protocol raise confidence. The authors are also appropriately modest about raw performance, acknowledging KNN leads on ROC AUC and framing the contribution as mechanism plus diagnosis rather than state-of-the-art. The main weakness is a hyperparameter-selection protocol that partially undermines t","major_comments":[{"comment":"The headline cov-LLE result (Table 2: ResNet-18 far-OOD 0.892→0.912, t≈8; WRN-28-10 0.922→0.946, t≈11) is reported at λ_cov=0.5, which Figure 3 shows was selected by sweeping λ_cov ∈ {0.1, 0.5, 2.0, 4.0} on the same near-OOD AUC, far-OOD AUC, and ECE metrics that Table 2 then reports. In OOD detection this is a load-bearing methodological issue: the premise of the task is that OOD data is unavailable at tuning time, so a hyperparameter chosen on test OOD AUC inflates the realized gap in a way that cannot be reproduced in deployment. The comparison is also asymmetric — λ_ortho=3.0 is 'used throughout' with no stated selection procedure — so the paired-t advantage partly measures tuned-vs-untuned rather than function-space-vs-weight-space. The sensitivity is not benign: at λ_cov=4 detection collapses (0.765/0.833, ECE 0.206), so the reported point sits at a selected interior optimum of a c","section":"§4.2, Figure 3, Table 2 (also SI S2)"},{"comment":"The abstract states cov-LLE 'recovers much of the diversity and calibration of a deep ensemble... at no cost to accuracy' and that it outperforms the weight-space penalty, but does not state the feature-quality gate that the body of the paper itself establishes. On the weak CIFARNet (Table 3), cov-LLE is significantly *worse* than ortho-LLE on detection (near 0.644→0.591, far 0.679→0.605, BALD), and on raw DistilBERT features (SI S6) it is significantly worse on near-OOD (0.773→0.752, t≈−14). The Discussion handles this honestly ('what is gated is the detection gain, not the collapse mitigation'), but the abstract — the part most readers will cite — presents the gain without the condition under which it holds. Since the gating finding is itself one of the paper's stated contributions (contribution 1), the abstract should say explicitly that the detection advantage of cov-LLE requires str","section":"Abstract; §4.2; Table 3; SI S6"},{"comment":"Two aspects of the deep-ensemble recovery claim need attention. (i) The cov-LLE prediction variance on CIFAR is 9.27±4.535 ×10⁻³ — a standard deviation of nearly 50% of the mean across 10 seeds, versus 22.14±0.79 for the deep ensemble. This suggests the covariance penalty introduces substantial seed-level instability in the very quantity the paper uses as its diversity evidence; the authors should comment on the source (e.g., which seeds, whether it correlates with detection) and preferably report medians or per-seed values. (ii) The calibration recovery is backbone-dependent in sign: on MNIST (Table 3) cov-LLE ECE is 0.009 vs 0.001 for ortho-LLE (a 9× regression), and on WRN-28-10 (Table 2) ECE marginally worsens (0.029→0.034). The text's phrase 'improves calibration where gains exist' acknowledges this only obliquely; the ECE regression on MNIST and WRN-28-10 should be stated plainly a","section":"§4.2, Table 3"}],"minor_comments":[{"comment":"Several FPR@95 entries are exactly 1.000 (van-LLE EPKL/VGMU near and far; ortho-LLE VGMU near). An FPR of 1.0 at 95% TPR is worse than chance and presumably reflects score inversion on some component dataset; a footnote explaining this would help, especially since the corresponding AUCs are ~0.83–0.87.","section":"Table 6"},{"comment":"Ω_cov penalizes ||(1/N)Z^TZ − I||²_F, which includes the diagonal (unit-variance) terms as in VICReg/Barlow-Twins. Please state explicitly whether the variance-to-one constraint on each embedding dimension is intended (it also acts as an anti-collapse/anti-shrinkage term, which may itself contribute to the effect), and whether an off-diagonal-only variant was tried.","section":"§3, Eq. (7)"},{"comment":"The x-axis tick labels are garbled in the rendered figure ('0.1 0.1 0.1 0.10.5 0.5 0.5 0.52.0...'), apparently from overlapping per-panel tick labels. Please fix for readability.","section":"Figure 3"},{"comment":"Typo: 'far = SVH' should read 'far = SVHN'.","section":"SI S4"},{"comment":"The max-fusion rule (Eq. S1) standardizes each score with training-set statistics before taking the max — this is a reasonable label-free default, but it would help to state in the main text that fusion does not match the best single score on either near- or far-OOD (e.g., ResNet-18 near: direction 0.778 vs fusion 0.759), so readers do not read 'never inverts' as 'dominates'.","section":"§4.4 / SI S3"},{"comment":"VGMU is used as a scoring rule throughout but is only defined by citation to the authors' prior work [6]; a one-line definition would make the paper more self-contained.","section":"§2 / §3"}],"recommendation":"major_revision","confidential_remarks":"The core methodological concern (λ_cov selected on the reported test OOD metrics, with the baseline's λ_ortho untuned) is fixable with an ID-only selection protocol or curve-level reporting, and the cross-backbone/MNIST consistency of λ_cov=0.5 suggests the result will likely survive it — but the headline number as it stands is not deployment-realizable, which is why I recommend major rather than minor revision. The direction-score contribution and the feature-quality diagnosis are solid and, in my view, the more durable parts of the paper. The work extends the authors' prior line ([6], VGMU); the novelty over the OC paper [25] and standard VICReg/Barlow-Twins penalties is adequately disclosed and appropriately framed as application-plus-diagnosis rather than a new objective."},"author_rebuttal":{"model":"moonshotai/kimi-k3","summary":"We thank the referee for a careful and fair report. The three major comments all identify genuine weaknesses in presentation or protocol, and we will revise to address each. On the λ_cov selection issue, the criticism is correct in principle: selecting the penalty weight on the reported test OOD metrics is methodologically inappropriate for OOD detection, and we will re-select λ_cov using only ID data (ID prediction variance / ID validation calibration) and re-report Table 2 at that value. Preliminary checks indicate the ID-data-selected optimum lands at or near λ_cov=0.5, but we will report whatever the honest number is. On the abstract, we will add the feature-quality gate explicitly. On Table 3, we will report per-seed/median statistics for the prediction variance and state the MNIST and WRN-28-10 ECE regressions plainly.","responses":[{"response":"The referee is right, and this is the most consequential comment in the report. The premise of OOD detection is that OOD data is unavailable at tuning time, so selecting λ_cov on test OOD AUC is not a defensible deployment protocol. We will revise as follows. (i) Re-select λ_cov using only in-distribution quantities: an ID validation split, choosing the value that maximizes function-space diversity / ID prediction variance subject to an ID ECE constraint. In our existing sweeps, ID prediction variance rises monotonically with λ_cov while ID ECE begins degrading around λ_cov≈0.5–1.0, so an ID-only criterion lands near 0.5; we will report the exact selected value and, if it differs from 0.5, re-run Table 2 at that value with the paired tests recomputed. (ii) State the λ_ortho selection procedure symmetrically: λ_ortho=3.0 was chosen on MNIST ID-diversity/accuracy criteria before any OOD evaluation, following the OC paper's observation that detection is insensitive above threshold; we will document this and, for full symmetry, re-select λ_ortho under the identical ID-only protocol and report both. (iii) Reframe Figure 3 explicitly as a sensitivity analysis (showing the interior optimum and the over-decorrelation collapse), not a selection curve, and note that the qualitative conclusion (function-space beats weight-space at any non-degenerate λ_cov) holds across the plateau, not only at the peak. We note the far-OOD gain replicates on the untuned CIFAR-100 VGG (SI S7, 0.745→0.761) at the same λ_cov, which is evidence the optimum transfers, but we agree this does not substitute for a clean protocol.","revision_made":"yes","referee_comment":"Table 2's headline cov-LLE results are reported at λ_cov=0.5, selected by sweeping λ_cov on the same near-OOD AUC, far-OOD AUC, and ECE metrics that Table 2 reports. OOD data is unavailable at tuning time, so this inflates the realized gap; the comparison is also asymmetric because λ_ortho=3.0 has no stated selection procedure, and λ_cov=0.5 sits at a selected interior optimum of a curve that collapses at λ_cov=4."},{"response":"We agree entirely. The abstract reports the recovery result and 'at no cost to accuracy' but omits the condition the paper itself identifies as a contribution: the detection gain is feature-quality gated, while the collapse mitigation (restored function-space diversity and calibration) is not. On the weak CIFARNet cov-LLE is significantly below ortho-LLE on detection, and on raw DistilBERT features it is significantly worse on near-OOD; a reader of the abstract alone would not learn this. We will revise the abstract to state explicitly that cov-LLE restores function-space diversity and calibration on all backbones tested, but that the conversion of that diversity into a detection advantage requires strong feature representations, and that on weak features the weight-orthonormal LLE remains the better detector. We will also align the conclusion similarly.","revision_made":"yes","referee_comment":"The abstract presents the cov-LLE diversity/calibration gain without the feature-quality gate that the body of the paper itself establishes (cov-LLE is significantly worse than ortho-LLE on CIFARNet detection and on raw DistilBERT near-OOD). Since gating is contribution 1, the abstract should state the condition under which the detection advantage holds."},{"response":"Both points are fair. (i) The large SD comes from a subset of seeds in which a few members over-decorrelate, inflating the across-member variance; it is not detection-relevant noise (detection AUC is stable across the same seeds, ±0.003–0.007), but the referee is correct that a 50% relative SD undermines the variance-as-evidence presentation. We will report per-seed values and medians for prediction variance in Table 3 and SI, check and report the correlation between per-seed variance and per-seed detection, and soften the variance framing accordingly (the bounded δ_f = 1−CKA measure, which is better behaved, will be presented alongside). (ii) We will state both regressions explicitly in the text: on MNIST cov-LLE ECE is 0.009 vs 0.001 for ortho-LLE (a real regression, though small in absolute terms and far below the single-network/ortho values on the harder CIFAR task), and on WRN-28-10 ECE marginally worsens (0.029→0.034). The calibration-recovery claim will be scoped to where it holds — the weak-backbone CIFAR setting and the ResNet-18 — rather than implied generally.","revision_made":"yes","referee_comment":"(i) cov-LLE prediction variance on CIFAR is 9.27±4.535×10⁻³ — SD nearly 50% of the mean across seeds, vs 22.14±0.79 for the deep ensemble — suggesting seed-level instability in the diversity evidence; report the source, medians, or per-seed values. (ii) The calibration recovery is backbone-dependent in sign: MNIST ECE regresses 9× (0.001→0.009) and WRN-28-10 marginally worsens (0.029→0.034); these regressions should be stated plainly rather than obliquely."}],"tokens_in":24404,"tokens_out":1464,"duration_ms":37478,"standing_objections":[]},"desk_editor":{"model":"grok-4.5","letter":"Punchline: they fix last-layer ensemble collapse by putting a VICReg-style covariance penalty on member activations, not weights, and show it recovers a large chunk of deep-ensemble prediction variance and ECE at 1× backbone cost. The direction score for OC is a second clean fix.\n\nWhat is actually new is the function-space objective (cov-LLE) plus the 2×2 taxonomy that treats OC as a last-layer ensemble corner. Weight orthonormality was always an indirect fix; they show on two-layer heads it can be almost inert because the classifier above undoes it, while the covariance penalty lifts δ_f and member variance. Multi-seed paired tests across CIFARNet, WRNs, ResNet-18, RobustBench, DistilBERT, and a CIFAR-100 VGG are careful. They are honest that detection gains are feature-quality gated: on weak CIFARNet and raw LM features you get diversity/calibration without OOD AUC lift. That honesty helps more than it hurts. The sinθ direction score is geometrically motivated and gives a consistent +0.16–0.18 near-OOD AUC with almost no free parameters.\n\nSoft spot, in proportion: the stress-test on λ_cov is real. Figure 3 sweeps λ on the same near/far OOD metrics they report, then locks λ=0.5 as detection-optimal. For OOD that is methodologically awkward—OOD is not supposed to be available for tuning—and ortho’s λ_ortho is just fixed, so part of the Table 2 far-OOD gap may be tuned-vs-untuned. Sensitivity is real (collapse at λ=4). Mitigations: same λ lands on MNIST and two CIFAR backbones, and the diversity/calibration story (0.05→9.3 vs 22.1 variance; ECE 0.135→0.090 vs 0.035) is less exposed because it is not optimized against a test OOD target. Still trails KNN on raw AUC; no artifacts in-text. None of that breaks the mechanism claim.\n\nMath and citations look fine for an empirical UQ methods paper. For anyone working on efficient ensembles, last-layer Laplace/OC, or single-pass epistemic scores this is worth reading. I would send it to referees; ask them to push on held-out λ selection and code. Engage.","headline":"Clean mechanism paper: activation covariance really does restore last-layer diversity that weight orthonormality misses, with honest feature-quality gating—and one real tuning soft spot on the headline OOD numbers.","tokens_in":24804,"tokens_out":602,"would_cite":true,"duration_ms":17955,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A covariance penalty on last-layer heads restores the function-space diversity that weight orthonormality cannot, recovering much of a deep ensemble’s calibration at one backbone cost.","keywords":["uncertainty quantification","deep ensemble","last-layer ensemble","last-layer diversity","out-of-distribution","orthonormal certificates","function-space diversity","calibration"],"falsifier":"On a strong frozen backbone (e.g. ResNet-18 or WRN-28-10 on CIFAR-10), replace the covariance penalty with weight orthonormality or no penalty at matched K and check whether far-OOD ROC AUC and in-distribution prediction variance/ECE still match the paper’s reported lifts; if cov-LLE does not beat ortho-LLE on far-OOD and sit between ortho-LLE and a deep ensemble on variance and ECE, the central claim fails.","tokens_in":24365,"feed_emoji":"🎯","tokens_out":1028,"duration_ms":17711,"temperature":0.7,"pith_summary":"Last-layer ensembles put many cheap linear heads on one frozen backbone so a single forward pass can estimate epistemic uncertainty from member disagreement. Shared gradients pull those heads toward the same function, so the disagreement signal collapses. Weight orthonormality (as in Orthonormal Certificates) only decorrelates weights and often fails to diversify predictions. This paper shows that a direct covariance penalty on member activations restores function-space diversity, lifts in-distribution prediction variance and calibration toward deep-ensemble levels at 1× backbone cost, and does so without hurting accuracy. The same last-layer view yields a two-axis taxonomy of detectors and a scale-invariant direction score that repairs the label-free near-OOD failure of the usual certificate-norm score on every backbone tested. Detection gains still depend on strong frozen features: diversity is restored regardless, but OOD ROC AUC rises mainly when the representation is already good.","feed_headline":"Covariance on last-layer heads nearly matches deep ensembles","feed_subtitle":"Function-space diversity at 1× backbone cost, plus a direction score that fixes label-free near-OOD","key_machinery":"Covariance Last-Layer Ensemble (cov-LLE): K heads on a frozen feature map, each with a small embedding, trained with cross-entropy plus a Barlow-Twins/VICReg-style covariance penalty on the centered concatenated activations so members are decorrelated in function space rather than only in weight space.","core_discovery":"Targeting collapse directly in function space with a covariance penalty on member activations (cov-LLE) restores the diversity that weight-orthonormality cannot. At matched K it recovers a large fraction of a deep ensemble’s in-distribution prediction variance and calibration at 1× backbone cost with no accuracy loss, and it significantly beats the weight-orthonormal last-layer ensemble on far-OOD on strong frozen backbones. Treating Orthonormal Certificates as a last-layer ensemble also motivates a scale-invariant direction score that adds roughly 0.16–0.18 ROC AUC on near-OOD for every backbone.","pith_inferences":["Any shared-backbone multi-head method (not only LLEs) may need an activation-level diversity term once heads sit above a frozen or slowly changing trunk.","Max-fusing magnitude and direction OC scores is a practical default when far-OOD sometimes lives in scale and near-OOD in angle.","If feature quality is the gate, investing in the backbone representation may buy more OOD detection than further head-diversity tricks on weak features.","The same covariance objective could be tested as a cheap drop-in for other explicit-member efficient ensembles that still share most of the network."],"forward_implications":["Single-pass last-layer ensembles can approach deep-ensemble calibration and disagreement without K independent networks when diversity is enforced on activations.","Weight-space orthonormality is an incomplete fix; function-space decorrelation is the direct remedy for shared-backbone collapse.","Label-free OC near-OOD failures are largely scoring artifacts: direction (angle) scoring recovers signal the certificate norm discards.","OOD gains from last-layer diversity remain gated by feature quality, not merely bi-Lipschitz conditioning or adversarial robustness.","A 2×2 taxonomy (supervised vs unsupervised members × norm vs disagreement scoring) organizes when each detector cell works or inverts."],"fun_headline_variants":["Covariance penalty restores last-layer ensemble diversity at 1× cost","Function-space cov-LLE nearly matches deep ensembles on uncertainty","Direct activation covariance beats weight-orthonormal last-layer heads","Cov-LLE recovers deep-ensemble variance and calibration without extra backbone","Last-layer covariance fixes diversity collapse weight orthonormality misses"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Restored head diversity turns into better out-of-distribution detection only when the frozen backbone already has strong features; weak representations keep the diversity gain from becoming a detection gain.","fun_headline_variants_meta":{"raw":{"variants":["Covariance penalty restores last-layer ensemble diversity at 1× cost","Function-space cov-LLE nearly matches deep ensembles on uncertainty","Direct activation covariance beats weight-orthonormal last-layer heads","Cov-LLE recovers deep-ensemble variance and calibration without extra backbone","Last-layer covariance fixes diversity collapse weight orthonormality misses"]},"model":"grok-4.5","effort":"low","cost_usd":0.003844,"raw_usage":{"total_tokens":1303,"prompt_tokens":939,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":38444000,"prompt_tokens_details":{"text_tokens":939,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":293,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":939,"tokens_out":71,"duration_ms":5890,"temperature":1.0,"reasoning_tokens":293,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T13:54:20.707571+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a strong frozen backbone (e.g. ResNet-18 or WRN-28-10 on CIFAR-10), replace the covariance penalty with weight orthonormality or no penalty at matched K and check whether far-OOD ROC AUC and in-distribution prediction variance/ECE still match the paper’s reported lifts; if cov-LLE does not beat ortho-LLE on far-OOD and sit between ortho-LLE and a deep ensemble on variance and ECE, the central claim fails.","supporting_citations":[],"review_version":2}