{"id":"c970ec87-5537-4ecf-a826-4350f97a728b","arxiv_id":"2607.26582","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Zero-shot OOD detector rankings do not transfer across domains or models; a complementary-evidence wrapper (CEG) cuts FPR95 without using OOD samples.","lead":"Across 17 datasets and three vision-language models, zero-shot out-of-distribution detector rankings keep changing, so there is no universal best detector. The paper explains why via complementary 'level' and 'sharpness' evidence and proposes a wrapper, CEG, that reduces failure rates without needing OOD samples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CEG's headline gains are measured on the same task domains used to select its hyperparameters and protected weight; the few prospective checks show mixed regressions, so the deployment-robustness claim is not yet established.","rationale":"The reader's verdict of CONDITIONAL is appropriate. The negative audit (rankings reverse) is robust and well-supported by the extensive controlled study. The positive CEG claim is the vulnerable part: hyperparameters and the protected weight were selected using the same task domains on which the headline results are reported, and the genuinely prospective checks show mixed regressions. This is an acknowledged, disclosed limitation, but it means the central 'deployment robustness' contribution is not yet confirmed by out-of-development evidence. My concern matches the reader's weakest assumption, and the proposed new-domain test would settle whether CEG transfers. Since the reader already assigned CONDITIONAL, no verdict change is needed.","tokens_in":33570,"tokens_out":5271,"duration_ms":48429,"concrete_test":"Freeze CEG-P with λ=1/3 and all other hyperparameters exactly as specified, and evaluate it on two ID domains that were not part of the development suite – Oxford Flowers-102 and Stanford Cars – using their OpenOOD-v1.5 near/far OOD pools, CLIP-B/16, and the Appendix B protocol (1,000 ID calibration, disjoint 1,000 operating split, full test images). If CEG-P fails to improve family-balanced FPR95 over the raw MCM baseline on both new domains, or if it regresses on either, the deployment-robustness claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central negative finding – detector rankings reverse across ID domains and VLMs – is well supported by the seventeen-domain audit and cross-VLM port, so I do not object to it. The load-bearing weakness is in the positive CEG claim. Appendix B 'Evidence tiers' discloses that the channel forms, k=10, Q0.9, 3×3 kernel, and minimum fusion were developed on capped versions of the same five strict tasks used for the headline results; the uncapped reruns are sampling-level stability evidence, not confirmatory. Appendix E 'Protected-weight provenance' states the 2:1 λ ratio was chosen after inspecting the transfer failures it was designed to mitigate. The only genuinely prospective checks are the reciprocal ImageNet-10/20 experiment (still ImageNet, hence same domain) and the SigLIP2/PE port, and the port shows five of fifteen point estimates worsen, with PE CIFAR-100 unresolved and CEG(MaxLogit) worsening the SigLIP2 aggregate. Thus the paper's headline 'reduces deployment sensitivity' rests on development/evaluation overlap, not on held-out task evidence. This is an acknowledged, disclosed limitation, but it materially weakens the central positive claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether zero-shot OOD detector rankings transfer across deployment settings. Across seventeen ID datasets, three VLMs, and seven detectors, the authors find that rankings reverse with ID domain and VLM, that every detector exceeds 80% FPR95 on at least one domain, and that the preferred detector depends on both the ID data and the encoder. They attribute these reversals to complementary evidence channels: absolute match level, relative/spatial sharpness, and, for WordNet-based methods, external semantic coverage. A proposition shows that MCM discards level and approximates sharpness, while Energy preserves level, explaining complementary failures. The authors then propose CEG, a wrapper combining a base detector with level and sharpness channels through an empirical-percentile minimum fusion. On five strict OpenOOD tasks, CEG lowers family-balanced FPR95 for all wrapped detectors, and a broader audit shows contraction in detector spread. The manuscript is unusual in that it explicitly labels evidence tiers and discloses development/evaluation overlap in the appendices.","tokens_in":33850,"tokens_out":3002,"duration_ms":32321,"significance":"If the negative portability finding is accepted, it is a timely and important caution for the zero-shot OOD detection community: benchmark rankings on ImageNet-centred suites do not transfer across ID domains or VLM families. The paper's strengths include a controlled audit with frozen implementations, paired resampling intervals, a benchmark-composition probe, and an unusually transparent evidence-tier disclosure. The positive CEG proposal is well motivated by the evidence-channel diagnosis and the channel-substitution controls are informative; however, as discussed below, the deployment-robustness claim for CEG is not yet established by held-out task evidence. The paper's contribution is therefore strongest as a diagnosis/audit and weaker as a validated new method.","major_comments":[{"comment":"The headline CEG results in Table 1a are measured on the same five strict task domains whose capped versions were used to develop the channel forms, k=10, Q0.9, 3x3 kernel, and minimum fusion. As Appendix B states, the uncapped rerun is sampling-level stability evidence, not new-task confirmation. The only prospective checks are the reciprocal ImageNet-10/20 experiment (still ImageNet, hence same domain and correlated subsets) and the SigLIP2/PE port. The port shows five of fifteen point estimates worsen, including both SigLIP2 ImageNet tasks and PE CIFAR-100 unresolved, and CEG(MaxLogit) worsens the SigLIP2 aggregate. Thus the claim in the abstract that CEG 'reduces detector sensitivity' across deployments rests on development/evaluation overlap rather than on held-out task evidence. The claim should be narrowed or supported by a genuinely out-of-development evaluation (e.g., new ID dat","section":"§5 / Appendix B (Evidence tiers)"},{"comment":"The 2:1 shrinkage ratio for CEG-P was selected after inspecting the transfer failures it was designed to mitigate, and the appendix correctly labels this as post-result development evidence. However, Section 5 and Figure 7 present CEG-P as a protected variant and report broad all-domain improvements, and Table F1 is titled 'protected audit.' Because the weight was chosen after observing the failures, all CEG-P comparisons in Tables F1-F4 are circular for that variant. The paper should either clearly present CEG-P as development evidence in the main text, or provide a prospective validation of the protected weight on tasks not used in its selection.","section":"Appendix E (Protected-weight provenance)"},{"comment":"The broad CEG-P audit claims improvements in all 15 near/far deployment summaries and that every interval excludes zero. But the underlying point estimates include multiple regressions, and the intervals are described as nominal and unadjusted in Appendix B. The symmetric fold-wise control is a good sanity check, but it does not address the more serious problem that the audit is post-result development evidence. The strong language 'deployment robustness' and 'reduces deployment sensitivity' is not warranted by the disclosed evidence tier. I would ask the authors to separate confirmatory results from descriptive development evidence in the main-text claims.","section":"§5 / Appendix F (Broad deployment audit)"},{"comment":"The expansion for MCM at large temperature is correct and helpful, but the claim that MCM 'preserves only the sharpness channel' should be stated with the caveat that the expansion is asymptotic and holds for T >> R; the empirical approximation is verified in the appendix, but the main-text phrasing overstates exactness. This is not a blocking issue, but it affects the interpretive precision of the mechanism argument.","section":"Proposition 1 / §4"}],"minor_comments":[{"comment":"Several rows appear malformed or have misaligned column counts (e.g., the CUB-200 row). Please check the table rendering and ensure each detector has a value in every row.","section":"Table B1"},{"comment":"The notation z_ID for standardization is used before the mean/standard deviation are defined; state explicitly that z_ID denotes (value - mean)/std estimated on the unlabeled ID calibration set.","section":"Eq. (1)"},{"comment":"'Family-balanced' is used in the abstract and Section 5 but defined only in the appendix; give a one-sentence definition in the main text.","section":"Main text"},{"comment":"'All-five winner' and 'frozen global subset' are referenced in the caption before these terms are defined in the text; define or rephrase.","section":"Figure 2 caption"},{"comment":"The text reports 896/1,185 point improvements for CEG-P; given the large number of comparisons, a multiplicity correction or explicit statement that these are unadjusted descriptive counts would be useful.","section":"Appendix F"}],"recommendation":"major_revision","confidential_remarks":"The paper's honest evidence-tier appendix actually strengthens its credibility, but it also exposes that the central positive CEG claim is not yet confirmed by out-of-development data. The negative portability result is solid and could support acceptance if the positive claim were reframed as a hypothesis with preliminary evidence. I would encourage the editor to ask for either a prospective validation (e.g., a new VLM or ID dataset not touched during CEG development) or a substantial rhetorical downgrade of the deployment-robustness claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — two things to know. The negative result is solid: a controlled 17-dataset, 3-VLM audit shows zero-shot OOD detector rankings reverse across in-distribution domains and across backbones, and every detector exceeds 80% FPR95 on at least one domain. That finding is well-supported and worth taking seriously. The proposed fix, CEG, is not yet established as a deployment-robustness claim. The headline gains are measured on the same task domains used to select its hyperparameters and protected weight; the few prospective checks show mixed regressions. The paper discloses all this, which is a point in its favor, but it means the positive claim rests on development/evaluation overlap.\n\nWhat's new: the portability audit itself, the level/sharpness decomposition (Proposition 1 is elementary but gives a useful lens on MCM and Energy), and the CEG wrapper idea. The benchmark-composition probe with the verified Places/ImageNet duplicate is careful. The paper also does something rare: it explicitly tiers its evidence and labels what is prospective vs. development. That makes the paper useful even where I disagree.\n\nSoft spots: the evaluation of CEG. Hyperparameters (k, Q, kernel, min fusion) were developed on capped versions of the same five strict tasks used for headline results; CEG-P's lambda was selected after seeing transfer failures. The prospective SigLIP2/PE port shows five of fifteen point estimates worsening, and CEG(MaxLogit) worsens the SigLIP2 aggregate. The reciprocal ImageNet-10/20 check is still ImageNet, same domain. So the 'deployment sensitivity' reduction is not yet shown to hold out of development. No public code or data either.\n\nNone of this undercuts the central negative finding, which stands independently. The worst-domain risk reporting and the call to revalidate detectors when ID or VLM changes are sensible and actionable. The paper is honest about its own evidentiary limits, and the audit is broad enough that even a skeptical reader learns from it.\n\nWho this is for: anyone working on zero-shot OOD detection or benchmark evaluation. They should read the audit and weigh the negative result heavily. The CEG method is worth watching but not ready for deployment claims. I'd send this to peer review: the negative finding alone justifies referee time, and the method, while questionably evaluated, is clearly presented and its limitations candidly named. I'd ask for held-out task validation and code release before accepting.","headline":"Solid negative audit of OOD detector ranking transfer, but the CEG fix is developed and evaluated on the same tasks — treat the positive claim as promising but unproven.","tokens_in":34371,"tokens_out":1919,"would_cite":true,"duration_ms":20870,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Zero-shot OOD detector rankings reverse across domains and models, so benchmark winners are not deployment winners.","keywords":["zero-shot OOD detection","vision-language models","detector portability","complementary evidence","non-compensatory fusion","level and sharpness","FPR95","CLIP"],"falsifier":"Run CEG with frozen hyperparameters on a new, non-ImageNet-like ID domain (for example, medical imaging or aerial scenes) using CLIP features and show that the family-balanced FPR95 does not improve over the best raw base detector; or demonstrate that a detector explicitly preserving both level and sharpness through learned fusion outperforms CEG on a prospective benchmark, contradicting the claim that the minimum rule's gains come from evidence preservation.","tokens_in":33429,"feed_emoji":"🛡️","tokens_out":3143,"duration_ms":36784,"temperature":0.7,"pith_summary":"This paper tries to establish that the common practice of choosing a zero-shot OOD detector by benchmark ranking is unsafe: rankings reverse when the in-distribution dataset or the underlying vision-language model changes, and every detector exceeds 80% FPR95 on at least one domain. It explains these reversals by showing that detectors rely on complementary evidence channels — absolute match level, relative or spatial sharpness, and external semantic coverage — and that no single channel is universally informative. A simple proposition shows that level and sharpness cannot generally be recovered from one another, which is why no fixed detector transfers reliably. The paper then proposes the Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves all channels by taking the minimum of their empirical in-distribution percentiles, and reports that this reduces FPR95 for every wrapped detector while shrinking the spread between detectors.","feed_headline":"OOD detector rankings do not transfer across domains","feed_subtitle":"Every detector fails somewhere; a level-sharpness guard cuts worst-case error for all.","key_machinery":"The central objects are the complementary evidence channels extracted from CLIP-style cosine logits: level (the maximum cosine match to ID anchors) and sharpness (how much the best match stands out from the mean, plus a spatial sharpness term). Proposition 1 identifies MCM as a shift-invariant approximation to the centred max, preserving only sharpness, while the log-sum-exp decomposition shows Energy preserves level to leading order. CEG maps the base score, level, and sharpness each to empirical in-distribution percentiles and then takes their minimum, enforcing a non-compensatory veto: any single atypical channel rejects the sample. The minimum is the load-bearing mechanism that directly","core_discovery":"The central claim is that zero-shot OOD detector rankings are not portable across deployments. Through a controlled audit across seventeen in-distribution datasets, three vision-language models, and seven detectors, the paper shows that the best detector changes with both the ID domain and the VLM, and that no detector stays below 80% FPR95 everywhere. The explanation is that detectors preserve different evidence: softmax-based scores like MCM discard absolute level and keep only sharpness, energy and MaxLogit keep level, and WordNet-based methods depend on external corpus coverage. Proposition 1 shows MCM is shift-invariant and approximates the centred maximum, while Energy preserves level","pith_inferences":["The level/sharpness decomposition likely applies beyond vision-language models to any logit-based classifier, so a similar non-compensatory guard could stabilize OOD detection in standard deep networks.","The strong dependence of WordNet-based methods on corpus coverage suggests that future external-semantic detectors should mine or adapt the negative vocabulary to the target ID domain rather than reuse a fixed pool.","The paper's own appendix flags that most headline CEG results were developed and evaluated on capped versions of the same five tasks; the only truly prospective checks are the ImageNet-10/20 reciprocal test and the SigLIP2/PE port, several cells of which regress, so the transferability claim is weaker than the headline numbers suggest.","The non-compensatory veto principle could be useful in other safety-critical fusion settings where distinct cues fail independently and a single atypical channel should be enough to raise an alarm."],"forward_implications":["Benchmark averages alone are insufficient for detector selection; worst-domain risk and revalidation across ID domains and VLMs should be reported.","Detector rankings should be re-checked whenever the in-distribution data or the underlying vision-language model changes, since both flip the preferred method.","Preserving complementary evidence rather than committing to a single score is a more robust deployment strategy, with CEG improving every wrapped detector without OOD samples or auxiliary corpora.","The vocabulary-size confound (growing K alone changes detector rankings) implies that benchmark comparisons should control for label-set size.","CEG reduces sensitivity to detector choice, but it does not crown a universal winner; the best guarded detector still shifts across evaluation weightings."],"fun_headline_variants":["OOD detector rankings reverse across deployments","Which OOD detector wins? Depends on the domain","No detector stays under 80% FPR95 everywhere","Complementary evidence guard cuts worst-case OOD error"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that CEG's hyperparameters and the level/sharpness channel definitions, tuned on capped versions of the same five strict tasks, transfer to new ID domains and VLMs; the paper's own evidence tiers show that most headline results come from development/evaluation overlap, with only a handful of genuinely prospective tests.","fun_headline_variants_meta":{"raw":{"variants":["OOD detector rankings reverse across deployments","Which OOD detector wins? Depends on the domain","No detector stays under 80% FPR95 everywhere","Complementary evidence guard cuts worst-case OOD error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1233,"prompt_tokens":808,"completion_tokens":425,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":552,"tokens_out":425,"duration_ms":4696,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T12:52:31.559423+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CEG with frozen hyperparameters on a new, non-ImageNet-like ID domain (for example, medical imaging or aerial scenes) using CLIP features and show that the family-balanced FPR95 does not improve over the best raw base detector; or demonstrate that a detector explicitly preserving both level and sharpness through learned fusion outperforms CEG on a prospective benchmark, contradicting the claim that the minimum rule's gains come from evidence preservation.","supporting_citations":[],"review_version":1}