{"id":"65c604eb-0ab2-4c5a-9095-899ad74b5d50","arxiv_id":"2607.24878","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new auditable benchmark that jointly measures target retrieval, hard-confounder discrimination, and evidence access in rare-disease diagnosis.","lead":"GraphRareBench is a new benchmark for phenotype-driven rare-disease diagnosis that keeps the correct disease hidden while forcing systems to rank it against graph-defined 'hard' lookalikes, and that records which evidence each system examines. It shows that retrieval accuracy, confounder discrimination, and evidence-seeking behavior are distinct abilities—a finding that matters for building trustworthy diagnostic AI.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hard-confounder clinical validity is the load-bearing assumption: no external validation shows graph-defined alternatives are plausible differentials, so ToC-based complementarity could measure graph similarity rather than diagnostic difficulty.","rationale":"The reader's weakest assumption is exactly the most load-bearing point: graph-defined hard confounders are assumed to be clinically plausible alternatives, and every ToC-based claim inherits that assumption. The paper's own limitation statement acknowledges incomplete coverage but does not provide external validation, so the conditional verdict is appropriate. I considered whether feature-label overlap in the supervised 21-feature interface is a stronger concern, since those features are never enumerated and could encode the same graph relations used to define confounder labels; however, the primary load-bearing uncertainty remains the clinical validity of the confounder construct itself. The proposed clinician-study test would directly resolve this uncertainty. The reader's CONDITIONAL verdict already reflects this condition, so no change in verdict is needed.","tokens_in":11631,"tokens_out":8746,"duration_ms":83368,"concrete_test":"Select 150 target-confounder pairs stratified by the seven confounder mechanisms. Present two clinical geneticists with the coarsened HPO query and candidate list, without graph labels, and have them independently mark which candidates they would include in a differential diagnosis. Pre-specify agreement thresholds (e.g., at least 70% of graph-hard pairs judged clinically plausible and Cohen's kappa >= 0.6). If actual agreement falls below threshold, ToC-based complementarity conclusions should be reframed as measuring graph similarity rather than clinical differential difficulty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that full-pool retrieval, hard-confounder discrimination, and observable evidence access are complementary rests on interpreting target-over-confounder (ToC) accuracy as clinically meaningful discrimination. GraphRareBench mines hard confounders from HPO overlap, SapBERT semantic neighbors, shared MONDO parents, disease-family membership, and causal-gene relations. These are graph-similarity heuristics; the paper provides no evidence that they correspond to real differential-diagnosis difficulty. The Limitations section acknowledges that the relations 'may not cover every distinction encountered in clinical practice,' but that understates the issue: no expert validation, no comparison against curated differential-diagnosis fields from OMIM/Orphanet, and no clinician study supports the 'hard' label. Consequently, AnyConfAbove and ToC values (0.898–0.916) may partly reflect semantic/ontological closeness rather than clinical plausibility. This is especially important because the headline complementarity finding includes the 22.1–43.7% of Hit@10 cases with a graph-defined confounder above the target; if those confounders are not realistic alternatives, that failure is not necessarily a diagnostic failure. The issue is not internal inconsistency—the construction is deterministic and auditable—but external validity of the core construct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GraphRareBench, a provenance-preserving benchmark for phenotype-driven rare-disease diagnosis. It constructs 2,365 ontology-derived cases with coarsened HPO queries, fixed candidate pools, graph-defined hard confounders, and source-linked evidence records. The benchmark is split by causal-gene components to reduce leakage, yielding 237 test cases. The authors evaluate negative controls, phenotype-driven tools (Exomiser, LIRICAL, HPO Similarity), a prompted LLM (DeepSeek-V4-Flash), supervised graph-evidence rankers (RAG2D, PPP, LPP) sharing a 21-feature interface, and tool-using agents. They report MRR, Hit@k, and target-over-confounder (ToC) metrics, and find that full-pool retrieval, hard-confounder discrimination, and evidence-access behavior capture complementary aspects of model performance. The headline results include MRRs of 0.640–0.740 for supervised rankers, ToCcase of 0.898–0.916, and an agent MRR difference of 0.029 that is not statistically significant while evidence coverage differs by 0.561.","tokens_in":11952,"tokens_out":7356,"duration_ms":70409,"significance":"If the construct validity of the 'hard confounder' labels is accepted, GraphRareBench would fill an important gap by adding differential-diagnosis-focused metrics and auditable evidence traces to phenotype-driven ranking benchmarks. The paper is careful in several respects: it performs a label-leakage screen, uses a gene-component-disjoint split, mines confounders deterministically, reports bootstrap confidence intervals for MRR, checks tail-ranking sensitivity, and provides deterministic compliance checks for agent tool use. These controls strengthen the internal validity of the comparisons. The main uncertainty is external: the graph-defined hard confounders are not shown to correspond to clinically plausible differential diagnoses, which is load-bearing for the interpretation of ToC and the complementarity claim. The omission of the 21-feature interface specification also limits reproducibility and prevents assessment of a possible circularity between feature construction and confounder-label generation.","major_comments":[{"comment":"The label 'hard' is operationalized by seven graph/semantic relations (HPO overlap, SapBERT, MONDO parents, disease-family membership, causal-gene relations). The abstract and Q1 interpretation use ToC and the 22.1–43.7% Hit@10 failures to draw conclusions about 'clinically plausible alternatives.' However, no evidence shows that these graph-defined alternatives correspond to real differential-diagnosis difficulty; the Limitations only concede that the relations 'may not cover every distinction encountered in clinical practice.' This is a construct-validity problem, not merely a coverage gap. Please add a validation against curated differential diagnoses (e.g., OMIM/Orphanet differential fields or expert review of a sampled set of pairs). Without this, ToC should be described as 'graph-defined alternative ordering,' and the complementarity claim should be correspondingly qualified.","section":"GraphRareBench Construction / Limitations"},{"comment":"The supervised rankers RAG2D/PPP/LPP are defined by their 'shared 21-feature interface,' and Q4's score-time feature-channel isolation (Table 2) is central to the evidence-audit claim. Yet the main text never defines the 21 features; Table 2 only names feature groups such as 'case-specific evidence' and 'aggregate graph scores.' Without a precise list (which HPO-overlap features, semantic similarity scores, graph-distance features, gene-context features; how they are normalized; which resources they consume), the supervised results are not reproducible and the circularity analysis below cannot be performed. Please provide the full feature specification in the main text or a named appendix.","section":"Compared methods / Table 2"},{"comment":"ToC for the supervised rankers is at risk of being partially self-fulfilling. Hard confounders are mined using HPO overlap, SapBERT semantic similarity, shared MONDO parents, disease-family membership, and causal-gene relations; the supervised models are trained on graph-evidence features plausibly derived from the same relations. If so, the reported ToC values of 0.898–0.916 may reflect, in part, learning the label-generation rule rather than a general differential-diagnosis ability. The complementarity conclusion is partly supported by unsupervised HPO Similarity and agent evidence, so the issue is not fatal to all claims. Still, the paper should either ablate feature channels that correspond to confounder-defining relations or evaluate on a held-out confounder mechanism to show the ToC result is not an artifact of label-feature overlap.","section":"Q3 / GraphRareBench Construction"},{"comment":"The agent audit is based on one fixed protocol run per agent. The main text acknowledges this, but the Abstract reports 'their target-evidence coverage differed by 0.561' without the fixed-run caveat. Because the evidence-access component of the complementarity claim rests on this single run, the observed tool-use differences could reflect decoding variability, prompt sensitivity, or candidate order rather than stable agent properties. Please provide repeated-run variance estimates (e.g., multiple decoding seeds/temperatures) or explicitly present these results as an illustrative case study and soften the corresponding conclusion in the Abstract and Discussion.","section":"Experimental Protocol / Table 3"}],"minor_comments":[{"comment":"Typo: 'avaliable' should be 'available'.","section":"Abstract"},{"comment":"Bootstrap confidence intervals are reported only for MRR. Since ToCcase is a central metric, please report CIs for ToCcase as well.","section":"Table 1"},{"comment":"The reason-sliced ToCpair values are presented without confidence intervals, and some slices are small (e.g., semantic slice has 100 pairs). Please show case-clustered CIs or at least give the pair counts in the figure itself.","section":"Figure 3"},{"comment":"A few citations appear malformed: 'AlDin, Z. E.' should likely be 'Al Din, Z. E.', and 'Ma, G.; NM, B.' appears to have a formatting artifact. Please check the reference list for consistency.","section":"References"},{"comment":"The tool-use differences (evidence calls, switches, coverage) are descriptive and no formal multiple-comparison correction is applied. Since the primary comparison of MRR is not significant, the coverage differences should be explicitly labeled as exploratory.","section":"Q4 / Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a careful and potentially valuable resource, with strong internal controls (leakage screening, gene-disjoint split, deterministic construction, bootstrap CIs). The main risk is construct validity of the 'hard confounder' label; without external validation, the ToC-based complementarity claim is fragile. A targeted validation against curated differentials or an ablation of confounder-defining features would substantially strengthen the manuscript. The missing 21-feature specification is also a blocking issue for reproducibility and for judging circularity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"GraphRareBench is a solid, carefully constructed benchmark for phenotype-driven rare-disease diagnosis, and it does something genuinely new: fixed candidate pools, graph-defined hard confounders, and tool-trace auditing. The central finding—that full-pool retrieval, confounder ordering, and evidence access capture different behaviors—holds up.\n\nThe construction is careful: gene-component-disjoint splits, label-leakage screening, deterministic confounder mining, bootstrap CIs, and a score-time feature isolation that exactly reproduces stored rankings. The agent audit is the most instructive result: two agents with statistically indistinguishable MRR (0.718 vs 0.746) differed by 0.561 in target-evidence coverage. That alone justifies the paper's argument that rank metrics miss what matters.\n\nThe soft spot is the clinical validity of the hard confounders. They are mined from HPO overlap, SapBERT neighbors, MONDO parents, disease-family membership, and gene relations, but there is no expert validation or curated differential-diagnosis reference set. So ToC values could partly measure graph similarity rather than diagnostic difficulty. The paper acknowledges this in Limitations, but it's understated: the 'hard' label is a design choice, not a clinical fact. For the supervised rankers there is also a genuine overlap between the 21 features and the confounder definitions, which makes their high ToC partly self-fulfilling. That said, the complementarity claim does not rest only on those rankers—HPO similarity also shows high ToC with low MRR, which is consistent with the claim.\n\nMinor issues: the 21-feature interface is not described in the main text, and the candidate-cap/threshold choices are arbitrary but documented in the supplement. The dataset is linked on GitHub but not inspectable from the text, so reproducibility is conditional on the repo.\n\nVerdict: worth a serious referee. A revision should either validate the confounders against curated clinical data or carefully relabel them as 'graph-defined alternatives' and avoid claiming clinical plausibility. I'd cite this and bring it to a reading group.","headline":"A solid benchmark that adds evidence auditing to rare-disease diagnosis evaluation, but the clinical validity of its 'hard' confounders needs external support.","tokens_in":12408,"tokens_out":3367,"would_cite":true,"duration_ms":28419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark shows that rare-disease diagnostic systems must be judged not only on whether the true disease ranks high, but also on whether they outrank plausible alternatives and on what evidence they use.","keywords":["rare disease diagnosis","phenotype-driven ranking","benchmark","target-over-confounder accuracy","hard confounders","HPO phenotype queries","evidence audit","LLM agents"],"falsifier":"Take a random sample of 200 target–confounder pairs from GraphRareBench and ask clinicians to rate whether each pair is a realistic differential-diagnosis dilemma. If a large fraction of the pairs are judged implausible as competing diagnoses, or if pairs without any graph edge are rated equally plausible, then the graph-defined difficulty assumption fails and the ToC metrics measure artifacts rather than clinical competence.","tokens_in":11550,"feed_emoji":"🧬","tokens_out":4473,"duration_ms":38811,"temperature":0.7,"pith_summary":"GraphRareBench argues that phenotype-driven rare-disease diagnosis systems should not be evaluated solely by top-k retrieval of the true disease. It provides 2,365 ontology-derived cases with fixed candidate pools, graph-defined hard confounders, and source-linked evidence records, enabling audits of a system's ability to distinguish the target from close alternatives and to access relevant evidence. On a gene-component-disjoint test split, supervised rankers reach MRRs of 0.640–0.740 and target-over-confounder accuracies of 0.898–0.916, while two tool-using agents achieve statistically indistinguishable MRRs yet differ by 0.561 in target-evidence coverage. The paper shows that 22.1% to 43.7% of Hit@10 successes still have at least one hard confounder ranked above the target, indicating that full-pool retrieval, hard-confounder discrimination, and evidence access measure complementary failure modes.","feed_headline":"Hit@10 success still hides a confounder above the true diagnosis","feed_subtitle":"GraphRareBench shows 22–44% of top-10 hits keep a plausible alternative above the target, so retrieval scores miss real failures.","key_machinery":"The benchmark's load-bearing elements are (1) the fixed candidate pool, split into a full pool and a hard pool containing only the target plus graph-defined hard confounders; (2) the coarsened HPO query, where 97.7% of terms are one ontology step up from the source phenotype, reducing direct phenotype matching while preserving provenance; and (3) the ToC metrics, which compute the fraction of target–confounder pairs where the target outranks the confounder. Hard confounders are mined deterministically through seven graph mechanisms: high HPO overlap, SapBERT semantic neighbors, shared MONDO parents, disease-family membership, shared causal-gene families, shared causal-gene components, and ex","core_discovery":"The central claim is that conventional top-k metrics mask clinically meaningful failures in phenotype-driven rare-disease diagnosis: a model can achieve Hit@10 while ranking a plausible graph-defined alternative above the true disease. To expose this, the authors construct GraphRareBench, a provenance-preserving benchmark of 2,365 cases and 18,093 target–confounder pairs built from HPO, MONDO, disease-family, and causal-gene relations, and propose target-over-confounder (ToC) metrics that measure whether the target outranks its hard confounders. They also record which evidence a tool-using model requests, allowing an evidence-access audit separate from final ranking. The results show that tw","pith_inferences":["If real clinical differentials involve features not captured by ontology/semantic/gene relations (e.g., overlapping imaging, treatments, or age-of-onset patterns), the ToC scores may overstate or understate true diagnostic difficulty; this could be tested by comparing graph-defined confounders with clinician-selected alternatives.","The large evidence-coverage gap between the two agents suggests a testable hypothesis: systems that access more target-specific evidence before ranking may be more robust to perturbations like candidate-order scrambling or restricted tool budgets.","The benchmark format could be extended to a phenotype-disjoint split (no shared HPO terms between train and test) to probe generalization across phenotype space rather than just gene components.","Because evidence is source-linked, GraphRareBench could be adapted to measure shortcut reliance: a model that ranks correctly without querying the evidence tool may be relying on memorized disease–phenotype associations, which the framework makes detectable."],"forward_implications":["If top-k scores are not supplemented with ToC metrics, failures where a plausible alternative outranks the true disease will remain invisible.","Tool-using agents with similar final ranks can have very different evidence-seeking behavior; trace-based evidence coverage should be part of the evaluation.","The fixed candidate pools and evidence interfaces allow controlled interventions to isolate whether a failure is due to missing evidence, poor retrieval, or poor integration.","GraphRareBench can serve as a regression suite for updates to models, retrievers, or knowledge bases across confounder mechanisms.","The gene-component-disjoint split tests generalization beyond the causal genes seen during training, a specificity not offered by most existing benchmarks."],"fun_headline_variants":["Hit@10 can still rank a confounder above the true diagnosis","Top-10 hits hide a confounder above the target in rare-disease AI","GraphRareBench: 22–44% of top-10 ranked diagnoses misorder","Rare-disease retrieval: top-10 rank doesn't ensure correct diagnosis","Evidence audit reveals confounder blind spots in top-10 hits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the graph-defined hard confounders—mined from HPO overlap, semantic neighbors, shared ontology parents, disease-family membership, and causal-gene relations—represent the clinically plausible alternatives that a diagnostic system should rule out; if they do not correspond to real differential-diagnosis difficulty, the target-over-confounder metrics lose clinical meaning.","fun_headline_variants_meta":{"raw":{"variants":["Hit@10 can still rank a confounder above the true diagnosis","Top-10 hits hide a confounder above the target in rare-disease AI","GraphRareBench: 22–44% of top-10 ranked diagnoses misorder","Rare-disease retrieval: top-10 rank doesn't ensure correct diagnosis","Evidence audit reveals confounder blind spots in top-10 hits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1339,"prompt_tokens":836,"completion_tokens":503,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":403}},"tokens_in":580,"tokens_out":503,"duration_ms":5802,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:00:34.026097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 200 target–confounder pairs from GraphRareBench and ask clinicians to rate whether each pair is a realistic differential-diagnosis dilemma. If a large fraction of the pairs are judged implausible as competing diagnoses, or if pairs without any graph edge are rated equally plausible, then the graph-defined difficulty assumption fails and the ToC metrics measure artifacts rather than clinical competence.","supporting_citations":[],"review_version":1}