{"id":"0c5bc342-fca1-4e57-bde8-f1127ea8a5c2","arxiv_id":"2606.16149","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A clinician-audited diagnostic policy plus public tools lets a single unmodified LLM reach high phenotype-first rare-disease Recall@1 and modestly beat baselines on real UDN patients.","lead":"LiteOdyssey turns rare-disease diagnostic reasoning into a portable natural-language policy that steers an unmodified LLM with public biomedical tools. On phenotype-only benchmarks and 515 hard Undiagnosed Diseases Network cases, the structured policy lifts accuracy without fine-tuning or large case banks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Primary LLM-as-judge scoring can credit synonym-level disease matches that exact-OMIM rejects, so the SOTA R@1 claim may overstate identifier-level diagnostic correctness.","rationale":"The reader correctly flags evaluation softness (LLM judge vs exact-OMIM), phenotype-only scope, curated benchmarks, and modest UDN absolute accuracy as the main risks, and already lands on CONDITIONAL pending stricter scoring transparency and release. That is the right load-bearing concern for the strongest claim: absolute SOTA R@1 numbers and cross-system comparisons depend on a permissive primary judge, while the paper’s own sensitivity shows lower exact-OMIM rates. I do not invent a stronger objection—the ablations, held-out splits, Qwen transfer, and UDN McNemar gains still support a real structured-environment effect. No change of verdict category is warranted; the stress test sharpens the same condition the reader already imposed (stricter scoring transparency / exact-OMIM primary reporting). Confidence remains moderate for the same reasons (closed APIs, private UDN, URL TBD).","tokens_in":17035,"tokens_out":649,"duration_ms":5829,"concrete_test":"Re-score all public full-system and parametric-baseline top-1 outputs under pure exact-OMIM (and, if available, a fixed synonym-normalized OMIM map) for LIRICAL n=370 and PhenoPacket n=873; report R@1/R@5 side-by-side with judged scores and recompute the DeepRare comparison on the same strict metric. If exact-OMIM full-system R@1 falls below the best published comparator under identical scoring, or the environment lift shrinks by >10 absolute points, the SOTA absolute claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is state-of-the-art phenotype-first disease Recall@1 of 58.6% (LIRICAL) and 59.6% (PhenoPacket) for a single unmodified LLM under a PIHF policy plus public tools. Primary scoring is an LLM-as-a-judge that can accept synonym-level agreement; exact-OMIM is only a sensitivity analysis (Methods §5.6). Appendix Table A1 / Figure A1 show the stricter exact-OMIM numbers are lower (e.g., held-out LIRICAL full-system R@1 54.4% vs 57.5% judged; UDN baseline also drops). DeepRare-style comparators and the headline SOTA framing appear to rest on the more permissive judged metric. If a non-trivial fraction of “correct” top-1s are alias/synonym credits rather than OMIM-identity matches, the absolute SOTA claim and the size of the structured-environment lift are softer than presented—especially given phenotype-only inputs and much lower absolute UDN yield. The direction of the tool/policy lift is still supported by ablations and McNemar tests, but the load-bearing absolute ranking claim is not fully secured by the primary metric.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"LiteOdyssey is a single-agent, phenotype-first rare-disease diagnostic system that guides an unmodified reasoning LLM with a natural-language eight-phase clinical genetics policy (developed via Policy Iteration with Human Feedback, PIHF) and eight public/cached biomedical tools. On LIRICAL (n=370) and the PhenoPacket Store (n=873), the system reports disease Recall@1 of 58.6% and 59.6%, large lifts over a same-backbone tool-free baseline (e.g., +23.5 and +36.7 points), transfer of the environment gain to Qwen3.6-35B without re-tuning, and statistically significant but smaller gains on 515 UDN patients (R@1 20.4% vs 16.7%, McNemar p=0.027). The authors argue that externalizing diagnostic reasoning as a portable, auditable policy can match or exceed heavier multi-agent/retrieval/fine-tuning systems while remaining deployable and inspectable.","tokens_in":17380,"tokens_out":1640,"duration_ms":25535,"significance":"If the core result holds under stricter, comparator-aligned scoring, the paper is significant for medical AI: it shows that much of rare-disease diagnostic yield can come from organizing inference-time reasoning and public tool use rather than from fine-tuning, multi-agent orchestration, or large solved-case banks. Strengths that should be credited include same-backbone tool ablations, development-excluded public cases, model-level holdout on an open-weights backbone, a private multi-site UDN cohort with paired McNemar tests, gene-level secondary endpoints, exact-OMIM sensitivity preserving direction, and worked reasoning-trace cases that illustrate rescues and regressions. PIHF as a weight-free, clinician-auditable policy iteration process is a useful methodological contribution for settings where model weights cannot be modified.","major_comments":[{"comment":"Methods §5.6 and Results §3.1: Primary disease Recall@1 is defined by an LLM-as-a-judge that can credit synonym/alias agreement, while deterministic exact-OMIM is only a sensitivity analysis. Appendix Figure A1 / Table A1 show material absolute drops under exact-OMIM (e.g., held-out LIRICAL full-system R@1 54.4% vs 57.5% judged). The headline SOTA claim (58.6% LIRICAL; 59.6% PhenoPacket; Figure 1 vs DeepRare 56.0%) is therefore not fully secured at identifier level. Please report exact-OMIM R@1/R@5 as co-primary (or primary) for all main tables/figures, quantify the fraction of judged-only top-1 credits, and restate absolute SOTA claims only where they survive the stricter metric.","section":"Methods §5.6; Results §3.1; Appendix Figure A1 / Table A1"},{"comment":"§3.1 comparator framing: DeepRare is the main published reasoning comparator (39.5% without retrieval; 51.6–56.0% with a 67,795-case retrieval corpus that can include benchmark-like cases). The manuscript scores LiteOdyssey with “DeepRare-style” matching but does not show that DeepRare’s published numbers were obtained under the same judge/alias protocol, nor re-score DeepRare outputs on the identical case set and metric. Without matched scoring (and an explicit statement of which DeepRare configuration/backbone is being compared), the absolute “higher still” SOTA claim is not load-bearing. Either re-evaluate under a shared protocol or demote absolute ranking language and emphasize the controlled same-model environment lift.","section":"Results §3.1; Figure 1–2"},{"comment":"Methods §5.3–5.4 / §5.8: PIHF and the eight-phase policy are central inventions, but only 50 LIRICAL and 50 UDN cases informed development, the actual policy artifact is not provided (code/demo “URL TBD”), and the iteration loop (how many rounds, which critiques were accepted/rejected, how phase-transition rules changed) is described at a high level. Because portability and auditability are core claims, the final natural-language policy, tool schemas, judge prompt, and a minimal PIHF change log should be released or included as supplementary material so that the weight-free policy—not only the model—can be inspected and re-run.","section":"Methods §5.3–5.4; §5.8"},{"comment":"§3.5 and Discussion: On the UDN cohort the absolute full-system disease R@1 is 20.4% (+3.7 points over baseline). The gain is statistically significant and directionally consistent on held-out cases, but the abstract and introduction lean heavily on this “external evaluation” as clinical support. Please keep absolute yield, selection for diagnostic difficulty, phenotype-only inputs, and the non-significant gene R@1 (p=0.21) co-equal with the p-values so readers do not over-read clinical readiness from a modest absolute top-1 rate.","section":"Abstract; Results §3.5–3.6; Discussion"}],"minor_comments":[{"comment":"Naming is inconsistent across abstract/title casing (liteOdyssey vs LiteOdyssey).","section":"Abstract / title"},{"comment":"Figure 1 deployment-burden axis (0.66 GB vs 100+ GB) needs a short methods note on what is counted (cached indices, tool DBs, model weights excluded/included) so the comparison is reproducible.","section":"Figure 1"},{"comment":"§3.2 states PhenoPacket parametric baseline 22.9% while Figure 2 panels use subset-specific baselines (e.g., 10.7% unmapped); clarify full-corpus vs subset baselines in one place.","section":"Results §3.2; Figure 2"},{"comment":"Appendix A worked cases are valuable; consider one additional regression/rescue pair from PhenoPacket or UDN (de-identified) to show the trace behavior outside LIRICAL.","section":"Appendix A; §3.7"},{"comment":"References include several 2025–2026 arXiv preprints and a Nature DeepRare citation with 2026 date; verify bibliographic completeness and version pins for reproducibility manifests.","section":"References; §5.8"},{"comment":"Minor prose issues: missing spaces in concatenated author/affiliation strings in references; “pointprevalenceofrare diseases” etc. in ref [1] rendering.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The controlled ablations and UDN evaluation make this stronger than many rare-disease LLM papers, but the absolute SOTA marketing depends on a permissive primary judge and an incompletely matched DeepRare comparison. If the authors re-center on environment lift + exact-OMIM + released policy artifact, this could become a solid methods/systems contribution; if they insist on unadjusted SOTA without metric alignment, I would remain negative. Scope fit for a serious AI/medicine venue is good if clinical claims stay proportionate to the 20% UDN top-1 yield."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: they show you can get strong phenotype-first rare-disease ranking from one unmodified reasoning model by wrapping it in a clinician-audited natural-language diagnostic policy and eight public tools, without fine-tuning, multi-agent stacks, or a solved-case retrieval bank. That is the real contribution.\n\nWhat is new is not tool use or agentic diagnosis—DeepRare and friends already exist—but PIHF: iterative policy revision from model critiques that a domain expert vets, producing a portable eight-phase workflow artifact. They do the hard empirical work. Same-backbone tool-free ablations give large lifts (e.g. +50 points R@1 on the unmapped PhenoPacket subset). Held-out LIRICAL (320) and all of PhenoPacket preserve the gain. The policy transfers to Qwen3.6 without retuning. On 515 real UDN patients the lift is smaller but significant (R@1 20.4% vs 16.7%, McNemar p=0.027). Gene-level secondary endpoints and exact-OMIM sensitivity keep the same direction. Citations are fair; they position against DeepRare honestly, including the retrieval-corpus issue.\n\nSoft spots, in proportion. Primary scoring is LLM-as-judge that can credit synonym-level matches; exact-OMIM is only sensitivity and is a few points lower (held-out LIRICAL ~54% vs ~57%). So the absolute SOTA framing is a bit softer than the headline, though the structured-environment lift is not an artifact of the judge. Phenotype-only inputs, curated public cases, and low absolute UDN yield are real limits they mostly own in the Discussion. Code/URL still TBD and closed APIs matter for full reproducibility. None of that collapses the central claim that the policy-plus-tools scaffold is doing real work.\n\nThis is for people building deployable medical AI and clinical genetics tooling who care about auditability and infrastructure cost. It deserves a serious referee. I would engage: cite the PIHF idea and the ablation design, and push for stricter primary scoring and public-benchmark release.","headline":"Solid systems paper: external clinician-vetted policy + public tools lifts phenotype-first rare-disease ranking without fine-tunes or case banks; SOTA numbers rest partly on a permissive LLM judge, but the direction of the effect is real.","tokens_in":18070,"tokens_out":544,"would_cite":true,"duration_ms":5598,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A portable clinical policy plus public biomedical tools lets one unmodified language model diagnose rare diseases at high accuracy without fine-tuning or large case banks.","keywords":["rare-disease diagnosis","large language models","policy iteration with human feedback","phenotype-first reasoning","tool-augmented agents","interpretable clinical AI","Undiagnosed Diseases Network","Human Phenotype Ontology"],"falsifier":"A prospective bedside trial that supplies the system with incomplete or evolving phenotypes plus laboratory and genetic data, then measures whether physician-adjudicated top-1 and top-5 accuracy still exceeds the same unmodified model without the policy and tools.","tokens_in":17886,"feed_emoji":"🧬","tokens_out":1017,"duration_ms":10939,"temperature":0.7,"pith_summary":"Rare-disease diagnosis is hard because it demands long chains of phenotype interpretation, evidence gathering, and differential refinement that most AI systems either split into isolated tools or support only by adding heavy infrastructure around the model. This paper asks whether the same end-to-end reasoning can instead live in a lightweight, human-auditable policy that steers a single general-purpose language model and a handful of public biomedical tools. The authors build LiteOdyssey by Policy Iteration with Human Feedback: clinicians inspect the model’s reasoning traces and iteratively revise a natural-language diagnostic policy without ever changing model weights. On two large phenotype-first benchmarks heavy with ultra-rare diseases, and on a private cohort of 515 real Undiagnosed Diseases Network patients, the structured policy produces large, transferable gains over the same models run without tools, while leaving every diagnosis open to step-by-step clinical review. The claim is that expert diagnostic reasoning can be externalized as a reusable policy layer rather than baked into weights or multi-agent machinery.","feed_headline":"One policy plus public tools lifts rare-disease AI to 59% top-1","feed_subtitle":"No fine-tuning, no case bank: a human-auditable clinical workflow steers unmodified models","key_machinery":"Policy Iteration with Human Feedback (PIHF): a weight-free loop in which a frozen model runs a natural-language diagnostic policy, clinicians review scores and full reasoning traces, and the retained critiques revise the policy itself—producing a portable, inspectable eight-phase workflow that dictates tool use, evidence weighing, and reflective adjudication.","core_discovery":"LiteOdyssey shows that a single unmodified reasoning language model, guided by an eight-phase clinical-genetics policy developed through Policy Iteration with Human Feedback and eight public or cached biomedical tools, reaches state-of-the-art phenotype-first disease Recall@1 of 58.6% on LIRICAL and 59.6% on the PhenoPacket Store, with large structured-environment lifts over the identical model without tools, transfer to an open-weights model never used in development, and statistically significant gains on 515 Undiagnosed Diseases Network patients.","pith_inferences":["If the policy is the true carrier of performance, later work can distill a stabilized PIHF document into a small adapter while keeping the original policy as the human-readable source of truth.","The same external-policy pattern could be tried on other multi-step clinical tasks that currently rely on multi-agent orchestration or heavy retrieval, such as oncology staging or complex infectious-disease workups.","Because the policy is natural language, institutions could maintain local forks that encode site-specific testing pathways without retraining foundation models.","Absolute accuracy remaining low on real undiagnosed cohorts suggests the next binding constraint is richer multimodal inputs rather than further policy refinement alone."],"forward_implications":["Diagnostic support for rare disease can be shipped as a readable policy document plus public tools rather than as fine-tuned weights or multi-gigabyte retrieval corpora.","The same policy can be executed unchanged across closed and open model families, so improvements travel without re-training.","Every differential comes with a step-by-step reasoning trace that a clinician can audit, contest, or revise.","Groups without large compute or curated case banks can still field competitive phenotype-first rare-disease AI.","PIHF offers a general pattern for turning expert diagnostic practice into portable, weight-free policy artifacts."],"fun_headline_variants":["Policy plus tools lifts unmodified LLMs to 59% rare-disease top-1","Human-feedback policy steers LLMs to better rare-disease accuracy","LiteOdyssey: auditable policy yields 59% top-1 rare-disease calls","One reusable policy improves diagnosis on 515 UDN rare-disease cases","No fine-tuning needed: policy layer lifts rare-disease AI to 59% top-1"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That clinical features encoded as HPO terms plus public curated knowledge tools, scored mainly by a language-model judge, are a fair enough stand-in for real diagnostic performance even when laboratory, biochemical, and genetic-variant evidence are mostly absent and absolute accuracy falls sharply on hard real-world cases.","fun_headline_variants_meta":{"raw":{"variants":["Policy plus tools lifts unmodified LLMs to 59% rare-disease top-1","Human-feedback policy steers LLMs to better rare-disease accuracy","LiteOdyssey: auditable policy yields 59% top-1 rare-disease calls","One reusable policy improves diagnosis on 515 UDN rare-disease cases","No fine-tuning needed: policy layer lifts rare-disease AI to 59% top-1"]},"model":"grok-4.5","effort":"low","cost_usd":0.008246,"raw_usage":{"total_tokens":1899,"prompt_tokens":739,"num_sources_used":0,"completion_tokens":112,"cost_in_usd_ticks":82460000,"prompt_tokens_details":{"text_tokens":739,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1048,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":739,"tokens_out":112,"duration_ms":7993,"temperature":1.0,"reasoning_tokens":1048,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T13:52:44.731121+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A prospective bedside trial that supplies the system with incomplete or evolving phenotypes plus laboratory and genetic data, then measures whether physician-adjudicated top-1 and top-5 accuracy still exceeds the same unmodified model without the policy and tools.","supporting_citations":[],"review_version":1}