{"id":"b9cfe7bf-1323-4751-a1d9-b72193952c3c","arxiv_id":"2607.05401","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"ICLR peer-review scores are orthogonal to future disruptiveness (EDM); EDM identifies highly cited catalysts far better than CD, node2vec, or an LLM rater, and catalyst types precede large topic-share and cross-topic flow growth.","lead":"Across 36,113 ICLR submissions, peer-review scores are essentially uncorrelated with later trajectory-changing impact measured by direction-aware citation embeddings. Topic-bridge and topic-initiator papers precede large subsequent growth in topic share and cross-topic citation flow.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The orthogonality claim rests on EDM as the sole trajectory measure, yet EDM ranks are pipeline-sensitive and the ICLR-internal graph is incomplete.","rationale":"The reader correctly isolates the measurement assumption behind the headline orthogonality result: that the ICLR-internal graph plus EDM adequately capture multi-generational redirection. That is the single most load-bearing point. The paper supplies extensive null tests and propensity-matched growth factors, and the authors themselves document the pipeline sensitivity and coverage gaps, so the concern does not overturn the result; it simply keeps the claim conditional on the chosen operationalization of disruptiveness. A multi-pipeline re-estimation of the exact statistics in §7.1 would settle whether the null survives the documented rank instability. No stronger internal inconsistency is present, and the growth-factor and taxonomy contributions remain intact. Hence the verdict stays CONDITIONAL with no change in direction.","tokens_in":25821,"tokens_out":577,"duration_ms":5367,"concrete_test":"Recompute EDM under three independent walk-generation pipelines (canonical sequential R=80, parallel R=20, and a third seed-averaged ensemble) on the full matched subgraph, then re-estimate the four Spearman correlations and the accepted-vs-rejected Mann–Whitney test of Table 4 / §7.1. If any |ρ| exceeds 0.05 or the p-value for mean EDM difference falls below 0.05 under any pipeline, the orthogonality claim is pipeline-dependent and weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that peer-review scores are essentially orthogonal to future disruptiveness, with |ρ|≤0.005 and indistinguishable mean EDM for accepted vs. rejected papers (p=0.11). That claim is only as strong as EDM itself. Appendix C shows that sequential vs. parallel walk generation at matched hyperparameters yields cross-pipeline Spearman ρ near zero (mean −0.145), so individual paper ranks are not reproducible across ordinary pipeline choices; only distribution-level AUC is stable. In addition the citation graph is ICLR-internal with a 77.1% S2 match rate, unmatched papers concentrated among rejected submissions, and sparse coverage for 2024–2025 cohorts. If the null correlation is an artifact of rank noise or of the missing external edges that rejected papers often acquire, the orthogonality result would not hold under a more complete multi-generational measure. The paper already reports that accepted and rejected papers have nearly identical EDM distributions, but that comparison itself inherits the same measurement fragility.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies 36,113 ICLR submissions (2017–2025) with public OpenReview scores, defining catalyst papers as those whose descendants redirect later research. It compares four disruptiveness measures (CD, node2vec, direction-aware EDM, and an LLM rater), introduces a five-type multi-label catalyst taxonomy (TI, TB, WR, SC, RM), and links catalysts to subsequent topic-share and cross-topic flow growth. EDM best recovers highly cited ICLR papers (ERS AUC 0.83 vs 0.60/0.49/0.42). Topic initiators and bridges precede large multiplicative growth versus year-matched (and propensity-matched) controls. The headline recognition result is that reviewer scores are essentially orthogonal to EDM-based disruptiveness (|ρ|≤0.005; accepted vs rejected mean EDM indistinguishable, p=0.11), with residual miscalibration structured by catalyst type and topic.","tokens_in":26130,"tokens_out":1578,"duration_ms":24245,"significance":"If the results hold, this is a substantial science-of-science contribution for ML/NLP: it is the first corpus-scale pairing of per-paper OpenReview signals with a direction-aware, multi-generational trajectory measure, and the near-null review–disruption relationship is both surprising and programmatically relevant. Strengths include the head-to-head measure comparison on a common venue, year- and propensity-matched mechanism tests, threshold sweeps for TI/TB, explicit documentation of EDM pipeline sensitivity (Appendix C), and a portable operational taxonomy. The work is carefully instrumented (Firth logistic ORs, Mann–Whitney/Welch tests, multi-vendor LLM agreement) and does not overclaim peer-review “failure.” These are real assets for a cs.DL / science-of-science audience.","major_comments":[{"comment":"§7.1 / Table 4 and the abstract’s strongest claim (review–EDM orthogonality, |ρ|≤0.005; accepted vs rejected mean EDM p=0.11) rest on EDM as the sole trajectory outcome. Appendix C reports that sequential vs parallel walk generation at matched hyperparameters yields cross-pipeline Spearman ρ ≈ −0.145 for per-paper ranks, while only distribution-level AUC is stable. Because review-gap analyses (§7.2, Table 14) and top-decile membership use EDM percentiles/ranks, the null result needs an explicit robustness check under the multi-seed median-rank ensemble (or a second pipeline) already described in Appendix C: recompute Spearman ρ, Firth ORs, and accepted/rejected mean-Δ tests on ensemble EDM. Without that, the headline claim is only as strong as a pipeline-sensitive ranking.","section":null},{"comment":"§3 and Appendix A: the citation graph is ICLR-internal with a 77.1% S2 match rate; unmatched papers are concentrated among rejected and recent submissions. The accepted vs rejected EDM comparison (n_acc=8,586; n_rej=13,716 with Δ) and the claim that rejected papers are slightly over-represented in the top EDM decile therefore condition on being matched and having enough structure for EDM. Please quantify how match failure and 2024–2025 sparsity affect the null: e.g., sensitivity restricted to 2017–2022 cohorts with high match rates, or bounds under alternative external-citation graphs for reappearing rejects. This is load-bearing for RQ3’s interpretation that gatekeeping is orthogonal to trajectory change rather than that unmatched rejects are simply unobserved.","section":null},{"comment":"§4.1 / §6 and Tables 8–10: catalyst labels are threshold-defined (TI share growth ≥2.0×; TB top 10% flow and D_i≥2; WR above cluster-median centroid shift; RM Δ≥90th pct plus review boundary). Appendix H shows useful TI/TB sensitivity, but the multi-label co-occurrence claims (e.g., 36.6% of RM also TB; union 22.2%) and the “TB is the largest mechanism” ranking should be reported under the same threshold grid used for growth ratios, not only at the canonical cut. Otherwise the taxonomy risks reading as free-parameter-dependent rather than mechanism-stable.","section":null},{"comment":"§5.1 / Tables 1–2 and Appendix E: LAS has n=50 with only 9 union positives and run-to-run κ=0.291 for the LLM judge. The paper correctly frames LAS as cross-model semantic-rubric agreement, not human gold, but still uses LAS AUC/OR to position M4 against EDM. With 9 positives, the M4 LAS OR 1.41 (p=0.03) and the “complementarity” narrative are under-powered. Either enlarge LAS (or add a small human-annotated subset) or demote LAS-based ranking claims in the main text and keep M4 as an exploratory content-only baseline validated mainly by inter-LLM agreement (Table 3).","section":null}],"minor_comments":[{"comment":"Fig. 1 and §5: state clearly in the main text that EDM covers 62% of papers while CD/node2vec cover 35%/18%; readers can otherwise over-read head-to-head AUCs as same-support comparisons.","section":null},{"comment":"Eq. (3) and Appendix B: define past/future vectors and the single-side skip-gram objective earlier in §4.2 so that Δ_i is self-contained without jumping to the appendix for the training loss.","section":null},{"comment":"§6.2 simultaneous-discovery funnel: the drop from 483,809 to 162 candidates is important; a one-sentence main-text note that raw cosine≥0.9 is dominated by sparse-neighborhood artifacts would prevent over-reading the rarity claim relative to Kim et al. (2026).","section":null},{"comment":"Table 15 / topic bias: clarify whether topic gaps are residualized on year and acceptance; trendy topics (diffusion, ViT) may also differ in citation half-life, which could partially drive EDM percentiles.","section":null},{"comment":"References and related work: Tran et al. (2020) is cited for score–citation correlation; a brief quantitative contrast (their weak correlation vs your near-zero score–EDM ρ) in §2 or §7 would sharpen the novelty claim.","section":null},{"comment":"Ethics / §8: the caution against using EDM as a direct acceptance input is well taken; consider stating explicitly that ensemble ranks, not single-pipeline ranks, would be the minimum if anyone attempted operational use.","section":null}],"recommendation":"major_revision","confidential_remarks":"The contribution is real and well matched to a science-of-science / cs.DL venue; I would not reject on novelty grounds. The skeptic’s concern about EDM pipeline noise is substantive for the abstract’s strongest claim and is the main reason I ask for major rather than minor revision—but it is fixable inside the current scope with analyses the authors already partially ran in Appendix C. No integrity or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: on 36k ICLR papers, OpenReview scores do not track long-run trajectory change as measured by EDM. Accepted and rejected papers have the same mean EDM (p=0.11), and the Spearman with scores is essentially zero. That is the result people will remember, and it is backed by multiple null tests, not a single p-value.\n\nWhat is new is the join of two literatures that usually stay separate. The paper runs CD, node2vec, EDM, and a content-only LLM rater head-to-head on the same public corpus, introduces a five-type operational catalyst taxonomy (TI/TB/WR/SC/RM), and shows that topic initiators and bridges precede large subsequent growth (7.55\times topic share, 11.52\times cross-topic flow) that survives year-matched and propensity-matched controls. EDM’s ERS AUC of 0.83 is clearly better than the alternatives, and the authors correctly treat the LLM rater as a complementary semantic channel rather than a gold standard. The appendices are unusually thorough: hyperparameter sweeps, threshold sensitivity, co-occurrence matrices, and rejected-paper recovery rates.\n\nThe soft spots are real but do not sink the main claim. EDM per-paper ranks are pipeline-sensitive (cross-pipeline Spearman near zero in Appendix C); only distribution-level AUC and means are stable. The citation graph is ICLR-internal with 77% S2 match and sparse recent cohorts, so multi-generational reach is incomplete. LAS is tiny (9 positives, κ=0.291). Thresholds for the catalyst types are hand-chosen. None of this invents the orthogonality result; it just means the result is about EDM distributions, not about a perfectly reproducible ranking of every paper. The stress-test concern is therefore partly right about measurement fragility and partly overstated: the paper already shows the accepted/rejected comparison is distributional, and that is the claim that matters for peer-review design.\n\nThis is for science-of-science people and for anyone who designs ML conference review. It is not a theory paper and not a methods breakthrough, but the data work is careful and the central null is useful. I would send it to referees; the limitations are already on the page and can be tightened without rewriting the contribution. Worth engaging.","headline":"Solid large-scale ICLR study showing review scores are orthogonal to EDM-based trajectory change; the null is real at distribution level even if per-paper ranks are noisy.","tokens_in":26760,"tokens_out":576,"would_cite":true,"duration_ms":5621,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"ICLR peer-review scores do not predict which papers later redirect research trajectories.","keywords":["catalyst papers","disruptiveness measures","EDM","ICLR peer review","topic initiation","topic bridging","science of science","citation embeddings"],"falsifier":"Replicate the same EDM-versus-reviewer-score analysis on a multi-venue corpus (or a denser citation graph that includes arXiv and non-ICLR citers) and find a substantial positive correlation between mean review score and future EDM, or a clear mean-EDM gap between accepted and rejected papers.","tokens_in":26697,"feed_emoji":"🧭","tokens_out":603,"duration_ms":5138,"temperature":0.7,"pith_summary":"A handful of methods have redirected AI research, yet conference review still decides which submissions enter the record. This paper asks whether ICLR reviewer scores, available for every submission from 2017 to 2025, can identify those trajectory-changing papers at decision time. On 36,113 papers it defines catalysts as submissions whose later descendants measurably reorient topics or citation structure, compares four disruptiveness measures, and shows that a direction-aware embedding measure best recovers highly cited work. Topic-initiating and topic-bridging papers precede large subsequent growth in topic share and cross-topic flow. Review scores, however, correlate essentially at zero with future disruptiveness, and accepted and rejected papers look the same on that measure. The residual mismatch is structured by catalyst type and by topic rather than random noise.","feed_headline":"ICLR review scores miss the papers that redirect research","feed_subtitle":"On 36k submissions, scores are orthogonal to future disruptiveness; bridges and initiators drive later growth","key_machinery":"The Embedding Disruptiveness Measure (EDM): past and future vectors learned from direction-aware random walks on the ICLR-internal citation graph; disruptiveness is the cosine distance between those vectors, paired with a five-type operational catalyst taxonomy (topic initiator, topic bridge, within-topic redirector, simultaneous, recognition-misaligned).","core_discovery":"On the nine-year ICLR record, peer-review signals are essentially orthogonal to long-run trajectory change measured by direction-aware citation embeddings: Spearman correlations with EDM stay within |ρ| ≤ 0.005, accepted and rejected papers have indistinguishable mean EDM, and residual miscalibration concentrates on topic bridges and recognition-misaligned papers rather than on within-topic work.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["ICLR peer scores orthogonal to papers that redirect research trajectories","Review scores fail to flag catalyst papers across 36k ICLR submissions","Accepted and rejected ICLR papers show equal future disruptiveness","Topic bridges and initiators drive growth; review scores stay blind","EDM spots highly cited catalysts while peer scores do not"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That an incomplete ICLR-only citation graph and a pipeline-sensitive embedding score adequately capture multi-generational research redirection for both accepted and rejected papers.","fun_headline_variants_meta":{"raw":{"variants":["ICLR peer scores orthogonal to papers that redirect research trajectories","Review scores fail to flag catalyst papers across 36k ICLR submissions","Accepted and rejected ICLR papers show equal future disruptiveness","Topic bridges and initiators drive growth; review scores stay blind","EDM spots highly cited catalysts while peer scores do not"]},"model":"grok-4.5","effort":"low","cost_usd":0.002614,"raw_usage":{"total_tokens":1073,"prompt_tokens":860,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":26140000,"prompt_tokens_details":{"text_tokens":860,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":126,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":860,"tokens_out":87,"duration_ms":2054,"temperature":1.0,"reasoning_tokens":126,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T16:03:37.881376+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Replicate the same EDM-versus-reviewer-score analysis on a multi-venue corpus (or a denser citation graph that includes arXiv and non-ICLR citers) and find a substantial positive correlation between mean review score and future EDM, or a clear mean-EDM gap between accepted and rejected papers.","supporting_citations":[],"review_version":1}