{"id":"7c7e121a-02ca-4c49-8dbd-e056a7fe9296","arxiv_id":"2607.24165","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"On conjunctive cross-page retrieval, strong hybrids find all-condition golds for ~81% of queries but rank them before natural subset matches only ~36% of the time.","lead":"Current document retrievers often find a long document that matches a multi-part request without ranking one that covers every requested condition first. The study shows a large discovery–completion gap on a controlled cross-page benchmark and points to condition coverage, not model size, as the bottleneck.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline 81.1/35.8 gap rests on subset labels that certify absence only via inventory silence; the paper's own retrieval-conditioned challenge on SourceShift promoted 72% of top-ranked \"subsets\" to gold, and no equivalent audit exists for CrossPage's blocking subsets.","rationale":"The reader's weakest_assumption names inventory false negatives and the versioned-inventory scoring rule; I identify the same locus and sharpen it with the paper's own SourceShift challenge data (90/125 top-ranked subsets promoted to gold) plus the explicit absence-asymmetry policy in §10.2.2. This is the load-bearing point because the metric's numerator event — gold before *every* released subset — is maximally sensitive to false subset labels, the error direction is one-sided (promotions can only raise n-Clue Score, shrinking the gap), and the one population never audited on CrossPage (high-ranked blocking subsets) is exactly the population that determines the headline number.\n\nI nonetheless recommend UNCHANGED (CONDITIONAL) rather than a downgrade, for three reasons. First, the paper is unusually transparent about this exposure: it discloses the grade-0 rule, the asymmetric audit policy, the adversarial k-promotion bounds in Table S1h, and limits its claim to \"the fixed, versioned instrument\" rather than real-traffic prevalence (§7). Second, the operation-level conclusions — dense decomposition helps (+6.8/+7.3, corrected p=.005/.002), fusion helps, generic reranking hurts, within-family scale is null — were re-evaluated on SourceShift's post-challenge qrels and mostly held (Table 7), so label repair did not overturn the directional findings there. Third, the concern is testable with the authors' own released protocol and code, which is exactly what a CONDITIONAL verdict should attach to: acceptance should be conditioned on running the retrieval-conditioned subset audit on CrossPage. If that audit shows a low promotion rate, the paper's central claim stands essentially intact; if it mirrors SourceShift, the magnitude of the discovery–completion gap — though likely not its existence — needs revision.","tokens_in":29903,"tokens_out":3223,"duration_ms":115803,"concrete_test":"Replicate the SourceShift targeted challenge on CrossPage: collect every released subset appearing in the top-10 union of the five Table 4 systems (hybrid, ColQwen top-3, Qwen 0.6B dec., Qwen 8B, BM25-AND), run the same expanded-page adjudication (6–9 page high-recall pools, exact-quote validation, distinct-page gate), then recompute Table 3/Table 4 under promoted qrels. If the blocking-subset promotion rate is under ~5% and the 81.1/35.8 gap moves by <2 points, the concern does not land; if promotion is tens of percent as in SourceShift, the subset-first failure mass and the headline gap shrink materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — Gold Hit 81.1% vs n-Clue Score 35.8%, with \"subset-first\" the largest failure mass at 45.3% (Table 4) — requires the released subset qrels to be truly incomplete. But subset labels are built on absence: a document is grade g<n because the versioned page-fact inventory contains no page for the missing condition. The paper's own audit policy states this asymmetry explicitly (§10.2.2): \"failure to find a quote in three pages cannot establish document-wide absence.\" Construction reads only \"the pages where mined surface forms occur\" (§11), so any condition expressed without the mined section/entity surface form, or lost to OCR error across a median-67-page document, becomes a silent false negative — and each such false negative manufactures a blocking subset that directly deflates n-Clue Score while leaving Gold Hit untouched. This asymmetry means the measurement error runs in exactly the direction that inflates the headline gap.\n\nThe paper contains direct evidence that this risk is material. The SourceShift targeted challenge (§10.5) audited the 125 already-judged subsets sitting in the frozen fusion/BM25 top-10 union: 101/125 grades changed and 90 became gold — a 72% promotion rate among high-ranked subsets. Selection was contrast-conditioned, but selection only chose which pairs *mattered*; it did not cause the promotions, which required verbatim quotes and distinct-page gates. That is the best available empirical estimate of inventory false-negative rates in the population that determines the metric, and it is very high. For CrossPage, the only audit was a stratified random 100-pair sample (98 WB, 2 PMC), whose 20 repairs moved scores ≤0.2 points — but a random sample never probes the specific subsets that actually block the headline systems' top-10. Table S1h sharpens the worry: promoting just one judged top-10 subset per query (k=1) flips every positive controlled gap (e.g., complementarity +8.7 → −24.6). CrossPage's page-grounded pipeline is p","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript studies conjunctive cross-page document retrieval: queries with two or three explicit conditions that must be supported on different pages of a single long document. It introduces CrossPage (1,000 queries over 2,021 real documents with a page-grounded condition inventory) and the n-Clue Score@10, which requires a top-10 all-condition gold to precede every released subset qrel. Across 70 configurations, the paper reports a large discovery–completion gap (strongest hybrid: 81.1% Gold Hit@10 vs. 35.8% n-Clue Score@10), gains from condition-wise decomposition (+6.8–7.3 points on two dense backbones) and lexical–visual RRF fusion (+8.7), consistent harm from four generic rerankers, and a null effect of scaling Qwen3-Embedding 0.6B→8B. Robustness is probed via deterministic query renderings, LLM paraphrases, task-factor slices, a four-source SourceShift stress set, blind quote-backed audits with repair, adversarial promotion bounds, and multiplicity-corrected paired inference; code, qrels, rankings, and audit records are released with one-command reproduction. The intended claim is deliberately scoped to explicit, conjunctive, cross-page requests.","tokens_in":30307,"tokens_out":11607,"duration_ms":351944,"significance":"If the results hold, the paper makes a useful diagnostic contribution: it separates gold discovery from complete-before-subset ordering and shows that current retrievers largely fail the latter, with condition coverage — not recall or model size — as the bottleneck. Strengths that raise confidence: a controlled instrument with distinct-page gates and privileged-oracle/floor validity checks; operation-level paired contrasts rather than a leaderboard; blockwise max-|T| plus Holm correction; family bootstrap and crossed dependence analyses; quote-backed blind audits with released records; adversarial promotion bounds; and standard-library one-command reproduction over frozen rankings. These make the claims unusually checkable and give the community a falsifiable instrument plus concrete design targets (coverage tracking, weakest-condition verification).","major_comments":[{"comment":"The headline 81.1/35.8 gap and Table 4's 45.3% 'subset-first' mass depend on subset labels that certify absence by inventory silence (§7; §11: construction reads only fact-matched pages; §10.2.2 states the asymmetry explicitly). The paper's own SourceShift targeted challenge (§10.5) promoted 90/125 (72%) of high-ranked judged subsets to gold, and the CrossPage stratified audit revised 20/100 sampled pairs (19 promotions). No equivalent retrieval-conditioned audit exists for CrossPage's blocking subsets. Because subset-to-gold false negatives deflate n-Clue Score while leaving Gold Hit untouched, label error inflates the measured gap by an unquantified amount. Please run the §10.5-style challenge on (a sample of) subsets in the clean top-10 union of the 13 scorecard systems and report corrected score ranges.","section":"§10.5 vs. Tables 3–4"},{"comment":"Table S1h shows a single adversarial judged-subset promotion per query flips every positive contrast (e.g., complementarity +8.7 has bound −24.6 at k=1). As a worst-case bound this is expected, but combined with the empirical promotion rates above it leaves the signs of the Table 5 controlled contrasts unquantified under plausible label noise. Since the contrasts are paired on shared qrels, random promotions need not flip signs; please add a plausible-rate sensitivity analysis — recompute contrasts under subset-to-gold promotions sampled at audit-estimated rates, or report per-contrast break-even promotion rates. This is executable with the released machinery.","section":"Table S1h / Table 5"},{"comment":"The abstract and §4.3 state that operation directions 'replicate on a four-source stress set.' The n-Clue directions do replicate on final qrels (Table 7), but the frozen SourceShift fusion−BM25 gold-NDCG contrast that motivated the stress set collapsed from +.110 to +.016 (p=.381) after the targeted challenge, with both crossed intervals including zero (Tables S14, S1i). The main text should state this outcome: it bears directly on how much weight the stress-set replication can carry, and on the label-noise question in the first comment.","section":"§4.3 / Abstract vs. Tables S14, S1i"}],"minor_comments":[{"comment":"The indicator symbol renders as ⊮ in the PDF; check the glyph or define the notation explicitly.","section":"Eq. (1)"},{"comment":"Gold+Support is reported only for Qwen3-VL 2B/8B. Clarify whether ColQwen2.5 page scores were unusable for this metric, and soften the abstract's 'page-aware visual systems' wording, which generalizes from one model family.","section":"§4.4 / Table 8"},{"comment":"PMC contributes 1 of 2,176 gold qrels; the two-source framing slightly overstates diversity (§7's 'World-Bank-centered' is the accurate emphasis). The 'Single-page-solvable golds 0 / 2,176' row formatting is also ambiguous on first read.","section":"Table 1"},{"comment":"The dagger notes post-hoc selection of the hybrid on the raw matrix. It would defuse this concern to state that the same system is also the top raw-form fair row (Table S1a, 31.1), so the selection is among near-equivalent variants rather than cherry-picked.","section":"Table 3, footnote"},{"comment":"Gold+Support credits only the three highest-scoring surfaced pages, which is exactly tight for n=3 queries; a sweep over the number of surfaced pages would clarify how much of the 5.1–5.3% is metric strictness versus genuine delivery failure.","section":"§3.1"},{"comment":"Some cells are small (n=146, 185). The text notes the slices are descriptive, but adding interval estimates or explicit cell-size cautions would help readers avoid over-reading single cells.","section":"Table 6"}],"recommendation":"major_revision","confidential_remarks":"The post-hoc hybrid selection is disclosed and largely defused by the fact that the same system is also the top raw-form fair row (Table S1a); I would not weigh it heavily. The substantive open question is the unquantified false-negative rate among CrossPage blocking subsets, where the authors' own SourceShift challenge (72% promotion among high-ranked subsets) is the best available estimate and points in the direction that inflates the headline gap. The authors appear well positioned to resolve this with their existing audit apparatus at moderate cost, which is why I recommend major rather than minor revision. The release package (frozen rankings, audit records, one-command verification) is genuinely strong and should be encouraged."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is not another multi-condition leaderboard. They built a controlled instrument (CrossPage/n-Clue) where golds must cover two or three explicit conditions on distinct pages, competitors are natural corpus subsets, and the primary event is strict: a top-10 gold before every released subset. That estimand, plus the outcome decomposition (miss vs subset-first vs complete-first), is the real contribution.\n\nWhat they do well is the experimental hygiene. Seventy configs, same-backbone decomposition contrasts, fixed-pool rerankers, form ablations, SourceShift direction checks, family bootstrap, blockwise max-|T| plus Holm, quote-backed audit with repair sensitivity ≤0.2 points, and a full release of rankings and a one-command repro path. Dense condition-wise decomposition helps (~7 pts), lexical–visual RRF helps (~9), generic rerankers hurt Gold-NDCG, and Qwen 0.6B→8B is a flat 0.0 on complete-first. Those operation signs are measured, not vibes. The 81.1 Hit / 35.8 Score hybrid gap and the ~5% Gold+Support numbers make the coverage bottleneck concrete.\n\nThe soft spot that actually matters is label asymmetry on absence. Subsets are incomplete because the inventory is silent; their own policy admits three-page non-finds do not prove document-wide absence. On SourceShift, the retrieval-conditioned challenge promoted 90/125 high-ranked subsets to gold. CrossPage only got a stratified random sample, not an equivalent top-rank blocker audit, and Table S1h shows k=1 judged promotions can wipe the positive gaps. That does not invent the discovery–completion story, but it means the headline gap size is partly an inventory artifact until someone audits the actual blocking subsets the way they audited SourceShift. Post-hoc hybrid display and World-Bank-long-doc concentration are secondary and already scoped by the authors.\n\nMath and stats are standard and transparent; citations sit in the right neighborhood (MultiConIR, ComLQ, LongEmbed, MMDocIR, etc.) without pretending the joint controls already existed. This is for people building or evaluating multi-evidence RAG/long-doc retrievers who care about coverage state, not topical hit rate. I would send it to peer review, bring it to reading group, and cite the framing and the operation contrasts—with the inventory caveat attached.","headline":"Careful diagnostic that cleanly separates gold discovery from complete-before-subset ranking; the gap is real on their instrument, but inventory false-negatives on blocking subsets are the load-bearing soft spot.","tokens_in":31805,"tokens_out":612,"would_cite":true,"duration_ms":16844,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Finding a relevant long document is not the same as proving it covers every requested piece of evidence.","keywords":["conjunctive retrieval","cross-page evidence","discovery–completion gap","condition coverage","dense retrieval","lexical–visual fusion","reranking","long-document IR"],"falsifier":"A retriever that, on the same CrossPage queries and qrels, lifts complete-first success near its Gold Hit rate—especially closing the hybrid’s 81% hit versus 36% complete-first gap—or that surfaces stored support for every condition on far more than about 5% of queries.","tokens_in":31483,"feed_emoji":"📄","tokens_out":922,"duration_ms":17961,"temperature":0.7,"pith_summary":"Many real searches ask for a conjunction of evidence spread across different pages of one long document. A system can retrieve a topically relevant document and still fail if a natural partial match—supporting only some of the conditions—ranks higher. This paper builds a controlled test, n-Clue on CrossPage, where each query has all-condition golds and real subset competitors, and success means a top-10 gold beats every released subset. Across seventy configurations, condition-wise decomposition and lexical–visual fusion help, generic rerankers hurt gold placement, and scaling one dense family from 0.6B to 8B does not move complete-first success. The strongest hybrid finds a gold on 81.1% of queries but succeeds complete-first on only 35.8%, and page-aware systems surface stored support for every condition on roughly 5% of queries. The bottleneck is condition coverage and complete-before-subset ordering, not gold discovery alone.","feed_headline":"Retrievers find full evidence docs but rank partial matches first","feed_subtitle":"Best hybrid hits gold 81% of the time yet completes the request first only 36%","key_machinery":"n-Clue Score@k: a controlled complete-first event that succeeds only when a top-k all-condition gold precedes every released subset qrel, paired with Gold Hit and page-support diagnostics so discovery, ordering, and evidence delivery can be separated.","core_discovery":"On explicit conjunctive cross-page requests, representative retrievers often discover all-condition golds without ranking them ahead of natural subset matches. Condition coverage—not gold discovery or within-family scale—is the central bottleneck: the strongest displayed hybrid reaches 81.1% Gold Hit@10 but only 35.8% complete-first success, and page-aware systems deliver all stored support on only 5.1–5.3% of queries.","pith_inferences":["RAG pipelines that assume “hit a relevant doc” equals “evidence is available” will systematically under-serve multi-part legal, policy, and scientific requests.","Leaderboards that reward graded partial relevance may rank systems that flood the top with incomplete matches above systems that protect complete-first ordering.","A practical next system is a condition–page matrix plus a small verifier that re-queries only the missing condition rather than another end-to-end scalar reranker.","The same discovery–completion split likely appears in any corpus where partial matches are common and evidence is page-scattered, not only in this instrument."],"forward_implications":["Retrieval stacks for multi-part requests need an explicit coverage layer, not only a single relevance score.","Condition-wise candidate lists and cross-modal fusion are higher-leverage than scaling one dense encoder alone.","Generic rerankers trained for ordinary relevance can worsen complete-before-subset ordering.","Training should treat natural grade-(n−1) documents as hard negatives that must lose to full-coverage golds.","Page-level systems must match conditions to distinct pages and verify the weakest condition, not only max-pool pages."],"fun_headline_variants":["Retrievers hit gold 81% yet rank full evidence first only 36%","Condition coverage—not gold discovery—bottlenecks cross-page retrieval","Full-evidence docs lose to natural partial matches in top ranks","Dense scale from 0.6B to 8B adds 0 points on complete-first success","Page-aware systems surface all stored support on only 5% of queries"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That ranking a gold before known inventory-labeled subsets on this fixed, mostly World-Bank long-document test is a fair stand-in for real multi-evidence readiness.","fun_headline_variants_meta":{"raw":{"variants":["Retrievers hit gold 81% yet rank full evidence first only 36%","Condition coverage—not gold discovery—bottlenecks cross-page retrieval","Full-evidence docs lose to natural partial matches in top ranks","Dense scale from 0.6B to 8B adds 0 points on complete-first success","Page-aware systems surface all stored support on only 5% of queries"]},"model":"grok-4.5","effort":"low","cost_usd":0.00414,"raw_usage":{"total_tokens":1291,"prompt_tokens":847,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":41404000,"prompt_tokens_details":{"text_tokens":847,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":358,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":847,"tokens_out":86,"duration_ms":6902,"temperature":1.0,"reasoning_tokens":358,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T22:09:27.119363+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A retriever that, on the same CrossPage queries and qrels, lifts complete-first success near its Gold Hit rate—especially closing the hybrid’s 81% hit versus 36% complete-first gap—or that surfaces stored support for every condition on far more than about 5% of queries.","supporting_citations":[],"review_version":1}