{"id":"d79f8c41-3034-42cf-acda-6efff4307ab3","arxiv_id":"2608.01704","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Machines recover up to 53% of the crowd-highlight prediction headroom, and fusing five frontier models reaches about 60%, confirmed in a pre-registered replication.","lead":"This paper measures the lower and upper bounds for predicting which sentences a crowd of readers highlights, then places today's models and a multi-model fusion between those bounds. It shows the task is roughly half-solved, that fusing five different models is the cheapest known improvement, and that the gain survives an independent pre-registered replication.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The split-half ceiling in §4.1 rests on an explicitly unverified reader-independence assumption (§2); if readers influence each other or are non-exchangeable over time, the +0.2028 headroom and every '% of headroom' share shift, leaving the 'roughly half-solved' bracket unprotected.","rationale":"The reader's weakest assumption is exactly the one I identify: reader independence/exchangeability is structural, explicitly inherited rather than re-verified, and sits underneath every '% of headroom' claim. I considered the distillation/context confound in §4.7, but that is secondary to the main bracket claim, and the fusion advantage itself is protected by ablation and a pre-registered replication. The reader-independence assumption is the most load-bearing because the split-half ceiling is the denominator for the paper's headline 'roughly half-solved' assessment; a small bias in the ceiling changes the central conclusion even though the fusion-vs-best-single contrast may survive unchanged. The proposed temporal-split test is feasible from existing per-reader data and would directly probe whether the random-split ceiling is an artifact of non-independence. Since the reader already gave a CONDITIONAL verdict and my concern reinforces that conditionality, the verdict should remain CONDITIONAL—no adjustment is needed.","tokens_in":6655,"tokens_out":15427,"duration_ms":214661,"concrete_test":"Using the per-reader mark files, recompute the split-half ceiling with a temporal split: within each document order readers by first-highlight timestamp, split at the median into early and late halves, and score early-half counts against the late-half top-15% label under the same domain-clustered bootstrap. Compare this estimate to the random-split ceiling of §4.1. If the difference exceeds the bootstrap SE (95% CI on the difference excludes zero), the independence/exchangeability assumption fails and the reported headroom is not a stable label-reliability ceiling; if the estimates coincide within uncertainty, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central bracket is floor (lead AP 0.2410) and split-half oracle ceiling (0.4437). Every headline percentage—'no single model reaches 55%', 'fusion reaches 60%'—is a share of this ceiling, so the ceiling is the denominator of the paper's main quantitative claim. Section 2 states: 'Reader independence is an assumption inherited from the platform's design... not re-verified here.' The ceiling is meaningful as a label-reliability ceiling only if the two random halves are independent, exchangeable samples of the same latent salience. If readers affect one another (overlay enabled at least occasionally) or arrive in temporally correlated cohorts whose attention is shared, the split-half AP mixes reliability with social/temporal contagion. The sensitivity analyses in §4.3 perturb label percentile, reader gate, seed, and prompt, but not this structural assumption. A biased ceiling changes the headroom, 'percent solved', and 'unsolved half is semantic' conclusions. The fusion-vs-best-single comparison is independently supported, but the paper's central bracket is not robust until this assumption is tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines a benchmark bracket for predicting which sentences a crowd of readers highlights in web documents. The floor is naive position-based truncation (lead AP 0.2410); the ceiling is a split-half oracle in which one random half of the readers predicts the other half (AP 0.4437). Against this +0.2028 headroom, the paper reports that frontier language models recover 35–53% of the gap, while an unweighted Borda fusion of five models plus a position prior reaches 60%, significantly outperforming the best single model. A pre-registered replication on 217 new documents confirms the fusion gain. Distillation into an 8B open-weight student retains 90% of the fusion's edge and reaches statistical parity with the strongest frontier model. The paper concludes that the task is roughly half-solved, that the unsolved half is semantic rather than positional, and that cross-vendor fusion is the cheapest known improvement.","tokens_in":6993,"tokens_out":6576,"duration_ms":85349,"significance":"If the results hold, this is a valuable contribution to the measurement of machine prediction of human reading attention. The paper introduces a clear floor/ceiling framing that makes benchmark scores interpretable, and it backs the central fusion claim with a pre-registered replication, domain-clustered inference, multiple sensitivity analyses, and a transparent audit trail including verification scripts and a hostile audit record. The distillation result, showing that document-level context rather than local features carries the crowd signal, is also informative. The strengths are substantial: the paper reports reproducible aggregates, pre-registration commit timestamps, and explicit handling of granularity mismatches. The main weakness is that the ceiling, which is the denominator for all headline percentages, rests on an unverified reader-independence assumption explicitly acknowledged in §2.","major_comments":[{"comment":"The split-half ceiling is the denominator of the paper's central quantitative claim. Section 2 states: 'Reader independence is an assumption inherited from the platform's design, stated in that study's terms, not re-verified here.' If readers influence each other (the overlay is 'rarely enabled' but not never) or arrive in temporally correlated cohorts, the two halves are not independent, exchangeable samples of the same latent salience. Then the ceiling AP (0.4437), the headroom (+0.2028), and every '% of headroom' share (35–53%, 60%, 5%) are all biased. The sensitivity analyses in §4.3 perturb label percentile, reader gate, seed, and prompt, but not this structural assumption. This is load-bearing for the 'roughly half-solved' conclusion and for the interpretation that the unsolved half is semantic. I recommend either testing the assumption (e.g., split readers by time of arrival or co","section":"§2, §4.1"},{"comment":"The model-ranking cache mapping is recovered, not read. The paper provides validation: 17/17 exact-match documents agree, the minimum best-vs-second margin is 0.202, and a shuffle null gives at most 2 spurious matches in 200 permutations. This is reassuring, but the remaining 103 documents rely on consensus matching. A mismapped document would attach random rankings to real labels, attenuating all model/fusion APs. The argument that mismapping cannot manufacture the fusion-vs-single advantage is sound (it adds common noise), but the absolute AP values and '% of headroom' figures could be affected if mismapping is non-uniform across documents or models. Since the model arms are central to the main table, I ask that per-document mapping confidence be reported (or a sensitivity analysis that excludes low-confidence mappings) to confirm that the headline numbers are not attenuated by mapping","section":"§3, §5"}],"minor_comments":[{"comment":"The 'position prior' is used throughout but never explicitly defined. Please state its formula (e.g., score = position normalized by document length) and how it is combined with Borda scores.","section":"§3"},{"comment":"The statement that using 40 instead of 60 splits makes the comparison 'unaffected' needs a one-sentence justification: averaging within document first makes the point estimate stable, but the variance of the per-document AP estimates may differ. Please clarify.","section":"§4.3"},{"comment":"The H3 row reads 'surface features recover<20% of headroom−1% — pass.' This is confusing; presumably it means the recovered share was −1%, which passes the pre-registered <20% bound. Please reword.","section":"§4.6"},{"comment":"The text says 'validated three ways (§5.3),' but §5.3 does not exist; the validation appears in the Limitations bullet on the recovered mapping. Fix the cross-reference.","section":"§3"},{"comment":"The table row 'floor—lead0.2410 0 0' lacks spacing and is hard to read. Also, the column header 'AP vs floor' is ambiguous; clarify that the second column is absolute AP and the third is the difference from the floor.","section":"§4.1 table"},{"comment":"The 'pre-registered kill condition (split-half ceiling ≥ 2× the paired MDD) passed at 7.3×' is not defined. Please define the minimal detectable difference and how 7.3× is computed.","section":"§4.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful in its fusion and replication analyses, and the audit culture is exemplary. The main concern is the unverified reader-independence assumption, which is structural to the ceiling. This is fixable either by re-analysis or by careful re-framing, so I support major revision rather than rejection. The paper also relies heavily on companion studies by the same authors; the editor may wish to ensure that those are available for review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for two things: it builds a real floor and ceiling for crowd-highlight prediction (lead-truncation floor, split-half oracle ceiling), and it shows an unweighted fusion of five frontier models plus a position prior recovering about 60% of that headroom, with a pre-registered replication confirming the fusion gain. The practical claim — the task is roughly half-solved, and cross-vendor fusion is the cheapest way to climb — is probably sound, but you should read the ceiling with a grain of salt.\n\nThe empirical work is unusually careful. The domain-clustered inference, the four attacks (arm selection, ablation, sensitivity, prompt dependence), the explicit granularity rule, and the pre-registered replication on 217 new documents all make the fusion-vs-best-single comparison credible. The paper is also transparent to a fault: it reports the retracted claim, the multiplicity correction that killed an earlier result, and the recovered hash mapping with its three-way validation. That honesty earns real credit.\n\nThe soft spot is exactly where the stress-test lands. The split-half ceiling assumes readers mark independently and are exchangeable across halves. The paper states this is inherited from platform design, not re-verified. If readers influence each other or arrive in temporally correlated cohorts, the ceiling estimate — and therefore every \"% of headroom\" number — shifts. This does not threaten the fusion-vs-best-single contrast, which is a paired comparison against the same label, but it does weaken the interpretation that the unsolved half is semantic and the bracket's absolute meaning. The paper should either test this (e.g., temporal split or overlay-use analysis) or clearly label the ceiling as conditional on an unverified assumption.\n\nTwo smaller notes. The recovered model-ranking mapping is a genuine limitation, though the validation is convincing and mismapping would hurt all model arms equally rather than inflate fusion. The document-level structure conclusion from the distillation section is suggestive but rests on one architecture pair; I'd treat it as a hypothesis, not a demonstrated fact.\n\nWho should read this: anyone building or evaluating highlight AI, and anyone working on human-model agreement. It deserves a serious referee. My recommendation: send it to peer review, and ask for a sensitivity analysis on the independence assumption — or at least a sharp caveat in the ceiling definition. The main fusion result is likely to hold regardless.","headline":"This paper gives crowd-highlight prediction its first measured floor/ceiling bracket and a plausible fusion recipe; the central numbers are careful, but the ceiling rests on an untested independence assumption.","tokens_in":7419,"tokens_out":1563,"would_cite":true,"duration_ms":23639,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Crowd highlight prediction is roughly half-solved, and cross-vendor fusion is the cheapest known way to climb the gap.","keywords":["crowd highlighting","reading attention prediction","average precision","floor-ceiling benchmark","model fusion","knowledge distillation","language models","salience"],"falsifier":"Compare the split-half oracle under random splits versus splits stratified by reading order or by whether the highlight overlay was enabled. If early or overlay-exposed readers' marks predict later readers' marks better than random splits would imply, the independence assumption fails and the ceiling is inflated; if the two ceilings are equal, the assumption holds.","tokens_in":6587,"feed_emoji":"📊","tokens_out":6178,"duration_ms":65239,"temperature":0.7,"pith_summary":"This paper argues that a single benchmark number for predicting what a crowd of readers highlights is uninterpretable, and supplies the two numbers that make it interpretable: a floor (naive lead/position truncation, AP 0.2410) and a ceiling (a split-half oracle in which one half of the crowd predicts the other, AP 0.4437). Against that bracket, no single language model reaches 55% of the headroom, while an unweighted fusion of five frontier model rankings plus a position prior reaches 60% and beats the best single model by +0.0159 AP, a result confirmed on an independent 217-document corpus. The paper concludes the crowd-prediction task is roughly half-solved, that the unsolved half is semantic rather than positional, and that the cheapest known improvement is to average several different models rather than rely on one.","feed_headline":"Five-model fusion reaches 60% of the possible crowd-highlight score","feed_subtitle":"A floor/ceiling bracket shows the task is half-solved and averaging model rankings beats any single one.","key_machinery":"The load-bearing device is the bracket: floor = naive lead truncation score; ceiling = split-half oracle (half the crowd's per-sentence counts predicting the other half's top-15% label), both scored by average precision on the same label. The second device is unweighted cross-vendor fusion: five model rankings converted to normalised Borda scores and summed, optionally with a position prior, which is the mechanism that climbs from roughly 35–53% to 60% of headroom. The paper also uses distillation into an 8B whole-document student as the compression test that localises the signal to document-level structure.","core_discovery":"The central claim is that the crowd-highlight prediction task has a measurable, useful bracket, and today's best text-only predictors sit at about half its height. The floor is the lead heuristic: score sentences by position from the top, AP 0.2410. The ceiling is the split-half oracle: a random half of readers predicts the other half, AP 0.4437, gap +0.2028. Frontier language models reach 35–53% of that gap zero-shot; surface features recover only 5%. An unweighted Borda fusion of five vendor-diverse model rankings plus a position prior reaches 60% of the gap and beats the best single model by +0.0159, with a pre-registered replication on 217 fresh documents confirming a similar +0.0179 gai","pith_inferences":["Because the unsolved half is semantic, the paper's logic implies that further gains will come from models that read the whole document jointly, not from better local features; a targeted test would compare whole-document versus sliding-window variants on the same architecture.","The floor/ceiling bracket is transferable to other noisy human-label tasks such as relevance judgments or news salience, where most benchmarks still lack a measured label-reliability ceiling; applying the same split-half oracle there would calibrate how much of the headroom current systems actually fill.","The replication used a separate corpus, so the fusion gain appears robust, but the paper does not test whether the gain survives when the fused models all come from the same vendor family; that contrast would sharpen the claim that vendor diversity is the active ingredient.","The 66% per-document win rate and the ablation results suggest that a simple per-document selection rule, choosing between the fusion and the single best model on each document, could add a small further gain; the paper does not test this."],"forward_implications":["Any reported highlight-prediction score should be read relative to a measured floor and ceiling; a raw AP number by itself is uninformative.","Because no single model reaches 55% of headroom, the task is not saturated; meaningful headroom remains for text-based predictors.","Averaging rankings from several independent model vendors is a cheap, verified way to improve, and the gain arises because models' small residual disagreements are signal, not noise.","The fusion advantage is expected to shrink as models become more alike, so the benchmark should be remeasured against each new generation of models.","A single open-weight 8B student can serve most of the fusion's advantage, making document-level highlight prediction practical at low cost."],"supporting_citations":[{"why":"Supplies the 120-document corpus and establishes that models agree with each other more than with readers.","marker":"[1]"},{"why":"Establishes the individual-level ceiling (no model beats a second reader), the contrast for the crowd-level result.","marker":"[2]"},{"why":"Supplies the 217-document replication corpus with per-reader marks and freshly collected model rankings.","marker":"[3]"},{"why":"Prior art showing heterogeneous model families ensemble better than same-family, which this paper extends by measuring the gain against a human ceiling.","marker":"[4]"},{"why":"Explains why same-model polling fails because errors correlate, making the cross-vendor fusion gain informative.","marker":"[5]"},{"why":"The prompt-compression system used as the below-floor comparison arm.","marker":"[8]"},{"why":"The distillation method used to compress the fusion into a student model.","marker":"[9]"},{"why":"The parameter-efficient finetuning method used to train the 8B student.","marker":"[10]"},{"why":"The open-weight 8B model used as the whole-document student.","marker":"[11]"},{"why":"The local-context encoder used as the 150M student that retains only 63% of the edge.","marker":"[12]"}],"fun_headline_variants":["Fusion of five models reaches 60% of crowd-highlight ceiling","Crowd-highlight prediction: model fusion beats best single by +0.016 AP","Distilled 8B student keeps 90% of model-fusion edge on highlights","Five-model average captures 60% of possible crowd-highlight gain","Pre-registered: fusion beats best model on crowd-highlight prediction"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The ceiling assumes readers marked sentences independently and are exchangeable across random splits; if readers influenced each other or the sample of readers is not representative, the measured headroom and the conclusion that the unsolved half is semantic would be biased.","fun_headline_variants_meta":{"raw":{"variants":["Fusion of five models reaches 60% of crowd-highlight ceiling","Crowd-highlight prediction: model fusion beats best single by +0.016 AP","Distilled 8B student keeps 90% of model-fusion edge on highlights","Five-model average captures 60% of possible crowd-highlight gain","Pre-registered: fusion beats best model on crowd-highlight prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001075,"raw_usage":{"total_tokens":4434,"prompt_tokens":937,"completion_tokens":3497,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":681,"completion_tokens_details":{"reasoning_tokens":3396}},"tokens_in":681,"tokens_out":3497,"duration_ms":27363,"temperature":1.0,"reasoning_tokens":3396,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:30:50.341411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the split-half oracle under random splits versus splits stratified by reading order or by whether the highlight overlay was enabled. If early or overlay-exposed readers' marks predict later readers' marks better than random splits would imply, the independence assumption fails and the ceiling is inflated; if the two ceilings are equal, the assumption holds.","supporting_citations":[{"cited_title":"Language Models Agree With Each Other, Not With Readers","cited_arxiv_id":"2607.29274","evidence_quote":"Supplies the 120-document corpus and establishes that models agree with each other more than with readers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior art showing heterogeneous model families ensemble better than same-family, which this paper extends by measuring the gain against a human ceiling."}],"review_version":1}