{"id":"2f98c8f9-1ffd-4ba4-8e65-d9f56d35990c","arxiv_id":"2607.27739","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Language-model importance rankings retain 38.4% of crowd-highlighted sentences versus 19.9% of matched non-highlighted sentences, an enrichment of +0.196 net of position and length confounds.","lead":"A study measured whether AI language models select the same sentences that human readers actually highlight on web pages, after removing the advantages of sentence position and length. It found that an off-the-shelf LLM matches crowd highlights about as well as a single human reader does.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pretraining contamination is not ruled out by the mark-recency test; the +0.196 enrichment may reflect retrieval of crowd highlights rather than text-based alignment.","rationale":"The reader's CONDITIONAL verdict is driven mainly by label tie-break reproducibility. That concern is real but secondary: even under alternative tie-breaks the effect remains positive and significant, so it shifts the headline magnitude, not the existence of the effect. The pretraining contamination concern is more load-bearing because it threatens the construct itself: if the model has memorized which sentences readers marked, the +0.196 is not 'alignment with reader highlights' but retrieval of a publicly rendered label. The paper's mark-recency split cannot rule this out because all documents may have entered pretraining long before their marks accumulated. A truly held-out subset—documents whose text and marks postdate the model's training cutoff—would settle the question. The paper is unusually transparent and its internal controls are careful, which is why I do not recommend moving to REJECT; the right verdict remains CONDITIONAL, with the contamination test as a condition for interpreting the headline as alignment rather than retrieval. My agreement_with_reader is 'partial' because the reader's weakest-assumption pick is not the same as mine, though both concern the label's construction and reproducibility.","tokens_in":12211,"tokens_out":11359,"duration_ms":132401,"concrete_test":"Restrict the evaluation to documents whose first marking event (or first publication) postdates each model's reported training-data cutoff and whose text was not publicly indexed before that cutoff; recompute the Table 2 llm(GPT-5.4) and llm(Claude Opus 5) rows with the same matched estimator. If the enrichment remains near +0.196 with a confidence interval excluding zero, contamination is ruled out for the central claim; if it falls toward the classical baseline (+0.098) or zero, the headline effect is not separable from retrieval.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that an LLM importance ranking aligns with crowd highlights beyond position and length. For that to be a measure of reader alignment rather than memorization, the model must not have stored the highlights themselves. The paper's own test in §6 splits documents at the median mark month and finds no consistent recency effect, but as the paper admits, this compares mark recency, not document recency: a document that entered pretraining years ago but accumulated marks recently is placed in the 'after' group. Since the corpus is deliberately the well-read tail of a public platform whose highlight renderings are crawlable, the most plausible contamination channel is not mark date but pretraining exposure to the document-plus-highlight rendering. Cross-vendor agreement, offered as evidence of generality, is exactly what shared web pretraining predicts. The randomization test and bootstrap intervals treat the ranking as fixed and therefore cannot distinguish inference from retrieval. This is not an accusation of leakage; it is an unclosed identification gap in the construct the paper claims to measure. The model-versus-human scaling claim inherits the same gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a matched estimator for measuring whether an LLM's importance ranking aligns with naturalistic crowd highlights, after removing position and length confounds. Each crowd-marked sentence is compared against unmarked sentences at equal relative depth and equal within-document length rank; the estimator is calibrated on synthetic nulls built from position and length only. On 120 web documents, the authors report a matched enrichment of +0.196 (95% CI [+0.148, +0.239]) for GPT-5.4 against a 19.9% keep rate for matched neighbours, with p=0.0005 under an exact within-stratum randomization test, replicated with Claude Opus 5. A single human reader scores +0.182 on the same task, and the model—human paired difference is +0.002 (CI spanning zero). Classical baselines are not null: Luhn's heuristic reaches +0.088 and lexical degree centrality +0.098. The paper also reports extensive robustness checks, a clean-control analysis, a covariate balance table, a prompt-framing contrast, and a list of withdrawn claims from internal review.","tokens_in":12343,"tokens_out":5222,"duration_ms":63523,"significance":"If the identification gap is closed, this is a significant contribution to compression evaluation and to the study of LLM—reader alignment. The paper is unusually transparent: it ships a de-identified per-sentence artifact and runnable code, measures the false-positive rate of its own confound control, reports a non-replication of the authors' prior result, and explicitly bounds residual confounds. The human-reader calibration gives the headline number an interpretable scale. However, the central construct—'alignment with reader highlights'—is threatened by pretraining contamination: the corpus is deliberately drawn from the well-read tail of a public platform whose highlight renderings are crawlable, so a model could retrieve stored highlights rather than infer them from text. The paper's mark-recency test does not close this gap, because mark recency is not document recency. The tie-break non-reproducibility is a second, independent load-bearing issue. These are fixable with additional analyses or re-scoped claims, but they currently prevent full acceptance.","major_comments":[{"comment":"The mark-recency split does not test the contamination channel the corpus most plausibly exposes: a document may have entered pretraining long ago and accumulated highlights recently, so the 'after' group is not a held-out group. Since the corpus is the well-read tail of a public platform whose per-URL highlight renderings are crawlable, shared web pretraining can produce exactly the cross-vendor agreement the paper cites as evidence of generality. The randomization and bootstrap intervals condition on the ranking as fixed and therefore cannot distinguish inference from retrieval. To keep the central claim, please report a split by document publication date relative to each model's training cutoff (or a post-cutoff subset), or re-scope the conclusion to 'ranking behavior, mechanism unspecified.' Without such a control, the sentence 'an off-the-shelf language model predicts the crowd abou","section":"§6 (Pretraining contamination)"},{"comment":"At q=0.15, 1,407 of 2,095 label slots lie at the cut value and are decided by seeded jitter from the Firestore document identifier; the de-identified artifact deliberately omits that seed. The headline +0.196 cannot be exactly reproduced: the label sweep's +0.186 on 1,301 pairs is one random redraw, and the earlier front-loading tie-break gave +0.222, outside the reported redraw band. Because two-thirds of the label is determined by the tie-break, the exact p-value and enrichment are partly artifacts of an unreleased random stream. Please either include a de-identified deterministic seed (or a reproducible hash) so the exact label can be reconstructed, or report the tie-break ensemble (mean and interval over many redraws) as the headline. The paper's transparency is commendable, but the current presentation overstates the precision of the point estimate.","section":"§5.3 (Label definition and tie-break)"}],"minor_comments":[{"comment":"The notation ⊮ is nonstandard; please define it explicitly or replace with \\(\\mathbb{1}\\) or an indicator variable.","section":"§4.1, Eq. (1)"},{"comment":"Typo: 'W e state that as not established' should be 'We state...'.","section":"§5.2"},{"comment":"Reference [9] lists 'Hardy, Shashi Narayan, Andreas Vlachos'; the first author's given name appears to be missing or misformatted.","section":"References"},{"comment":"The distinction between the production tie-break stream and the sweep's random tie-break is important and clearly stated, but the row label '+0.186' could be misunderstood as a reproduction of Table 2. Consider adding a footnote or changing the row label to 'random redraw (one draw)'.","section":"§5.3 (Label sweep)"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest and artifact-rich, and I would not reject it on the current evidence. The main barrier is the pretraining contamination identification gap; the mark-recency test is not sufficient because it conflates mark recency with document recency. I would ask the authors to either provide a post-cutoff or publication-date-based analysis, or to explicitly weaken the interpretation. The tie-break reproducibility issue is also load-bearing for exact reproduction of the headline. Both are fixable within the manuscript's scope, so major revision is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a real paper, worth a serious read. The core measurement - language-model importance rankings retain crowd-marked sentences at about twice the best classical baseline, +0.196 against +0.098, p=0.0005 under an exact randomization test, replicated cross-vendor - is supported by a lot of careful work. The paper does something genuinely rare: it calibrates the confound control on synthetic nulls and shows that depth-only stratification false-positives at 20-36%. That is an important methodological warning, and it is now the first thing I would cite from this paper. The single-human-reader baseline, computed on the same estimator with the crowd label recomputed to exclude the reader, is also a thoughtful scale check, and the model-human difference of +0.002 is a striking, credible result. The transparency is not cosmetic. The appendix listing eleven withdrawn claims, including the failed covariate ladder and the lead-calibration false positive, is the kind of honesty that should make referees trust the surviving numbers. The paper also reproduces its own non-replication of prior work. That is exactly the behavior you want in this literature. The soft spots are real but not disqualifying. First, the label tie-break: at q=0.15 two-thirds of the label slots sit at the cut value and are resolved by seeded jitter, and the production seed is not in the artifact. The front-loading alternative gives +0.222, outside the redraw band, so the exact headline is not reproducible from what is shipped. The qualitative conclusion probably survives, but the paper should present the effect as a function of tie-break form, not hide behind one seed. Second, the pretraining contamination gap is genuine. The mark-recency test compares mark recency, not document recency, so a document that entered pretraining years ago but accumulated marks recently lands in the after group. The paper admits this, but the admission does not close the gap. Cross-vendor agreement is exactly what shared web pretraining predicts. The randomization test and bootstrap treat the ranking as fixed, so they cannot distinguish inference from retrieval. That limits the construct being measured: alignment with readers is not fully identified as text-based alignment. Third, the headline ratio r=0.20 sits high in a range where the effect varies by a factor of 2.7; at r=0.05 the effect drops to 42% of the headline. The abstract should not lead with the most favorable point. Still, the central claim - at least as a measurement of association, net of position and length - holds up under the perturbations the authors actually ran. This paper deserves a rigorous referee. I would send it to review, and I would ask the authors to either secure a truly held-out corpus or soften the interpretation to retrieval-robust agreement. The method and the calibration result will be useful regardless of the final interpretation.","headline":"A serious, unusually transparent measurement paper whose headline survives most of its own stress tests, but the contamination gap and tie-break sensitivity keep the result from being definitive.","tokens_in":759,"tokens_out":783,"would_cite":true,"duration_ms":28362,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language-model rankings retain crowd-marked sentences at roughly twice the rate of the best classical extractive baseline, net of position and length.","keywords":["social highlighting","crowd-marked sentences","position bias","length bias","language model ranking","confound control","context compression evaluation"],"falsifier":"Reproduce the label using a tie-break-independent rule (for example, a continuous mark count without thresholding) and recompute the matched enrichment; if it falls below the null's 97.5th percentile (+0.054), the headline is an artifact of the label construction. A simpler check: on the same 120 documents, shuffle mark assignments within depth/length strata and confirm the estimator returns enrichment near zero — if it exceeds the calibrated threshold, the control is still leaking.","tokens_in":11961,"feed_emoji":"📌","tokens_out":6518,"duration_ms":60678,"temperature":0.7,"pith_summary":"Context compression is usually judged by another model's downstream accuracy, which makes the evaluation circular. The paper uses naturalistic social highlighting — many readers independently marking the same web page — as a non-circular reference, and asks whether a language-model importance ranking keeps the same sentences readers kept. The central result is that after matching each crowd-marked sentence to unmarked neighbours at equal relative depth and equal within-document length rank, a language-model ranking keeps 38.4% of marked sentences against 19.9% of matched neighbours, an enrichment of +0.196 (95% CI [+0.148, +0.239], exact randomization p=0.0005), replicated across two vendors. The paper further shows this estimator's own false-positive rate is 4.5–7% on synthetic nulls, whereas depth-only stratification — the control one would reach for first — fires on 20–36% of nulls containing no effect. It also reports that on the same budget, a single human reader scores +0.182, indistinguishable from the language model's +0.184.","feed_headline":"Language models keep twice as many reader-highlighted sentences","feed_subtitle":"Net of position and length, a model retains 38% of crowd-marked sentences vs 20% for matched neighbours: twice the classical baseline.","key_machinery":"The argument is carried by a matched-enrichment estimator: for each crowd-marked sentence, the estimator finds unmarked sentences in the same document within 0.05 relative depth and 0.05 within-document length rank, subtracts the compressor's keep rate on those neighbours from its keep rate on the marked sentence, and aggregates across documents with domain-clustered bootstrap intervals. The companion mechanism is calibration on synthetic nulls — keep sets generated from position and length alone over the real corpus geometry — which turns the control's error rate into a measured quantity rather than an assumption. Matching discards 37.9% of marked sentences that lack an admissible comparato","core_discovery":"On the paper's own terms, the discovery is that readers' unprompted highlighting is substantially predictable by an off-the-shelf language model once the two dominant confounds — sentence position and sentence length — are removed by within-document matching. The headline number is +0.196 enrichment, with the model keeping 38.4% of crowd-marked sentences against 19.9% of matched unmarked neighbours. The same estimator gives classical word-frequency heuristics +0.088 and lexical centrality +0.098, so the model doubles rather than categorically exceeds cheap lexical selection. Scored identically, a single human reader reaches +0.182, and the paired model-minus-human difference is +0.002 (CI sp","pith_inferences":["If the matched enrichment survives on a broader, non-platform corpus, social highlighting could become a continuous, low-cost benchmark for prompt compression; the platform's convenience sample is the main external-validity limit.","Because two-thirds of the label slots at q=0.15 sit at the cutoff and the production tie-break seed is not available, the exact headline (+0.196) is not reproducible from the artifact; future label definitions should either use a continuous mark count or report a tie-break-insensitive interval.","The paper's explicit non-replication of its own earlier finding suggests that previously reported 'weak' model-human salience correlations may partly reflect uncontrolled position and length, and re-running existing datasets with a matched estimator could settle the disagreement.","The prompt-framing result — asking what is important to keep beats asking what a reader would highlight — hints that language models have implicit theories of 'importance' that differ from their theories of reader behaviour; that asymmetry is a testable handle on what the models actually learned."],"forward_implications":["Compression systems can be evaluated against reader behaviour instead of only downstream task accuracy, giving a non-circular yardstick for what a compressor preserves.","The language-model advantage over classical lexical selection is about a factor of two, so cheap word-counting methods are not null controls; an evaluation that pits a model only against a random baseline overstates the gain.","Any claim of model-human alignment on highlights must control for position and length; depth-only stratification is shown to be an unsafe control, and matching tolerance plus label contamination can move weak contrasts materially.","On the same budget, an off-the-shelf language model is statistically indistinguishable from one member of the crowd it is predicting, which gives the +0.196 number a human-scale reference.","The headline is configuration-sensitive: at compression ratio r=0.05 the effect is 42% of the r=0.20 headline, and prompt framing changes it by a factor of 2.1."],"fun_headline_variants":["Debiased LMs keep 2x as many reader-highlighted sentences","After matching position and length, AI doubles reader-marked picks","Net of length and position, LMs spot 2x more reader highlights","LM ranking keeps 2x reader marks vs matched neighbours","Reader-highlight alignment: LMs beat lexical baseline 2x after debiasing"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The crowd label is the top 15% of sentences by mark count with at least two marks, and because roughly two-thirds of the label slots sit exactly at the cutoff, the seeded-jitter tie-break decides much of the label — if that tie-breaking were meaningfully different, the headline enrichment could shift outside the reported range.","fun_headline_variants_meta":{"raw":{"variants":["Debiased LMs keep 2x as many reader-highlighted sentences","After matching position and length, AI doubles reader-marked picks","Net of length and position, LMs spot 2x more reader highlights","LM ranking keeps 2x reader marks vs matched neighbours","Reader-highlight alignment: LMs beat lexical baseline 2x after debiasing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001168,"raw_usage":{"total_tokens":4757,"prompt_tokens":921,"completion_tokens":3836,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":3739}},"tokens_in":665,"tokens_out":3836,"duration_ms":26324,"temperature":1.0,"reasoning_tokens":3739,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:12:55.715564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the label using a tie-break-independent rule (for example, a continuous mark count without thresholding) and recompute the matched enrichment; if it falls below the null's 97.5th percentile (+0.054), the headline is an artifact of the label construction. A simpler check: on the same 120 documents, shuffle mark assignments within depth/length strata and confirm the estimator returns enrichment near zero — if it exceeds the calibrated threshold, the control is still leaking.","supporting_citations":[],"review_version":1}