{"id":"ef39ec17-3a14-4330-96c6-d3f098356027","arxiv_id":"2607.27595","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A tool-constrained LLM extracts span-grounded, typology-labeled intertextual pairs; expert-adjudicated validation and a 65,380-comparison run across the Twenty-Four Histories yield stable citation composition but declining literal fidelity.","lead":"An AI reads classical Chinese histories and marks each reuse of the Analects—quote, paraphrase, or silent borrowing—with exact character spans and a five-part label of how and why it was reused. After expert validation of twelve models, the best extractor scanned all Twenty-Four Histories, finding 5,766 reuses and revealing stable citation style but decreasingly literal wording across 18 centuries.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recall is unmeasured and the deployment model surfaces only ~15% of the validation gold pairs; if recall varies with composition date, the diachronic findings in Tables 3 and Fig. 4 are confounded.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the corpus-level diachronic conclusions assume that the extractor's precision transfers and that recall is time-invariant. I agree, and the paper's own numbers make the concern sharper than the reader's phrasing alone suggests: the deployment model's recall on the validation gold is only 390/2,533 ≈ 15%. This is a concrete basis for the worry that the model's output is a highly selective sample. The paper's confidence-subset analysis and its framing as 'calibrated model judgments' do not address recall representativeness. I considered alternative concerns, such as indirect transmission through intermediate histories or span-boundary effects on fidelity; these could testably bias the interpretation, but the unmeasured-recall issue is more fundamental because it threatens the existence of the empirical patterns, not just their explanation. The proposed concrete check—per-period recall measurement via a stratified gold build—would directly test whether recall varies with composition date, and the instruction to weight or restrict the analysis provides a clear path to assess sensitivity. Since the paper is already marked CONDITIONAL and this is the same concern the reader raised, no verdict change is warranted; the condition should be resolved by the suggested validation. I also credit the paper's protocol and benchmark as genuine contributions, which is why the concern is framed as a call for additional validation rather than a rejection.","tokens_in":13998,"tokens_out":10862,"duration_ms":119077,"concrete_test":"Construct a per-period gold standard by running the same pooled twelve-model pipeline on a stratified sample of scrolls from early- and late-composed histories (e.g., 20 scrolls each from Shiji, Hanshu, Xin Tangshu, and Songshi), with the same expert adjudication, then compute deepseek-v4-flash's recall against this gold for each period. If the early-vs-late recall difference exceeds a few percentage points (or if recall differs by form or source-marking category), re-estimate Tables 3 and Figure 4 with inverse-probability weighting or by restricting to pairs independently proposed by at least two models; if the diachronic patterns shift materially, the recall-transfer assumption fails and the headline finding is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—stable interpretive composition with declining literal fidelity across eighteen centuries—rests on corpus-wide output from a single extractor (deepseek-v4-flash). Precision was validated on the Analects–Book of Han pair, but recall was never measured. In the validation set, that model recovers only 390 of the 2,533 adjudicated gold pairs (15.4%; Table 2 vs. the gold size), so the 5,766 corpus pairs are a narrow, uncharacterized slice of the true intertextual space. The paper explicitly describes the output as 'a distribution of calibrated model judgments' (Scaling section) and the limitations address only precision transfer ('First, corpus-scale extraction uses a single model calibrated on the Book of Han'), not recall. If detection sensitivity varies with composition date—for example, higher sensitivity for verbatim quotations in later formulaic histories and lower for early paraphrases, or vice versa—then both the null result in Table 3 (stable label mix) and the negative fidelity correlations in Figure 4 could be artifacts of which pairs the model surfaces, rather than properties of the textual tradition. The confidence≥0.9 subset analysis (55.3% of pairs) does not mitigate this: confidence is a validity-correlated score, not a recall-bias diagnostic. A low-recall extractor can still produce high-precision labels, but the distribution of those labels over time is trustworthy only if the missed pairs are missing at random with respect to century, genre, and interpretive dimension—an assumption for which no evidence is given.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper recasts fine-grained intertextuality extraction as an agentic LLM task: a model reads two full text units, grounds each proposed reuse in exact character spans through a constrained tool interface, and labels it along five interpretive dimensions. The method is validated on the Analects–Book of Han pair, where two experts plus an arbiter adjudicate a pooled multi-model candidate set into a 2,533-pair gold standard. Twelve LLMs are compared on precision, cost, latency, and calibration. The validated extractor is then scaled to all Twenty-Four Histories (65,380 chunk-pair tasks, 5,766 extracted pairs), and the paper claims that the interpretive composition of Analectscitation is stable across eighteen centuries while literal fidelity to a given passage declines. The protocol, benchmark, and code/data are released.","tokens_in":14339,"tokens_out":8123,"duration_ms":83649,"significance":"The methodological contribution is strong and valuable. The tool-enforced grounding with exact character spans, write-time schema/duplicate checks, hermetic execution, and explicit abstention produce auditable, verifiable annotations rather than parsed model prose. The expert-adjudicated gold standard with a reported agreement gradient provides a careful evaluation resource, and the twelve-model precision/cost/calibration comparison is informative. The oracle retrieval analysis convincingly shows the limitations of similarity-based prefiltering for paraphrase. If the corpus-scale diachronic findings are robust, the paper would be a significant advance for digital humanities and computational text reuse. However, those findings currently rest on an unmeasured recall assumption, and the 'no systematic change' claim is a null result without equivalence testing; the central historical contribution therefore needs additional support or more careful qualification.","major_comments":[{"comment":"The corpus-scale historical claims rest on an unmeasured and, on the validation set, low recall. The deployment model (deepseek-v4-flash) contributes 390 of the 2,533 adjudicated gold pairs (15.4%; Table 2 vs. the gold-size figure), and because the gold pool is itself the union of the twelve models' proposals, the benchmark cannot bound what the models miss. Precision alone cannot validate a diachronic distribution: if detection sensitivity varies with composition date, genre, or fidelity, the stable-composition result in Table 3 and the declining-fidelity result in Figure 4 could be artifacts of which pairs the extractor surfaces. The confidence≥0.9 subset is a precision-correlated filter, not a recall-bias diagnostic. The authors acknowledge precision transfer but not recall transfer. I ask for a recall estimate on at least one additional adjudicated source pair or a sampling-based det","section":"Scaling to the Twenty-Four Histories / Table 2 / Discussion and Limitations"},{"comment":"The 'no systematic change' conclusion is an unsupported null. Table 3 reports Spearman p-values of .12–.28 and a permutation test for Jensen–Shannon divergence, but failure to reject the null is not evidence of stability; no confidence intervals, equivalence bounds, or power analysis are given. With 24 histories and one composition date per multi-century work, the test is coarse. Please report effect sizes with confidence intervals or an explicit equivalence margin, and in the Abstract use 'no significant trend was detected' rather than 'shows no systematic change'.","section":"Scaling to the Twenty-Four Histories / Table 3"},{"comment":"The headline claim, 'Across eighteen centuries the interpretive composition of citation shows no systematic change while the same passage is quoted ever less literally,' is stronger than the evidence supports. The composition claim inherits the recall/representativeness issue (Major 1) and the null-result issue (Major 2); the fidelity decline is estimated on model-surfaced pairs from a single extractor without a recall correction. Additionally, Table 3 includes function and stance, which the paper itself labels exploratory (Discussion and Limitations) because of low inter-annotator agreement; the Abstract should specify that the stable-composition claim rests on form and source-marking, or should mark the other dimensions explicitly as exploratory.","section":"Abstract / Conclusion"}],"minor_comments":[{"comment":"The phrase 'exhaustive comparison of the Analects with the Book of Han' should be clarified: the chunk-pair comparison is exhaustive, but the gold standard is a candidate pool, not an exhaustive reference. Suggest 'exhaustive chunk-pair comparison' to avoid overstatement.","section":"Abstract"},{"comment":"The fidelity measure is described only as 'character-level longest common subsequence' in one sentence. Please specify the normalization (e.g., divided by source or target length) and how span-boundary differences are handled, since Figure 4's quantitative claims depend on it.","section":"Scaling to the Twenty-Four Histories"},{"comment":"Table 1 is informative, but the 'Chance' column would be clearer if it stated that chance is computed from the annotators' marginals, as the text already says. Consider adding a footnote to the table so the column is self-contained.","section":"Expert-Adjudicated Evaluation / Table 1"},{"comment":"The statement 'corpus totals are lower bounds' is too weak: the issue is not only totals but distributional representativeness. Please state explicitly that lower-bound status applies to counts, not to proportions or trends.","section":"Discussion and Limitations"}],"recommendation":"major_revision","confidential_remarks":"This is a strong paper with a valuable, carefully engineered protocol and benchmark. The main risk is the unmeasured recall of the corpus-scale extractor, which directly affects the paper's headline historical claim. I believe this is fixable in revision: additional recall evidence on a second source pair, or a clear reframing of the diachronic conclusions as properties of extracted pairs, would move it to acceptance. The paper does not have novelty or attribution concerns; it is a good fit for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The verifiable commitment protocol is the real contribution: span-grounded, tool-constrained, write-time checked annotations, with explicit abstention and stop guards. That turns an LLM's prose into auditable structured data. The twelve-model study is also well done — cost/precision/calibration separation, 51x spread, and the oracle retrieval baseline showing that retrieve-then-judge can't recover trigram-free reuse for a large chunk of the gold. Credit where due: the benchmark construction, adjudication, and agreement gradient analysis are careful and honest.\n\nThe soft spot is exactly what the stress-test note says: recall is unmeasured. The gold standard is the union of model proposals, so precision is the only quantity you can compute, and the corpus run inherits that. The deployment model surfaces only 390 of the 2,533 gold pairs on the validation set — 15.4% recall by the paper's own numbers, though the paper never frames it as recall. That matters because the headline historical claims (stable composition, declining fidelity) are statements about the population of intertextual pairs, not about what one extractor happens to surface. If detection sensitivity varies by century or genre — and nothing rules that out — the diachronic patterns in Table 3 and Figure 4 could be artifacts of which pairs the model commits. The confidence≥0.9 subset doesn't repair this; it's a precision filter, not a recall probe.\n\nI don't think this kills the paper. The systems contribution stands. And the authors are unusually explicit about limits: they frame the output as calibrated model judgments, flag the single-model transfer, and call function and stance exploratory. But they then interpret the distributional output as if it described the textual tradition rather than the model's calibrated judgment stream. That is a stretch, and the conclusion should be softened or backed with a second adjudicated evaluation on a later history, or a recall-oriented dense sample.\n\nThis paper is for digital humanists and NLP people working on text reuse; anyone building corpus-scale annotated extraction will learn from the protocol. The methodology is a genuine advance. The historical claims need to be reined in, but the work deserves a serious referee.","headline":"A careful, genuinely useful systems paper whose headline historical claims overreach the evidence: recall is never measured, and that weakens the diachronic conclusions.","tokens_in":14833,"tokens_out":1725,"would_cite":true,"duration_ms":18794,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that fine-grained intertextuality—how a text reuses another, not just where—can be extracted at scale by an LLM agent whose every proposed pair is grounded in exact character spans and labeled under a five-dimension typolo","keywords":["intertextuality","text reuse","LLM agent","span grounding","classical Chinese","Analects","Twenty-Four Histories","cultural attraction"],"falsifier":"A second expert-adjudicated benchmark on a late history (for example, the Ming History) with an independent enumeration of intertextual pairs—including a count of how many pairs share no character trigram—would settle the transfer assumption. If the model's recall on such a benchmark is measurably lower than on the Book of Han, the declining-fidelity and stable-composition patterns would be artifacts of differential detection sensitivity across time, not properties of the tradition.","tokens_in":13880,"feed_emoji":"📜","tokens_out":6962,"duration_ms":56923,"temperature":0.7,"pith_summary":"This paper recasts intertextuality extraction from a similarity-scoring problem into a grounded relation-extraction task: an LLM reads one Analects book against one history scroll in full and, through a constrained tool interface, must localize every reused fragment to exact character spans on both sides and label it on five dimensions—form, aspect, source-marking, function, stance. The authors validate the protocol on an exhaustive comparison of the Analects with the Book of Han, where three experts adjudicate 3,489 pooled candidates into a 2,533-pair gold standard, then scale the validated extractor to all Twenty-Four Histories, producing 5,766 pairs. The central finding is diachronic: the interpretive composition of Analects citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. The paper argues this stability-in-the-aggregate with drift-in-the-individual is exactly what a cultural-attraction account predicts, and that this structure is inexpressible in a similarity score.","feed_headline":"18 centuries of citation style held steady in Chinese histories","feed_subtitle":"A span-grounded AI extractor found 5,766 Analects reuses across 24 histories—wording drifted while practice held.","key_machinery":"The load-bearing mechanism is the span-grounded agentic extraction protocol: an LLM agent reads both text units in full, then interacts with tools (exact substring positioning, pair commitment with verbatim re-slice and schema validation, list/remove revision, and a guarded final submit) so that a prediction is a correct annotation only if both fragment texts occur verbatim at their stated offsets. The five-dimension typology (form, aspect, source-marking, function, stance) converts each grounded pair into a describable reuse event. The validate-then-scale design—adjudicated gold on one source pair, then deployment of the chosen model to the full corpus—is what turns a small expert effort in","core_discovery":"The paper's central claim is that a promptable LLM, constrained by a tool interface that verifies every annotation at write time, can transform the detection of text reuse into describable reception history. Each intertextual pair is committed only after exact substring lookup grounds both fragments and a schema check enforces the five-dimension typology; the final submission is checked against the committed set, so the output is auditable rather than parsed from model prose. Validated on an expert-adjudicated gold standard of 2,533 pairs, the protocol scales to 5,766 pairs across all Twenty-Four Histories. With this record, the paper establishes that the way the Analects is cited—its form,","pith_inferences":["If the protocol transfers beyond the Analects–histories pair, the same validate-then-scale design could be applied to other canonical corpora (the Odes, the Zhuangzi, or Latin and Greek texts), enabling cross-tradition comparisons of how citation practice evolves—or fails to.","The reliability gradient implies that future annotation designs should treat function and stance as multi-label or as targets for contested-judgment adjudication, rather than as forced single-choice dimensions, if these intent-laden categories are to carry load-bearing claims.","The fidelity decline held per passage suggests a testable 'audience effect': when an authority is named, wording fidelity rises; a cross-linguistic replication could reveal whether this is a property of canonical transmission generally or specific to the Chinese historiographic tradition.","A practical consequence left implicit: because precision and cost vary independently across models, deployment choices for scholarly extraction should be made on precision-per-dollar and calibration curves, not on benchmark rank alone."],"forward_implications":["Similarity scores cannot substitute for typology: distributions of form and source-marking are annotatable with confidence, while function and stance are contested; any corpus-scale claim about reuse should be built on the reliable dimensions.","Corpus-scale reception history becomes auditable: every one of the 65,380 tasks ended in a well-defined outcome, and every pair is a write-time-checked commitment, so a reader can verify a claim by inspecting any pair in context.","The null result for interpretive composition, combined with the fidelity decline, provides quantitative evidence for a cultural-attraction account of transmission: aggregate stability coexists with individual drift.","Marking and fidelity are distinct axes: named citations are more literal than unmarked ones even among direct quotations, so a similarity score cannot stand in for the marking label.","Fixed segmentation misses asymmetric reuse: many-to-one relations, where one history span answers to two or more disjoint Analects passages, are expressible only with free span localization."],"fun_headline_variants":["Stable citation style, drifting wording: 24 Chinese histories analyzed","Across 24 histories, citation practice stable but wording drifts","Expert-adjudicated AI finds citation stability amid wording drift","The Analects' reuse across 24 histories: style steady, wording shifts","Grounded AI extraction maps 5,766 intertextual pairs in Chinese histories"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the extractor's precision—and more importantly its sensitivity to paraphrase—measured on the Analects–Book of Han pair transfers unchanged to all twenty-four histories, so that the observed diachronic patterns are properties of the texts rather than artifacts of which pairs the model happens to surface.","fun_headline_variants_meta":{"raw":{"variants":["Stable citation style, drifting wording: 24 Chinese histories analyzed","Across 24 histories, citation practice stable but wording drifts","Expert-adjudicated AI finds citation stability amid wording drift","The Analects' reuse across 24 histories: style steady, wording shifts","Grounded AI extraction maps 5,766 intertextual pairs in Chinese histories"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001166,"raw_usage":{"total_tokens":4706,"prompt_tokens":835,"completion_tokens":3871,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":3778}},"tokens_in":579,"tokens_out":3871,"duration_ms":21461,"temperature":1.0,"reasoning_tokens":3778,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:01:36.546387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A second expert-adjudicated benchmark on a late history (for example, the Ming History) with an independent enumeration of intertextual pairs—including a count of how many pairs share no character trigram—would settle the transfer assumption. If the model's recall on such a benchmark is measurably lower than on the Book of Han, the declining-fidelity and stable-composition patterns would be artifacts of differential detection sensitivity across time, not properties of the tradition.","supporting_citations":[],"review_version":1}