{"id":"f867dec9-91c8-4721-a5e4-48eee8640002","arxiv_id":"2608.11843","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using a knowledge base both to inject person names into a translator and to grade it inflates entity accuracy by up to 2.7 times, and the entire apparent gain sits in the resource-overlapping segment.","lead":"This paper measures how much a machine translation score is inflated when the same knowledge base is used to inject person names into a model and to grade the output. Using expert annotations as an independent gold, it finds the apparent gain is confined to the segment where the grading key and the injected resource overlap.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Independence claim for expert gold lacks provenance evidence; the 31.1% figure anchors the loop measurement and the injected-gold reconstruction cannot substitute for it.","rationale":"The strongest claim of the paper is that the KB-injection gain is confined to the overlapping segment, so the KB-as-gold metric measures instruction compliance. That claim is only meaningful if the scoring gold is genuinely external to the injection resource. The paper provides several pieces of indirect support: the independence rate replicates across samples (31.1% vs 31.8% in §5.6), the negative spillover in the independent segment argues against a trivial artifact, and the §5.5 provenance-swap analysis shows the reading axis does not generate the inflation. These support the internal logic but do not establish the provenance assumption. The paper itself flags the reconstruction of the injection gold as a limitation, and that is honest, but the more dangerous assumption is the one the paper treats as an operational fact: that the NIKH gold was not built from the same resources. If the KB and the gold share provenance, then the 97.8% overlap fidelity is partly a construction artifact and the 31.1% independence figure is an upper bound rather than a measurement. This is an external, verifiable fact, so the appropriate action is to keep the paper conditional and require the provenance evidence, not to reject the paper. The difference-in-differences design, the replication sample, and the transparent limitations all support that the paper has identified a real mechanism; the open question is the size of the honest effect once the provenance is verified.","tokens_in":13485,"tokens_out":2528,"duration_ms":19370,"concrete_test":"Request from the National Institute of Korean History (or from the paper's authors) the construction history of the person_master KB: who built it, from which source lists, and whether any of its entries were derived from the idx_person expert annotations or from the same human translations used to establish ETS-REF. Concretely: compare the 363 overlapping mentions against the KB entries and the idx_person tags; if the KB entry text matches the idx_person tag text for a substantial fraction of the overlap by construction rather than by linguistic convergence, the 31.1% independence figure is overstated. A practical check: re-run the §4 decomposition after removing from the 'independent' segment any mention whose KB entry could have been sourced from the same translation memory or the same annotator pool, and see whether the 97.8% vs 70.1% gap and the DiD estimate survive.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central measurement depends on the expert gold being a genuinely independent resource, but the paper never documents the provenance of the NIKH idx_person annotations relative to the person_master KB. §1.4 and §3.1 assert independence as a design fact; §4.1 then builds the 31.1% independence figure on it. If the KB was compiled from the same NIKH annotations or the same human translations, then the 97.8% vs 70.1% fidelity split in §4.2 and the loop contribution in §5.2 would be partly definitional rather than empirical. The reader's concern is valid. My additional angle: the two resources are not shown to be epistemically independent either — the human translation used to settle ETS-REF in §5.5 could share sourcing habits with the NIKH gold, since both are produced by the same institute's editorial tradition, and the paper does not state whether the human translations were produced independently of the idx_person tags. The reconstructed injection gold (§3.1, §7) is a separate weakness but does not cut as deep: the loop effect is identified through the scoring gold, which survived. The load-bearing unverified assumption remains the provenance relation between the KB and the expert gold. This is externally checkable, so the paper should be accepted conditional on that check.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper measures a 'resource-shared evaluation loop' in entity-level machine translation for the Seungjeongwon Ilgi, in which the same person-name knowledge base is used both to inject names into the prompt and to score the output. Using expert person-name annotations from the National Institute of Korean History as a supposedly independent gold, the authors find that only 164 of 527 expert-annotated mentions (31.1%) lie outside the injection pipeline. In the overlapping segment the injected reading agrees with the human translation 97.8% of the time, versus 70.1% in the independent segment. Across four models, a difference-in-differences analysis shows that the entity-accuracy gain from KB injection is confined to the overlapping segment, with the independent segment at or below zero, and that post-injection ceilings cluster in a narrow 0.910–0.996 band. The main conclusion is that the KB-as-gold metric measures instruction compliance rather than translation quality, and that the reported gain is up to 2.7 times the value obtained under an independent gold.","tokens_in":13716,"tokens_out":5915,"duration_ms":65917,"significance":"If the measurements hold, this is a valuable and timely contribution: it turns a frequently asserted evaluation bias into a quantified, replicable measurement, and it provides a practical protocol for cultural-heritage archives. The paper ships useful transparency assets: exact McNemar paired tests with extreme counts, bootstrap confidence intervals, a replication on an unfiltered sample, released model outputs and evaluation code, and an unusually candid limitations section. It also carefully separates the set axis from the reading axis of circularity, and explicitly notes that the reported gain's complement structure follows from the near-invariant ceiling. The main load-bearing assumption, as the stress-test note correctly identifies, is the provenance independence of the expert gold from the injected KB; the manuscript asserts this independence but does not document it.","major_comments":[{"comment":"The central measurement—31.1% independent mentions, the 97.8% versus 70.1% fidelity split, and the resulting difference-in-differences—rests on the assumption that the NIKH idx_person annotations used as scoring gold are provenance-independent of the person_master KB used for injection, and of the human translations used in §5.5. The manuscript asserts this as a design fact but provides no evidence: it does not state whether person_master was compiled from the same NIKH annotations or from the same human translations, nor whether the human translations were produced independently of the idx_person tags. Because the same institute appears to be the source of both resources, this is not a trivial assumption. If the KB and the gold share provenance, then the 68.9%/31.1% split and the overlap/independent contrast are partly construction artifacts rather than empirical findings. This is externally checkable and should be documented before the 31.1% figure is used as the anchor of the paper.","section":"§3.1, §4.1, §4.2"},{"comment":"The injection gold is a reconstruction, because the original was overwritten by the scoring gold in the original experiment. The segmentation of the 527 expert mentions into overlap and independent segments—and therefore the DiD in §5.2—depends on which entities the NER+KB pipeline actually injected in the original run. If the reconstructed injection set differs from the original, the segment labels used in the analysis may be misassigned, and the reported 31.1% independence figure and the DiD values may not correspond to the experiment as run. The manuscript acknowledges the reconstruction but does not quantify its sensitivity. Please provide a robustness analysis (for example, how much would the DiD change under plausible perturbations of the reconstructed injection set) or argue explicitly why exact set identity is not needed for the qualitative conclusion.","section":"§3.1, §7"},{"comment":"The 'honest' independent-gold value of 0.348 and the '2.7×' inflation figure are computed with presence-based scoring, which the paper itself notes is vulnerable to over-generation (KoBE; Alam et al.). Because injection explicitly lists the target names, a model can satisfy presence-based scoring by copying or appending the injected block, inflating the overlap segment more than the independent segment. This means the headline inflation figure may conflate the resource-sharing loop with the known scoring-protocol artifact. The paper's choice to keep the protocol fixed is defensible for measuring current practice, but the '2.7×' claim should be either repeated under the hardened protocol as a robustness check or explicitly qualified as an upper bound that includes the over-generation artifact.","section":"§5.1, §10.7"}],"minor_comments":[{"comment":"Figure 2 is referenced in §3 but does not appear in the manuscript text; the pipeline diagram should be included so that the claimed separation is visually verifiable.","section":"Figure 2"},{"comment":"The abstract and §4 use 527 scored mentions, while §5 tables use 515 mentions (or 361 on the common set); the denominator distinction should be stated at the first occurrence in the abstract or in a footnote to Table 1 to avoid confusion.","section":"Abstract and §4.1"},{"comment":"The repeated value −0.026 for three models is coincidental (−3/117 for each), and while the text explains this, the table itself should carry a footnote or dagger so that readers do not infer a shared cause.","section":"§5.2 Table 2"},{"comment":"The bullet 'The style comparison is blind but conducted by the authors' is unclear; please specify what was blinded (e.g., condition labels, model identity) and how blinding was implemented.","section":"§7 Limitations"},{"comment":"The NER model 'SillokBERT-NER' is used for the injection pipeline but is not described or cited; a reference or a brief description of its training data and output format is needed, especially given the alias/instability discussion for gemma-4-26b in §8.","section":"§3.2, §8"},{"comment":"The phrase 'This separation is an operational fact before it is a methodological requirement' is misleading: the separation of the gold from the injection pipeline is an assumption about resource provenance, not an operational fact, and this distinction should be stated explicitly.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The provenance concern is the only substantive obstacle to accepting the paper. If the authors can supply documentation showing that the NIKH idx_person annotations and the person_master KB were constructed independently, and can quantify the sensitivity of the DiD to the reconstructed injection gold, I would support acceptance. The manuscript's candor about its own limitations and its release of outputs are strengths that should be recognized in the final decision letter."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should read this one. It quantifies something that people have described qualitatively: when the same knowledge base is injected into a translation prompt and then used to score entity accuracy, the metric measures instruction following, not translation quality. The paper does this on the Seungjeongwon Ilgi with expert person-name annotations, and the numbers are concrete: only 31.1% of the 527 expert-annotated mentions fall outside the injection pipeline, and the fidelity of the gold to human translation drops from 97.8% in the overlapping segment to 70.1% in the independent one. A difference-in-differences across four models shows the injection gain is confined to the overlapping segment and is at or below zero in the independent segment, and the post-injection ceiling clusters in a 0.910–0.996 band regardless of baseline. That is a genuine new result, and the replication on an unfiltered sample is the right way to answer the obvious sampling objection.\n\nThe execution is mostly careful. They use McNemar exact tests, bootstrap CIs, report counts, and are transparent that the original injection gold was overwritten and reconstructed. They also explicitly note the Δoverlap = ceiling − pre identity, which is the kind of self-check you want. The paper does not oversell priority; it cites the circularity literature properly.\n\nThe soft spot is the independence of the expert gold. The entire loop measurement rests on the claim that the NIKH person-name annotations are not derived from the same person_master KB used for injection. The paper asserts this in §1.4 and §3.1 but never documents the provenance. If the KB was compiled from the same annotations or the same human translations, the 97.8%/70.1% split and the 31.1% independence figure become partly construction artifacts. This is externally checkable — the authors can show the construction history or at least a statement from NIKH. The reconstructed injection gold is a second-order weakness but not fatal, since the scoring gold survived. The deliberate use of presence-based scoring is also a limitation, but keeping it fixed was the right call to isolate the loop.\n\nWho is this for? Anyone building entity-level MT evaluation in low-resource or cultural-heritage domains, and anyone who uses KB-derived golds. It deserves a serious referee. I would accept it conditional on the provenance check and, ideally, on releasing unmasked outputs for the BLEU portion. The conclusions are likely correct, but the key assumption needs to be nailed down.\n\nRecommendation: send to peer review. It's a solid methodological contribution with an honest limitation.","headline":"A careful measurement of a real evaluation loop in entity-level MT, with one load-bearing provenance assumption that needs checking.","tokens_in":14283,"tokens_out":3296,"would_cite":true,"duration_ms":30985,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When the same knowledge base is both injected into a translator and used to grade it, the score measures obedience, not quality.","keywords":["machine translation evaluation","knowledge base injection","entity accuracy","evaluation circularity","resource-shared evaluation loop","historical document translation","person name translation"],"falsifier":"Trace the construction history of the person-name KB and compare its entries with the expert person-name annotations used as gold: if a substantial share share provenance, the claimed independence is inflated. A complementary check is to run the same difference-in-differences design on a corpus where the KB and the gold are verifiably disjoint and see whether the independent-segment gain stays at or below zero.","tokens_in":13264,"feed_emoji":"📜","tokens_out":9175,"duration_ms":85907,"temperature":0.7,"pith_summary":"The paper tries to establish that entity-level machine translation scores become self-referential when the knowledge base (KB) injected into the translation prompt is also used to build the gold standard. Using expert person-name annotations as an independent gold, it finds that only 31.1% of 527 scored mentions lie outside the injection pipeline, and that the loop is uneven: in the overlapping segment the injected reading matches the human translation 97.8% of the time, versus 70.1% in the independent segment. Across four models, a difference-in-differences analysis shows the gain from KB injection is confined to the overlapping segment; in the independent segment the effect is at or below zero. Post-injection ceilings cluster in a narrow 0.910–0.996 band even though baseline capability differs fivefold, so the reported gain is mostly the complement of prior performance and weaker models look like they improve more. The conclusion a sympathetic reader should take is that a KB-as-gold metric in this setting measures instruction compliance rather than translation quality.","feed_headline":"Knowledge-base grading inflates translation scores by up to 2.7x","feed_subtitle":"Across four models, the gain appears only where the grading key shares the injected resource.","key_machinery":"The central object is the resource-shared evaluation loop: a scoring gold built from the same person-name KB that is injected into the translation prompt, so a string counts as correct merely by surviving into the output. The analysis separates the two pipelines — injection uses an NER model plus KB readings, scoring uses expert person-name tags with a KB/hanja fallback — and splits the gold into an overlapping segment (gold shares the injected resource) and an independent segment (gold does not). The main instrument is a difference-in-differences contrast using the few-shot condition as the reference, which isolates the loop's contribution from the style effect of the prompt. A fixed-denominator contrast separating the set axis (which entities are scored) from the reading axis (what counts as correct) shows the inflation comes entirely from the set axis. The near-invariance of the post-injection ceiling across models is what turns the reported gain into the complement of prior performance.","core_discovery":"The central discovery is quantitative: of 527 expert-annotated person-name mentions, 31.1% lie outside the injection pipeline, and the remaining 68.9% are scored against a gold that shares the injected KB. In that overlapping segment the injected reading agrees with the human translation 97.8% of the time, while in the independent segment it agrees only 70.1% of the time, so the healthiest-looking part of the answer key is the part the loop is holding up. A difference-in-differences contrast across four models shows the entire measured gain from injection is in the overlapping segment (loop contributions from +0.294 to +0.722), with independent-segment changes at or below zero; the post-injection ceiling is nearly invariant (0.910–0.996) across baselines that differ fivefold. The paper therefore claims that the reported entity-accuracy gain is inflated up to 2.7 times the honest value, and that the loop, not translation quality, is what the metric records.","pith_inferences":["Editorial inference: If the person-name KB was assembled from the same expert annotations or the same human translations as the scoring gold, the independence rate and the 70.1% agreement would both be construction artifacts; a provenance audit settles this.","Editorial inference: The paper's linear projection from 70% KB coverage to lower coverages implies that simply reporting KB coverage alongside an entity metric would let low-resource archives estimate their own distortion before paying for gold.","Editorial inference: A hardened scoring protocol with over-generation and positional penalties might change absolute levels, but the confinement of the gain to the overlapping segment should persist because the loop operates on which entities are scored, not on how presence is checked.","Editorial inference: The same design — separate injected from scored resources, fix the denominator, difference by segment — transfers to any low-resource domain that substitutes a dictionary for expert gold, such as clinical or legal corpora."],"forward_implications":["Entity-accuracy gains reported from KB injection are inflated up to 2.7 times the value measured against an independent gold, so archives should discount headline figures.","Injection does not generalize: across all four models the independent-segment change is at or below zero, with the strongest model slightly worse.","Weaker models appear to improve more dramatically because post-injection ceilings cluster in a 0.910–0.996 band while baselines differ fivefold; the reported gain is largely the complement of prior performance.","A KB-derived gold penalizes more capable models more heavily, so it is unsuitable for cross-model absolute comparison even though it tracks within-condition deltas.","Splitting the gold into overlapping and independent segments and reporting KB coverage reproduces the paper's central difference-in-differences and lets an archive estimate its own exposure."],"supporting_citations":[{"why":"Documents reference bias, establishing that evaluation keys pull scores toward themselves.","marker":"Fomicheva & Specia, 2016"},{"why":"Shows pseudo-reference bias, a prior form of circular evaluation this paper specializes.","marker":"Albrecht & Hwa, 2007"},{"why":"KoBE provides the presence-based entity scoring rule and the over-generation penalty diagnosis this paper keeps fixed.","marker":"Gekhman et al., 2020"},{"why":"Demonstrates that presence-only scoring fails to catch a system that appends the required term, motivating the protocol caveat.","marker":"Alam et al., 2021"},{"why":"Names circularity in RAG evaluation, the structure this paper transfers to entity-level MT with a shared KB.","marker":"Dietz et al., 2026"},{"why":"CALBC silver standard shows procedural blocking of resource sharing, the remedy against which the loop is measured.","marker":"Rebholz-Schuhmann et al., 2011"},{"why":"M-ETA defines manual entity translation accuracy against human gold, the kind of gold this paper uses as the independent check.","marker":"Conia et al., 2025"}],"fun_headline_variants":["Self-referential grading inflates entity translation scores up to 2.7x","Evaluation loop, not quality, drives reported translation gains (2.7x)","Only 31% of entity test escapes the KB loop; rest inflate scores 2.7x","Grading with injected KB inflates translation gains up to 2.7x","When the answer key shares the KB, translation scores inflate 2.7x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the expert gold used for scoring is genuinely independent of the KB used for injection; if the KB was built from the same annotations or the same human translations, the 31.1% independence figure overstates how much of the gold escapes the loop.","fun_headline_variants_meta":{"raw":{"variants":["Self-referential grading inflates entity translation scores up to 2.7x","Evaluation loop, not quality, drives reported translation gains (2.7x)","Only 31% of entity test escapes the KB loop; rest inflate scores 2.7x","Grading with injected KB inflates translation gains up to 2.7x","When the answer key shares the KB, translation scores inflate 2.7x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002192,"raw_usage":{"total_tokens":8560,"prompt_tokens":1090,"completion_tokens":7470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":7359}},"tokens_in":706,"tokens_out":7470,"duration_ms":51448,"temperature":1.0,"reasoning_tokens":7359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:24:27.363408+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Trace the construction history of the person-name KB and compare its entries with the expert person-name annotations used as gold: if a substantial share share provenance, the claimed independence is inflated. A complementary check is to run the same difference-in-differences design on a corpus where the KB and the gold are verifiably disjoint and see whether the independent-segment gain stays at or below zero.","supporting_citations":[],"review_version":1}