{"id":"f4016415-4d70-4677-9cfe-10222eaad130","arxiv_id":"2607.02587","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 12 checkpoints of Yi, Qwen, Mistral and Gemma, mean absolute adjacent-generation trust-score drift is 8.00 pp—3.6× an independence-based no-drift null—and remains elevated under leave-one-out and strict-scoring checks.","lead":"Open-source chat LLM release lines show large trust-benchmark score drift between successive generations, far above a no-drift null. Model-card trust numbers therefore cannot safely be carried forward without re-measurement and should be treated as dated, checkpoint-bound artefacts.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The independence-based null remains the softest link for the 3.60× claim, but the paper already treats it as a reference and the governance conclusion does not require a formal p-value.","rationale":"The Reader correctly isolates the independence-based reference null as the weakest assumption supporting the strongest numerical claim. The paper itself is transparent about the three independence layers and the unsigned net bias (Appendix G, Eq. 3). Because the authors already frame the comparator as a reference rather than a formal significance test, and because every leave-one-benchmark / leave-one-line / strict-scoring / constant-size perturbation still leaves |SDR| well above 3.30 pp, the operational conclusion (checkpoint-bound reporting, re-audit on material release) does not collapse even if the exact multiplier shrinks. No internal inconsistency or circular derivation is present; the remaining gaps (fixed 200-item subsets, non-canonical variants, artefacts not yet public) are already scoped out. A paired re-run with the newly persisted item traces is the single check that would convert the reference comparison into a design-valid one; until that is done the CONDITIONAL verdict with high confidence remains appropriate. I therefore leave the Reader’s verdict unchanged.","tokens_in":27033,"tokens_out":762,"duration_ms":6323,"concrete_test":"Re-run the 180 evaluations with the now-persisted per-item traces (item_id, template_id, is_correct) and recompute the primary endpoint under a genuine paired item-level permutation null that swaps generation labels inside item×template cells while jointly resampling the three templates as a block. If the observed |SDR| still exceeds the 99.9-percentile of that design-valid null, the 3.60× claim is confirmed; if it falls inside the bulk of the paired null, the reference-null elevation is an artefact of independence assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claim is that observed mean |SDR| = 8.00 pp is 3.60× an independence-based pooled no-drift reference null (mean 2.22 pp, 99.9-percentile 3.30 pp) and therefore far above sampling noise. Appendix G and §3.1 explicitly note that the null ignores two opposing effects: (i) paired-item reuse across generations (which would lower true Var(Δ) by a (1−ρ_item) factor) and (ii) positive cross-template covariance inside a checkpoint (which the null sets to zero, understating Var(Δ̄)). Because item-level correctness vectors were not retained, neither ρ_item nor the template covariances can be estimated from the submitted records, so the net bias of the comparator is unsigned. The paper correctly labels the construction a “reference null” rather than a design-valid paired test. If the omitted positive template covariance dominates, the true sampling distribution of |SDR| could be wider than the reported null, shrinking the 3.60× ratio and weakening the “outside the top 10−3 tail” language. The governance recommendation (do not carry scores forward without re-measurement) is still supported by the absolute magnitude of the drifts and by the leave-one-out robustness bundle, but the headline multiplier itself rests on an unresolvable reference comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper audits four open-source chat-LLM release lines (Yi, Qwen, Mistral, Gemma) at three successive public generations each (12 checkpoints) on a fixed basket of five trust benchmarks under three prompt templates. The pre-specified primary endpoint is mean absolute adjacent-generation Score Drift Rate, |SDR| = 8.00 pp (count-level bootstrap 95% CI [7.57, 9.12]), reported as 3.60× an independence-based pooled no-drift reference null (mean 2.22 pp; 99.9-percentile 3.30 pp) over 40 transitions. The elevation persists under strict scoring, leave-one-benchmark, leave-one-release-line, drop-low-parse, and a constant-parameter-size restriction. From this the authors conclude that trust scores attached to a named release line should not be carried forward without re-measurement, and they package a longitudinal model card (LMC) that binds scores to checkpoint identity, evaluation date, and drift-vs-prior context. Scope is repeatedly limited to the audited open-source, non-canonical, ≤12B setup.","tokens_in":27480,"tokens_out":1808,"duration_ms":28139,"significance":"If the result holds under the stated scope, the paper supplies a concrete, operationally usable argument against treating model-card trust numbers as transferable certificates across successive open-source checkpoints. Strengths that should be credited: a pre-specified primary endpoint; an explicit robustness bundle (Table 4, Appendix I) that keeps |SDR| in [6.36, 9.43] pp; a constant-size restriction (Appendix I.1) that does not collapse the headline; repeated, non-hand-waving scope tables (Tables 2 and 5); and an illustrative LMC template that is useful even if one rejects the formal null comparison. The contribution is not a new statistic but a longitudinal multi-family open-source audit pattern that governance and procurement workflows currently lack. That is a real, if narrow, contribution for cs.SE / evaluation practice.","major_comments":[{"comment":"§3.1, Eq. (1), Table 3, and Appendix G: the headline 3.60× ratio and the “outside the top 10−3 tail” language rest on an independence-based pooled binomial reference null that the paper itself shows can over- or under-state true sampling variance (paired-item reuse lowers Var(Δ); positive cross-template covariance raises Var(Δ̄); net bias is unsigned because item-level correctness vectors were not retained). For a primary endpoint this is load-bearing. Either (i) re-run the primary comparator with the now-persisted per-item traces (paired item-level permutation or template-block bootstrap) so the direction of bias is empirical, or (ii) demote the ratio/tail language in the abstract, §5.1, and Table 3 and lead with absolute |SDR| plus the leave-one-out robustness band, treating the null only as a background sanity check. The governance conclusion does not require a formal p-value, but the","section":"§3.1, Table 3, Appendix G"},{"comment":"§4 (Sampling) and §7: every reported number sits on one fixed nb = 200-item cached subset per benchmark, with no multi-seed item-subset resampling. The cell-level leave-one-benchmark / leave-one-line / drop-low-parse perturbations (Appendix I) do not substitute for item-subset uncertainty. The paper correctly flags this as “the biggest gap,” but the count-level CIs and the 8.00 pp point estimate are still presented as if subset choice were fixed truth. At minimum, either run a reduced multi-seed sensitivity on a slice of the grid and report the range of |SDR|, or move the primary numerical claim to a form that does not imply subset-stable precision (e.g., “|SDR| remains several pp above the reference null under all reported protocol perturbations”). Without one of these, the precision of the primary endpoint is overstated.","section":"§4 Sampling, §7, Appendix I"},{"comment":"Table 1 and Appendix I.1: three of eight transitions mix scale change with recipe/tokenizer change (Yi G2→G3; Gemma G1→G2, G2→G3). The constant-parameter restriction (5 transitions, |SDR| = 7.68 pp) is reassuring but is relegated to an appendix and is not reflected in the abstract or the primary Table 3. Because the central claim is about “named release line” non-transferability rather than pure recipe drift, the main text should either (a) headline the constant-size result alongside the full-grid result, or (b) explicitly redefine the estimand as “any materially new public checkpoint of a named line,” including scale, so that scale-mixing is not a confound but part of the target. As written, readers can reasonably worry that scale alone drives part of the signal.","section":"Table 1, Appendix I.1, Abstract"}],"minor_comments":[{"comment":"Figure 2 caption and Table 3: the grey band is labelled “null 99.9%ile (+/− 3.3 pp)” while the signed histogram is of individual transitions; make explicit that the band is the aggregate-null comparator, not a per-transition critical value, to avoid readers treating ±3.3 pp as a cell-level significance threshold.","section":"Figure 2, Table 3"},{"comment":"§5.2 and Tables 8–11: exploratory diagnostics (rank persistence, compliance flips, dimension volatility) are correctly labelled non-load-bearing, but the abstract and introduction still allude to “the gap persists…” without reminding the reader that only |SDR| is primary. A one-sentence separation in the abstract would help.","section":"Abstract, §5.2"},{"comment":"Table 2 / Appendix L: the BBQ-mixed ambiguous/disambiguated mix and CrowS-Pairs-FC vs PLL caveats are thorough in the appendix but easy to miss. Consider a short “variant semantics” paragraph in §4 that states what each non-canonical score does and does not measure, so absolute score levels are not over-interpreted.","section":"§4, Table 2, Appendix L"},{"comment":"Parse-rate detail (Appendix H): eight low-parse cells on binary safety under T2/T3 are well documented; a single sentence in §5.1 noting that strict scoring moves |SDR| up (to 8.93 pp) rather than down would pre-empt cherry-picking concerns more visibly.","section":"§5.1, Appendix H"},{"comment":"Related Work cites several concurrent/near-concurrent arXiv pieces by overlapping author sets (Li et al. 2026a,b; Zhuang et al. 2026; Wang et al. 2026a,b). Ensure each is used only for the specific methodological point claimed, and that the novelty paragraph does not lean on unpublished concurrent work as established prior art.","section":"§2 Related Work"},{"comment":"Terminology block in §3 is clear but long; a small glossary table (release line / generation / checkpoint / SDR / LMC) would improve skimmability for practitioners who are the intended LMC audience.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is careful, well-scoped, and more honest about limitations than most LLM-evaluation papers. The main risk for the journal is that the headline 3.60× claim is statistically softer than the abstract suggests; if the authors reframe around absolute drift + robustness and either re-run a paired null or demote the ratio, this becomes a solid, useful contribution. Citation density of concurrent same-group arXivs is high; worth a light editorial check that novelty is not circular. Fit for a software-engineering / evaluation-practice venue is good; less so for a pure ML-theory venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: on four open-source chat lines (Yi, Qwen, Mistral, Gemma) at three public generations each, mean absolute adjacent-generation drift on a fixed trust basket is 8 pp and stays well above their independence-based reference null under every leave-one-out and scoring perturbation they report. That is enough to stop treating a model-card trust number as a transferable certificate for the next checkpoint of the same brand.\n\nWhat is new is the multi-family open-source longitudinal design plus the packaging. Chen et al. already showed large closed-system snapshot change; this paper runs the open-source analogue with a pre-specified primary endpoint, a pooled no-drift reference null, and an explicit longitudinal model-card template (checkpoint hash, date, cached-sample hash, |SDR| vs prior, null summary). Scope discipline is unusually careful: non-canonical variants, 4-bit NF4, greedy decoding, fixed 200-item caches, and out-of-scope claims (closed APIs, >12B, month cadence, canonical protocols) are stated repeatedly and not smuggled into the conclusion.\n\nThe soft spots are real but proportionate. The null is independence-based and unsigned on net bias (paired-item reuse lowers variance; cross-template covariance raises it; item traces were not retained), so the 3.60× and “outside the top 10−3 tail” language is a reference comparison, not a design-valid paired test. They say so themselves. The family set is small, scale and recipe are confounded on some transitions, and multi-seed item-subset uncertainty is missing. None of that collapses the absolute drifts or the leave-one-out range [6.36, 9.43] pp, and the constant-size restriction still sits far above the null. Math and citation pattern look clean; concurrent self-cites are adjacent audit work, not circular.\n\nThis is for people who write or consume model cards, procurement reviews, or open-source governance checklists. I would bring it to reading group, cite the non-transferability result and the LMC fields, and send it to peer review. The artefacts are promised rather than public, so a referee should demand the release and an item-level re-run of the null, but the paper already earns that time.","headline":"Solid open-source longitudinal audit: trust scores on named release lines do not transfer, and the process recommendation holds even if the 3.60\times null multiplier is only a reference.","tokens_in":28065,"tokens_out":586,"would_cite":true,"duration_ms":5864,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Trust scores on open-source chat LLMs shift far more across successive generations than sampling noise would allow, so they cannot be carried forward without re-measurement.","keywords":["trustworthiness drift","open-source LLMs","model cards","longitudinal audit","score drift rate","release-line evaluation","benchmark stability"],"falsifier":"A re-run that keeps item-level correctness vectors and uses a paired item-level permutation null, or multi-seed resamples of the 200-item subsets, yielding mean absolute SDR inside or near the new null’s 99.9-percentile would falsify the claim that drift is far above sampling noise.","tokens_in":27926,"feed_emoji":"📉","tokens_out":903,"duration_ms":15726,"temperature":0.7,"pith_summary":"Model cards often quote trust-benchmark scores for a named release line as if the same number still holds when the next checkpoint ships under the same brand. This paper audits four open-source lines—Yi, Qwen, Mistral, and Gemma—at three successive public generations each on a fixed basket of truthfulness, fairness, and safety benchmarks under multiple prompt templates. Mean absolute adjacent-generation score change is 8 percentage points, 3.60 times an independence-based no-drift reference, and the elevation survives dropping any benchmark or line and switching to strict scoring. The authors therefore conclude that a trust score attached to a release line is not a transferable certificate; it must be reported as a checkpoint-bound, dated artefact. They package that practice as a longitudinal model card that records the evaluated checkpoint, the evaluation date, and the drift relative to the prior audited release.","feed_headline":"Trust scores drift 8 points between LLM generations","feed_subtitle":"Open-source release lines cannot reuse model-card numbers without re-measurement","key_machinery":"Score Drift Rate (SDR): the percentage-point change in template-mean benchmark score between two adjacent generations of one release line, aggregated as mean absolute SDR over forty transitions and compared to a pooled no-drift reference null that draws binomial scores from rates pooled across generations.","core_discovery":"Across twelve public checkpoints of four open-source chat-LLM release lines, mean absolute adjacent-generation Score Drift Rate on a fixed trust-benchmark basket is 8.00 percentage points—3.60 times an independence-based pooled no-drift reference null—and remains elevated under strict scoring and leave-one-out perturbations. A trust score attached to a named release line should therefore not be presumed to transfer to the next generation without re-measurement; it should be treated as a checkpoint-bound, time-stamped artefact.","pith_inferences":["If closed APIs show similar drift, contracts that treat a model-card score as valid until the next named version would systematically understate risk between updates.","Systems that ship continuous or silent weight changes may still need calendar-based re-evaluation even if discrete generation re-audit is enough for open release lines.","Item-level paired nulls and multi-seed subset resampling are the natural next checks that would turn the reference comparison into a design-valid test of sampling noise."],"forward_implications":["Model cards must record the exact checkpoint hash and evaluation date for every trust score.","A materially new release (weights, tokenizer, scale, or recipe) should trigger re-audit rather than carry-forward of prior scores.","Open-source release-line evaluations used in procurement or governance should not assume transfer across later generations without re-evaluation.","Longitudinal model cards should include prior-release scores, observed absolute SDR, and a reference-null summary so readers can decide whether re-audit is warranted."],"fun_headline_variants":["Trust scores drift 8 pts between open-source LLM generations","LLM trust scores fail to transfer across successive checkpoints","Remeasure: trust scores shift 8 points between LLM gens","Open-source release lines show trustworthiness score drift","Trust scores are checkpoint-bound, not reusable across gens"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The paper treats an independence-based reference null that ignores item-level pairing and cross-template correlation as a good enough yardstick for saying the observed drift exceeds sampling noise.","fun_headline_variants_meta":{"raw":{"variants":["Trust scores drift 8 pts between open-source LLM generations","LLM trust scores fail to transfer across successive checkpoints","Remeasure: trust scores shift 8 points between LLM gens","Open-source release lines show trustworthiness score drift","Trust scores are checkpoint-bound, not reusable across gens"]},"model":"grok-4.5","effort":"low","cost_usd":0.002946,"raw_usage":{"total_tokens":1059,"prompt_tokens":756,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":29460000,"prompt_tokens_details":{"text_tokens":756,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":239,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":756,"tokens_out":64,"duration_ms":2681,"temperature":1.0,"reasoning_tokens":239,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T09:30:23.698033+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A re-run that keeps item-level correctness vectors and uses a paired item-level permutation null, or multi-seed resamples of the 200-item subsets, yielding mean absolute SDR inside or near the new null’s 99.9-percentile would falsify the claim that drift is far above sampling noise.","supporting_citations":[],"review_version":1}