{"id":"3f2c9f52-4d69-4ef1-a7be-a22d503400f2","arxiv_id":"2608.11803","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"No AI provider in the studied sample exposes a publicly verifiable link between the model it serves and the model it evaluated, so published safety results cannot be reliably tied to deployed systems.","lead":"The paper examined public documentation and APIs from nine AI model providers and seven hosting services, and found that none let an outside party verify that the model being served is the one described in safety evaluations. It introduces a scorecard for measuring this disclosure gap and proposes rules for when providers should announce deployment changes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"M2=0/9 is an absence-of-evidence finding: the paper searches for provider-published mappings but never attempts the constructive round-trip that its own definition allows.","rationale":"The reader's weakest assumption correctly identifies the exclusion of third-party behavioral fingerprinting as the main soft spot in the operationalization of M2. I agree that this is the most load-bearing concern, because the chain-of-custody conclusion rests on the M2 zero count. My emphasis is slightly different: rather than claiming the paper definitionally excludes fingerprinting, I read the Table 1 definition as broad enough to include any public information, including measured API behavior. The problem is that the empirical investigation never attempts the constructive round-trip that the definition describes; it searches for a documented mapping but does not try to build one from behavioral fingerprints. This is an absence-of-evidence finding, not a tested absence. A successful fingerprint-based match for a single provider would overturn the headline zero. However, the concern is not sufficient to reject the paper: the evidence tables, pre-registered rubric, and documented provider statements support the narrower claim that no provider publishes a binding primitive, and the paper's own limitations section acknowledges the single-rater and observability constraints. The reader's CONDITIONAL verdict already captures this residual uncertainty, so no verdict adjustment is needed. The proposed concrete test would settle whether the concern lands by attempting the missing constructive direction.","tokens_in":909,"tokens_out":1554,"duration_ms":82022,"concrete_test":"For each of the nine first-party providers, choose the most recent system card or model card that reports quantitative safety metrics. Using only standard API access, collect a fixed prompt battery and response log-probabilities or output distributions from the provider's stable alias and from any dated snapshot identifier named in the card. Pre-register a matching rule, e.g., a distributional distance below a threshold calibrated to a false-positive rate below 1%. If any stable alias can be uniquely matched to the specific evaluated artifact on the basis of public API behavior, re-score M2=1 for that provider; if no match clears the threshold, the M2=0/9 finding survives this constructive test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that no provider exposes a verifiable API-to-evaluation round-trip (M2=0/9) and therefore no externally checkable chain of custody exists. Table 1 defines the round-trip as: 'An external party can map a model's API response back to the specific evaluation artifact describing the served snapshot using only public information.' The appendix scoring note similarly credits M2 when an external evaluator could connect the deployed model to a specific published safety evaluation using publicly available information. This definition does not restrict the mapping to provider-published binding primitives: behavioral fingerprints obtained from standard API calls are public information, and the paper states in its Scope section that it collects 'behavioral fingerprints obtained through standard API calls.' Yet the empirical implementation searches provider documentation, archived snapshots, and API references for counterexamples, rather than attempting the constructive direction: actually querying a provider's stable aliases and dated snapshots, measuring their behavior, and testing whether any served system can be uniquely matched to a specific evaluated artifact. If such a fingerprinting-based round-trip succeeds for even one provider, the headline M2=0/9 result and the chain-of-custody conclusion would be false. The concern is not that third-party fingerprinting is philosophically excluded; it is that the evidence base never tests the definition's full scope, so the zero is an unsupported absence claim rather than a tested absence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether an externally verifiable chain of custody connects published safety evaluations to the model artifacts actually served by foundation-model providers. It defines silent updates as post-deployment changes without public disclosure, introduces a Silent Updates Scorecard with twenty-nine questions across seven sections, and applies it to nine first-party API providers and seven inference hosts. The main empirical findings are that providers publish substantial safety documentation but none exposes a verifiable API-to-evaluation round-trip (M2=0/9) or a version-bound content policy (M6=0/9), and only Anthropic partially names a pinned evaluated snapshot (M1). The paper also documents Terms-of-Service restrictions on benchmarking and proposes a Three-Part Behavioral Trigger System together with a regulatory safe harbor for public-interest benchmarking. The evidence base includes archived documentation, more than 270 evidence rows with URLs and access dates, a pre-scoring rubric, a zero-as-claim-to-be-falsified procedure, and a public repository with the full scorecard and dataset.","tokens_in":16964,"tokens_out":9946,"duration_ms":104920,"significance":"If the empirical claims hold, the paper makes a useful and timely contribution to AI governance: it operationalizes the verification gap, provides a reproducible measurement instrument, and identifies specific binding primitives that are missing across major providers. The strengths include the pre-specified rubric, the documented evidence rows, the independent evidence audit, the explicit treatment of zero scores as claims to be falsified, and the public code and data repository. The paper also credits partial successes such as Anthropic's pinned snapshot and Together AI's quantization disclosure, which makes the finding more credible. The central caveat is that the M2 zero-count is measured as an absence of provider-published documentation, even though the paper's own definition of the round-trip is broader; the strong conclusion that external verification is impossible is therefore not fully established by the presented evidence.","major_comments":[{"comment":"The operational definition of M2 is broader than the empirical implementation. Table 1 defines an API-to-evaluation round-trip as one in which 'an external party can map a model's API response back to the specific evaluation artifact describing the served snapshot using only public information,' and the Scope section states that the evidence base includes 'behavioral fingerprints obtained through standard API calls.' The empirical search, however, looks for provider-published mappings in documentation, archives, and API references; it never attempts the constructive round-trip by querying stable aliases and dated snapshots, measuring their behavior, and testing whether any served system can be uniquely matched to a specific evaluated artifact. If black-box fingerprinting can establish such a match, the M2=0/9 result and the conclusion that 'it is not currently possible to connect evaluation results to deployed systems' overstate the verification gap. Please either run a constructive round-trip probe for at least a subsample of providers, or narrow the central claims to 'no provider publishes a documentation-based mapping.'","section":"§4 Methodology; Table 1; Appendix 'Scoring Notes' (M2)"},{"comment":"The chain-of-custody protocol states that each evaluation is coded on three dimensions and assigned to one of four outcomes (chain established, version unspecified, version inaccessible, methodology underspecified), but the results section never reports the distribution of these outcomes or the per-provider values for all three dimensions. Only the M-section scores and the M1/M2/M6 zero counts are presented. Because the chain-of-custody claim is a central result, the reader cannot see where the chain breaks for each provider beyond the two zero primitives. Please add a per-provider table reporting the three coded dimensions and the resulting outcome category.","section":"§5 Chain-of-Custody Protocol; §7 Chain-of-Custody Analysis"},{"comment":"The scorecard was coded by a single rater, and the independent evidence audit described in the paper is not a blinded second coding. For a newly introduced measurement instrument whose rankings and percentage scores are reported in Table 3 and used in comparative statements, the absence of any inter-rater reliability estimate is a load-bearing measurement concern. Please provide a reliability check on a subsample, such as two independent raters scoring a subset of providers, or explicitly label all comparative numeric claims as single-rater descriptive scores and avoid ranking language. The Limitations paragraph acknowledges the issue, but the acknowledgment does not resolve the concern.","section":"§4 The Scorecard Instrument; §5 Results; Table 3"}],"minor_comments":[{"comment":"The sentence 'Only four of nine publish version-level behavior change documentation (C2=3/9)' is internally inconsistent; either the prose number or the parenthetical count is wrong.","section":"§6 Failure Mode 2"},{"comment":"The statement that Cohere is 'the highest-scoring provider in our Scorecard on transparency artifacts' conflicts with Table 3, where Cohere ranks second behind OpenAI; please rephrase to 'among the highest-scoring providers' or specify the subset being compared.","section":"§6 Contractual Limits on External Verification"},{"comment":"The percentage column appears to normalize against provider-specific applicable maxima, since Cohere's 22 points and AI21's 12 points correspond to 62.9% and 34.3% only if their applicable denominator is 35 rather than 37. The caption should state each provider's applicable denominator or add a note explaining the normalization.","section":"Table 3"},{"comment":"The mean first-party score of 16.4 corresponds to 44.3% of the raw maximum of 37, while the average of the normalized percentages in Table 3 is about 45.0%; please clarify which quantity is being reported as 44.4%.","section":"§5 Results"},{"comment":"The phrase 'Across 9 AI providers providers' contains a duplicated word and should be corrected.","section":"Conclusion"},{"comment":"The definition of a silent update relies on 'materially affects user-facing behavior' without an operational criterion, and the trigger thresholds introduced later are explicitly illustrative. Please add a sentence stating whether any of the empirical scorecard findings would change under alternative materiality thresholds, since the central claims otherwise appear to depend on this undefined term.","section":"§3 Defining Silent Updates"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about M2 lands: the paper's zero-round-trip finding is an absence-of-evidence result based on documentation search, while the definition and scope text admit behavioral fingerprints as relevant public information. This is fixable by narrowing the claims or adding a constructive probe, so I do not recommend rejection. The paper is not circular: scores are assigned from public documentation and archived evidence under a rubric fixed before scoring, and the trigger thresholds are explicitly illustrative. The missing chain-outcome table and the lack of inter-rater reliability are also addressable in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Useful paper. It quantifies something people have been gestating at: providers publish lots of safety evaluation material but almost nothing that binds those evaluations to a specific deployed snapshot. The Silent Updates Scorecard is a real instrument, the evidence base is documented (270+ rows with URLs and dates), the rubric was pre-registered, and the zero-falsification pass is a good habit. Credit where due: Anthropic's M1 is scored partial rather than full, and the Replicate content-hash example is a nice feasibility proof that the missing primitives are choices, not technical impossibilities. The citation pattern is appropriate; FMTI and the drift literature are the right touchstones.\n\nThe main soft spot is the M2=0/9 headline. The operational definition in Table 1 includes behavioral fingerprints obtained through standard API calls, and the appendix says M2 is about whether an external evaluator could connect the served model to a published evaluation using public information. That is a capability statement. But the empirical implementation only searches provider documentation, archives, and API references; it never actually attempts the constructive direction—querying a snapshot, fingerprinting it, and trying to match it to a published evaluation artifact. So the zero is an absence-of-evidence result for the definition as written. The paper often phrases the finding more carefully (no provider published information allowing...), but the conclusion that it is not currently possible overstates what was tested. Either narrow M2 to provider-published round-trips or run a small pilot of the fingerprinting route.\n\nOther limits are more minor: single-rater coding without inter-rater reliability (the independent evidence audit partially mitigates this), and the sample, while reasonable, is nine providers and seven hosts, so the generalizability claims should stay modest. The trigger thresholds are explicitly illustrative, which is fine. The ToS benchmarking analysis is interpretive but clearly labeled.\n\nThis paper deserves a serious referee. The empirical core is useful for AI governance and audit-tooling audiences, and the scorecard is something others can apply and extend. I would bring it to reading group and would cite it. Main revision ask: fix the M2 framing and address the constructive test.","headline":"Useful, well-evidenced scorecard; the zero-round-trip headline is real as a statement about provider-published mechanisms, but the paper never tests the behavioral-fingerprinting route its own definition allows.","tokens_in":17479,"tokens_out":4305,"would_cite":true,"duration_ms":42507,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across nine first-party API providers, none exposes a verifiable chain from safety evaluation to served model.","keywords":["silent updates","chain of custody","post-deployment disclosure","safety evaluation traceability","model versioning","behavioral drift","AI governance","API transparency"],"falsifier":"Run the chain-of-custody protocol on a provider scored zero: find a public safety evaluation that names an immutable, API-callable snapshot identifier, then confirm through public documentation and standard API calls that a response from that identifier corresponds to the evaluated artifact and that the content policy is version-bound; one passing provider would refute the universal zero claim. Alternatively, mount an independent black-box comparison that reliably matches an API's output distribution to a specific evaluated snapshot, which would show that an externally verifiable round-trip exists even without provider-published primitives.","tokens_in":16497,"feed_emoji":"🔍","tokens_out":7335,"duration_ms":72635,"temperature":0.7,"pith_summary":"Deployed foundation models can change after release through fine-tuning, classifier updates, prompt revisions, retrieval changes, and routing changes, often without disclosure or a version increment. The paper argues that this \"silent update\" problem is not hypothetical: across nine first-party API providers and seven inference hosts, providers typically publish substantial safety documentation, but none publishes the binding information that would let an outsider verify that the served artifact is the one the documentation describes. The authors introduce the Silent Updates Scorecard, a 29-question instrument measuring versioning, changelogs, cross-surface transparency, deprecation, evaluation traceability, reproducibility, and monitoring support, and report that the primitives needed for a chain of custody are absent or nearly absent: no provider enables an external API-to-evaluation round-trip, no provider version-binds its content policy, and only one provider partially names the evaluated snapshot. The paper then proposes a Three-Part Behavioral Trigger System tying disclosure obligations to capability shifts, behavioral drift, and component changes. If the finding holds, current AI governance frameworks rest on an unverified assumption that evaluations and system cards describe the system users actually encounter.","feed_headline":"None of nine AI providers links safety tests to served model","feed_subtitle":"An audit of major API providers finds published safety evaluations cannot be tied to the deployed system.","key_machinery":"The central mechanism is the chain-of-custody protocol, defined as a verifiable link among three artifacts: the model identifier returned by an API, the snapshot named in a safety evaluation, and the content policy governing that snapshot. The Silent Updates Scorecard measures this through three binding primitives: snapshot binding, meaning an evaluation names an immutable, API-callable identifier; an externally verifiable API-to-evaluation round-trip; and policy version-binding, meaning a content-policy revision is tied to a specific snapshot. Because each criterion is scored only against publicly observable artifacts such as documentation, Terms of Service, and standard API responses, the instrument measures external verifiability rather than internal practice. The proposed remedy, the Three-Part Behavioral Trigger System, is a second mechanism: capability triggers initiate full re-evaluation, drift triggers initiate documentation updates, and component triggers initiate logged disclosure.","core_discovery":"The central discovery is that the chain of custody between published safety evaluations and deployed systems breaks at the same points for every provider studied. Providers do disclose: eight of nine maintain changelogs, seven publish quantitative safety metrics, six publish per-version safety comparisons. Yet naming an artifact is not verification. Only Anthropic names a pinned, API-callable evaluated snapshot, and even there the external round-trip fails; no provider offers a public mechanism that maps an API response back to the evaluated artifact, and no provider binds a content policy to a specific snapshot. The authors describe the resulting state as \"transparency without verifiability\": documentation is abundant, but the primitives that would let regulators, auditors, or downstream users confirm that documentation describes the deployed system are missing.","pith_inferences":["An implication the paper leaves implicit is that the scorecard can be run repeatedly to produce longitudinal gap metrics, such as the share of deployed versions without changelog entries, that quantify whether documentation is keeping pace with deployment changes.","The chain-of-custody concept plausibly transfers to other hosted machine-learning settings where evaluations are published for served models, including image generation, embedding models, and open-weight deployers that modify releases.","A testable extension would be for providers to adopt content hashes like the one observed at one host: an auditor could record the hash of a served deployment and later verify bit-identity against the evaluated snapshot, turning the currently zero round-trip result into a measurable primitive.","If black-box fingerprinting matures enough to identify a served snapshot from output distributions alone, the paper's zero-round-trip finding may understate the verification available to outsiders, shifting policy attention from provider-published primitives to independent detection."],"forward_implications":["Regulators and downstream users currently cannot confirm that published safety evaluations describe the model being served, so any governance obligation that assumes such a link is unenforceable in practice.","Stable API identifiers can silently resolve to different snapshots, so pinning an alias does not pin behavior across time.","Contractual restrictions on benchmarking in six of nine providers compound the technical gap by limiting the independent testing that could detect silent changes.","Immutable snapshot binding is technically feasible, since at least one host exposes content-hashed deployments, so the absence elsewhere reflects disclosure choices rather than technical limits.","The proposed trigger system would convert vague \"material change\" duties into concrete obligations: capability changes trigger re-evaluation, behavioral drift triggers documentation updates, and component changes trigger logged disclosure."],"supporting_citations":[{"why":"Documents that a deployed API's behavior shifts over time, providing the empirical premise that silent updates matter.","marker":"Chen, Zaharia, and Zou 2024"},{"why":"Shows the stable identifier deepseek-chat resolving to many snapshots, the clearest documented alias-drift case.","marker":"DeepSeek AI 2024"},{"why":"Demonstrates that immutable model identifiers exist, making it the single near-match to snapshot binding.","marker":"Anthropic 2026b"},{"why":"The GPT-5 System Card serves as the worked example of evaluations naming family-level identifiers rather than pinned snapshots.","marker":"OpenAI 2025"},{"why":"Prior transparency measurement that flags post-deployment disclosure as weak and supplies indicators the scorecard adapts.","marker":"Bommasani et al. 2025"},{"why":"The GPAI Code of Practice whose \"substantial modification\" language the trigger system seeks to operationalize.","marker":"European Commission 2025"},{"why":"The EU AI Act creates the post-market monitoring and material-change obligations that assume a verifiable chain of custody.","marker":"European Parliament and Council of the European Union 2024"},{"why":"Documents benchmarking restrictions and motivates the proposed safe harbor for public-interest evaluation.","marker":"Longpre et al. 2024"}],"fun_headline_variants":["AI providers fail to link safety docs to deployed models","No provider proves its AI model matches its safety report","Silent updates break AI's chain of custody, audit finds","New scorecard exposes missing proof in AI safety claims","AI transparency gap: docs don't match deployed models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The zero-verification finding depends on defining \"externally verifiable\" as provider-published binding primitives and refusing to credit independent black-box fingerprinting, so if outsiders can verify a served snapshot by behavioral matching alone, the reported gap is smaller than claimed.","fun_headline_variants_meta":{"raw":{"variants":["AI providers fail to link safety docs to deployed models","No provider proves its AI model matches its safety report","Silent updates break AI's chain of custody, audit finds","New scorecard exposes missing proof in AI safety claims","AI transparency gap: docs don't match deployed models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000711,"raw_usage":{"total_tokens":3183,"prompt_tokens":909,"completion_tokens":2274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":2196}},"tokens_in":525,"tokens_out":2274,"duration_ms":15598,"temperature":1.0,"reasoning_tokens":2196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:26:32.477015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the chain-of-custody protocol on a provider scored zero: find a public safety evaluation that names an immutable, API-callable snapshot identifier, then confirm through public documentation and standard API calls that a response from that identifier corresponds to the evaluated artifact and that the content policy is version-bound; one passing provider would refute the universal zero claim. Alternatively, mount an independent black-box comparison that reliably matches an API's output distribution to a specific evaluated snapshot, which would show that an externally verifiable round-trip exists even without provider-published primitives.","supporting_citations":[{"cited_title":"2024 , url =","cited_arxiv_id":null,"evidence_quote":"Shows the stable identifier deepseek-chat resolving to many snapshots, the clearest documented alias-drift case."}],"review_version":1}