{"id":"e9cf7be4-fed4-4afc-8403-af1be95aea76","arxiv_id":"2608.07913","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A differentially private, federated, anytime-valid selective-risk certificate for RAG: bounds held in 500-trial audits, but registered targets r*=0.10 and often r*=0.20 never certified.","lead":"Fed-SRC is a new certificate that aims to guarantee, privately and across federated clients, that outputs accepted by a retrieval-augmented system stay below a declared error rate. It never violated its own bounds in extensive simulations, but its registered error targets almost never certified, so the guarantee is valid yet often operationally near-vacuous.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Q* guarantee rests on per-client TV bounds γ_k that are declared but never externally justified; without an independent shift audit the deployment certificate has an unverified premise.","rationale":"I reviewed the proof structure of Theorem 1 and the supporting lemmas carefully. The martingale arguments, Gaussian envelopes, and the range-one TV transfer algebra appear internally consistent: the simultaneous event, the predictable-schedule handling, and the optional-stopping argument all follow standard machinery, and I found no clear fatal flaw in the theorem's derivation. The genuinely load-bearing premise is the existence of per-client TV bounds γ_k that transfer validity from the realized calibration mixture to the declared deployment mixture. The paper is admirably explicit in its Limitations that no γ_k is justified by an external shift audit and that shifted deployment risk is not reported; the experiments use γ_k = 0 or declared sensitivity values constructed for the audit. This means the central deployment claim is conditional on an assumption that has not been independently verified. That is exactly the reader's weakest assumption, and I agree with it. The concrete test would close the gap: run an honest external shift audit, feed the resulting γ_k into the certificate, and observe whether the certificate remains useful and valid. The verdict CONDITIONAL remains appropriate, so no change is needed.","tokens_in":32338,"tokens_out":12397,"duration_ms":145032,"concrete_test":"Take one RAGTruth cell and construct a genuinely shifted deployment split (e.g., a later temporal batch or a held-out domain) for which per-client TV(P*_k, P_k) can be computed exactly or bounded by a finite-sample upper confidence interval from an independent audit sample. Run Fed-SRC (both H and VA) with γ_k set to these audit-derived bounds and record firing rate, selected acceptance, and held-out selective risk. Then rerun with γ_k halved and doubled to show sensitivity. If the audit-derived γ_k certificate still fires with held-out risk below target, the transfer premise is practically usable; if it abstains everywhere or the risk breaches target under honest bounds, the deployment guarantee is vacuous without external shift estimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The theorem's deployment clause depends on Lemma 6/9's transfer from the realized calibration mixture Qbar_t to the declared deployment mixture Q* via η_t = ½∑|w_k − wbar_{k,t}| + ∑ w_k γ_k. The first term is observed participation; the second is the per-client TV bound TV(P*_k, P_k) ≤ γ_k, which must be independently justified. The paper's Limitations state plainly that no γ_k is justified by an external shift audit and no shifted deployment risk is reported. In the experiments γ_k is either zero (deployment identical to calibration) or a declared sensitivity value under a hand-constructed coupling (Table 10), so the empirical 'no violations' audits do not exercise the premise. If the true deployment shift exceeds the declared γ_k, the inequality d_{Q*,j} ≤ dbar_{j,t} can fail even though the frozen-population audit passes; the certificate would overstate safety. This is not an internal inconsistency in the martingale construction, which reads as sound, but it makes the central claim conditional on an unverified quantity that the paper itself identifies as missing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops Fed-SRC, a score-agnostic certificate for federated, differentially private, adaptively monitored selective-risk control of retrieval-augmented generation. Clients release Gaussian-perturbed score/loss histograms, and the server uses record-indexed and variance-indexed martingales to bound, simultaneously over all registered thresholds and rounds, the target-risk contrast and the accepted mass. Theorem 1 (Appendix M) is the central claim: under a predictable-stream assumption and per-client total-variation bounds TV(P*_k, P_k) ≤ γ_k, with probability at least 1 − α_s − α_n the declared-mixture risk contrast and acceptance are bounded for every threshold and round, so transcript-measurable threshold selection and optional stopping need no further correction. A pathwise zCDP accountant (Theorem 3) composes the Gaussian releases. Empirically, the paper audits the machinery with 500 trials per core cell (no simultaneous-bound violations), reports that the registered targets r* = 0.10/0.20 mostly abstain, and reports positive HaluEval QA results at r* = 0.20 without privacy and a full-support C1 result at r* = 0.30. The paper is candid about limitations: event-level privacy only, high record reuse, an unproved capital heuristic, and, critically, that the deployment transfer radii γ_k are never externally justified.","tokens_in":32452,"tokens_out":25404,"duration_ms":268756,"significance":"If the results stand, the contribution is a carefully engineered and honestly evaluated combination of known statistical tools (Hoeffding, Ville, Freedman, Gaussian-martingale bounds, zCDP, total-variation transfer) into an event-level private federated anytime certificate. The strengths are real: the proofs are detailed and follow standard machinery; the artifact contract (checksummed manifest with a single regeneration entry point) is exemplary; and negative results are reported rather than hidden (registered targets abstain; held-out intervals are not certified; the capital heuristic is unproved). The paper does not claim a superior hallucination score and explicitly disclaims the earlier lift-centered claims. The main value is as a reference architecture and a cautionary empirical study of where such certificates are not operationally useful under privacy. Because the deployment guarantee is conditional on externally justified TV bounds that the paper does not supply, the transfer mechanism is validated as arithmetic but not as a deployment guarantee; this limits but does not destroy the contribution.","major_comments":[{"comment":"The certificate's transfer to the declared deployment mixture Q* rests on per-client total-variation bounds TV(P*_k, P_k) ≤ γ_k that are nowhere externally justified. The theorem is explicitly conditional, but the abstract and Section 1 present the transfer as a contribution without flagging that the empirical audits (Tables 7 and 10) exercise only γ_k = 0 or hand-declared radii under a coupling constructed by the authors, and the Limitations concede that 'no γ_k is justified by an external shift audit and we do not report shifted deployment risk.' If a real deployment shift exceeds the declared γ_k, the inequality d_{Q*,j} ≤ dbar^M_{j,t} can fail even though the frozen-population audit passes. The paper should (i) state the conditionality of the Q* guarantee in the abstract and Section 3.2, (ii) label the transfer experiments as sensitivity analyses in the main text as the Limitations already do, and (iii) discuss a concrete route to justifying γ_k, such as a separate drift-audit sample yielding a confidence bound on TV, or an argument for why such bounds are obtainable in the intended deployments.","section":"Abstract and Section 7.1 / Table 4"},{"comment":"The abstract's statement that on HaluEval question answering the certificate certifies 'with held-out risk below the target' overstates the evidence: the paper's own Section 7.1 says held-out risk below target 'is a point-estimate claim, not a certified bound,' and Table 4's exact 95% Clopper-Pearson intervals have upper endpoints (0.241 at ε = ∞, 0.204 at ε = 8 and ε = 4) above r* = 0.20 even in the non-private row. The abstract should say 'held-out point-estimate risk below the target,' and the main text should state at the first occurrence of the claim that the selective-risk bound is not certified on the held-out set.","section":"Abstract / Section 7.1 / Table 4"}],"minor_comments":[{"comment":"The HaluEval guarantee is conditional on the frozen calibration support, as Proposition 14 explains; the Table 4 caption and the surrounding text should explicitly say 'conditional on the frozen calibration support' whenever the support-derived quantile grid is used, so that a reader cannot mistake it for a marginal guarantee.","section":"Section 7.1 / Table 4 / Proposition 14"},{"comment":"The displayed definition of H_c(n; α) renders as '2⌈log2 n⌉', which appears to mean 2^{⌈log2 n⌉}; this should be aligned with the proof's b_r = sqrt(U_r x_r / 2) in Lemma 10 so that the boundary formula matches the epoch argument.","section":"Equation (19)"},{"comment":"The embedded provenance notes ('withdrawn,' 'superseded as a headline,' 'disabled block above is retained only as a provenance record') are unusual for a journal submission; superseded material should be moved to clearly marked supplementary text so that the main narrative is not interrupted.","section":"Appendices F and G"},{"comment":"The annotation 'the displayed 6.00 is 10^3 d' is confusing; writing d = 0.00600 with an explicit standard-form note would be clearer.","section":"Table 13"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest, and I found no internal inconsistency in the martingale constructions. The main risk is scope rather than correctness: the deployment guarantee depends on per-client TV bounds that the paper explicitly admits are not externally justified, and the abstract's headline empirical claim ('held-out risk below the target') is contradicted by the paper's own careful interval language. A revision that re-scopes the abstract, labels the transfer experiments as sensitivity analyses in the main text, and adds a concrete discussion of how γ_k could be obtained would make the paper publishable in a strong security/privacy venue. The self-citations to the authors' information-lift score are not circular because that score is shown to fail on RAGTruth; the negative result actually mitigates self-promotion concerns. The paper's fit for the journal is good if the scope is adjusted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper is that it is a genuinely serious contribution draped in unusually honest packaging. The core construction — federated histograms with event-level zCDP, record- and noise-variance-indexed martingales, and a range-one TV transfer to a declared deployment mixture — is assembled from known parts, but the joint anytime-valid deployment contract is new, and the proofs are detailed and follow standard Hoeffding, Ville, Freedman, and zCDP composition arguments. I read through the main theorems and the appendices and found no fatal flaw. The empirical work is also refreshingly candid: the primary target never certifies, the secondary rarely does, and the paper explicitly retracts earlier overclaims. That is exactly how empirical negative results should be reported.\n\nThe main soft spot is the one the stress-test note flags, and I think it lands. Theorem 1's transfer to Q* depends on independently justified per-client total-variation bounds TV(P*_k, P_k) ≤ γ_k. In the experiments those γ_k are either zero or declared sensitivity values under a hand-constructed coupling. The paper itself admits no γ_k is justified by an external shift audit and that shifted deployment risk is not reported. So the deployment certificate has an unverified premise: if the real shift exceeds the declared γ_k, the bound can fail even though the frozen-population audit passes. This is not an internal inconsistency, but it means the headline guarantee is conditional on a quantity the paper does not supply. The HaluEval threshold grid is another honest soft spot: it is not fixed in advance, and the salvage via Proposition 14 gives only conditional validity on the frozen support. Also, the paper contains no proved private comparator; the capital heuristic is explicitly unproved, which limits the empirical comparison.\n\nNone of this sinks the paper. The central theorem is likely sound, the limitations section is more complete than most published papers' entire discussion, and the near-vacuous empirical power at registered targets is reported as a finding rather than hidden. The gap is between the theorem and its deployable form. A serious referee should push the authors to provide an external shift audit or an honest statement that deployment requires independently justified γ_k, and ideally a proved private baseline.\n\nI would send this to peer review. It deserves referee time, and the authors are clearly capable of responding to it. I'd bring it to a reading group only if the group cares about privacy-preserving selective prediction; it is somewhat niche, but the honesty of the empirical report is worth studying in its own right.","headline":"Honest, carefully built private anytime selective-risk certificate; the math is sound and the empirical limits are reported straight, but the deployment guarantee rests on per-client TV bounds that are declared rather than externally justified.","tokens_in":33051,"tokens_out":1704,"would_cite":true,"duration_ms":23203,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62L10","68T50"],"pacs":[],"model":"deepseek-v4-flash","headline":"A differentially private federated protocol can certify that the accepted outputs of a retrieval-augmented generator keep their expected error below a declared target, valid at every monitored round and threshold.","keywords":["selective risk certification","federated learning","differential privacy","retrieval-augmented generation","anytime-valid inference","martingale confidence sequences","hallucination detection","mixture transfer"],"falsifier":"Take a real deployment whose per-client shift is measured by an independent audit, run Fed-SRC with declared $\\gamma_k$ below the measured total variation, and check whether any trial both fires with $\\bar{d} \\le 0$ and $\\underline{a} \\ge a_{\\min}$ while exhaustive population risk exceeds $r^\\star$; the paper reports no such event, but its transfer experiments use claimed radii, not externally audited ones.","tokens_in":32062,"feed_emoji":"🔐","tokens_out":7595,"duration_ms":67085,"temperature":0.7,"pith_summary":"This paper claims that a differentially private, federated calibration protocol can certify that the expected error among accepted retrieval-augmented generation outputs stays below a declared target, and can keep that promise valid no matter when the operator stops or which threshold it picks. The certificate works by having each client release only Gaussian-noised histograms of scores and losses, then combining two types of anytime-valid martingale bounds: one that controls sampling error over the record stream and one that controls the privacy noise accumulated over rounds. A single total-variation term, costing at most one unit of probability mass, transfers the guarantee from the realized mixture of client data to the mixture the operator actually wants to deploy on. If the paper is right, selective-risk certification for RAG outputs can be made privacy-preserving and federated without sacrificing the anytime-valid, post-selection-invariant character of the guarantee. The authors' own experiments, however, show sharp limits: the registered tight targets mostly abstain, and only a looser exploratory target fires under common privacy budgets.","feed_headline":"Federated private audit certifies RAG risk at every stop","feed_subtitle":"Noised score histograms per client yield anytime-valid risk bounds, though tight targets mostly abstain.","key_machinery":"The machinery is a pair of simultaneous confidence envelopes built from martingales indexed differently: one record-indexed process that recenters at each chosen client's law under predictable recruitment, and one variance-indexed Gaussian process that charges only releases in which a given record participates, with the total-variation transfer term $\\eta_t$ carrying the declared deployment-mismatch cost. Two constructions are offered, a range-only one with Hoeffding-style record epochs and a variance-adaptive one using a Freedman bound on the contrast's predictable quadratic variation; both are stitched over dyadic epochs to stay simultaneous in time and threshold.","core_discovery":"The central claim, stated as Theorem 1, is that under predictable calibration streams and declared conditional laws, with probability at least $1-\\alpha_s-\\alpha_n$ the true target-risk contrast $d_{Q^\\star,j}$ is no greater than the certificate's upper bound $\\bar{d}^M_{j,t}$ and the true accepted mass $a_{Q^\\star,j}$ is no less than its lower bound $\\underline{a}^M_{j,t}$, simultaneously for every registered threshold and every monitored round. Consequently any transcript-measurable threshold and stopping time that satisfy $\\bar{d}^M_{j,t} \\le 0$ and $\\underline{a}^M_{j,t} \\ge a_{\\min}$ certify that the deployment selected risk $R_{Q^\\star,j} \\le r^\\star$ with no additional correction. The proof splits the error probability into a sampling budget and a Gaussian-noise budget, and transfers the calibration mixture to the declared deployment mixture through a range-one total-variation term that costs $\\eta_t$ for both numerator and denominator rather than $(1+r^\\star)\\eta_t$. The empirical companion claim is that across 500-trial audits over cells, privacy levels, and policies, no simultaneous-bound violation occurred, but the registered target $r^\\star=0.10$ never certified and $r^\\star=0.20$ certified only on HaluEval question answering without privacy.","pith_inferences":["The dependence on declared $\\gamma_k$ radii suggests a practical enhancement: pair Fed-SRC with an independent shift-audit stage that measures per-client total variation on a held-out sample and feeds audited radii into the certificate; the paper's experiments stop short of this.","The 30-to-78-fold event reuse means event-level privacy is weak protection for items or people that recur; a source- or person-level deployment would require bounded contribution and a recalibrated accountant, which the current guarantee does not provide.","The near-universal abstention at registered targets suggests the certificate's operational role may be to quickly retire a score/population combination rather than to certify often; this supports the paper's score-agnostic design, since the winning score is not knowable in advance.","If the range-one transfer is as tight as claimed, the same certificate could be extended to subgroup-conditional risk by pre-registering subgroup histograms and splitting the error and privacy budgets, at additional sample and communication cost."],"forward_implications":["An operator following the predictable-schedule rule can adaptively recruit clients, select thresholds, and stop without invalidating the certificate, because the simultaneous event covers all such choices in advance.","Privacy is event-level: each stream occurrence pays only for releases containing it, and the pathwise accountant converts the zCDP budget to $(\\varepsilon,\\delta)$-DP under adaptive composition.","The mixture-transfer term makes the deployment claim explicit: a mismatch between realized participation and declared deployment weights is priced into both acceptance and risk, so a large shift automatically forces abstention.","The empirical audit shows validity is not the bottleneck; score quality and privacy budget are. Tight targets ($r^\\star=0.10$) never fire, and at $\\varepsilon=4$ the exploratory target fires only 1% of trials on the information-lift score.","Naive privatization without the noise envelope breaks the bounds in 146 to 198 of 200 trials, demonstrating that the Gaussian envelope is load-bearing for validity."],"supporting_citations":[{"why":"Supplies the time-uniform confidence-sequence framework that the record-indexed martingales rely on.","marker":"Howard et al., 2021"},{"why":"Provides the gated target-risk contrast and anytime selective-risk control that Fed-SRC extends to privacy and federated settings.","marker":"Khosravi and Huo, 2026"},{"why":"Provides the joint variance-adaptive risk/acceptance certificate behind one of the two bound constructions.","marker":"Yu and Liu, 2026"},{"why":"Supplies the zCDP definition and Gaussian-mechanism composition used for the pathwise privacy accountant.","marker":"Bun and Steinke, 2016"},{"why":"Establishes the differential-privacy baseline and the $(\\varepsilon,\\delta)$ conversion context.","marker":"Dwork and Roth, 2014"},{"why":"Motivates per-observation total-variation transfer instead of per-block BK bounds, supporting the mixture-transfer term.","marker":"Barber et al., 2023"},{"why":"Demonstrates that privacy and anytime validity can coexist, background for the private confidence-sequence design.","marker":"Waudby-Smith et al., 2023"},{"why":"Provides the private e-value mechanism that inspires the exploratory capital heuristic, which the paper does not prove valid.","marker":"Csillag and Mesquita, 2025"}],"fun_headline_variants":["Private federated RAG certificates hold at every stop","Federated private audits keep RAG risk bounds valid anytime","RAG risk certificates survive privacy, but tight targets refuse to certify","No violations in federated private RAG audits, yet strict targets abstain","Anytime-valid risk bounds for RAG, under federation and privacy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The certificate's deployment guarantee depends on independently justified per-client total-variation bounds $\\mathrm{TV}(P^\\star_k, P_k) \\le \\gamma_k$, which in the experiments are declared by the authors rather than fixed by an external shift audit; if real deployment drift exceeds the declared radii, the certificate can fire while the true deployment risk exceeds the target.","fun_headline_variants_meta":{"raw":{"variants":["Private federated RAG certificates hold at every stop","Federated private audits keep RAG risk bounds valid anytime","RAG risk certificates survive privacy, but tight targets refuse to certify","No violations in federated private RAG audits, yet strict targets abstain","Anytime-valid risk bounds for RAG, under federation and privacy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1570,"prompt_tokens":1113,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":729,"completion_tokens_details":{"reasoning_tokens":366}},"tokens_in":729,"tokens_out":457,"duration_ms":5611,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:41:47.328609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real deployment whose per-client shift is measured by an independent audit, run Fed-SRC with declared $\\gamma_k$ below the measured total variation, and check whether any trial both fires with $\\bar{d} \\le 0$ and $\\underline{a} \\ge a_{\\min}$ while exhaustive population risk exceeds $r^\\star$; the paper reports no such event, but its transfer experiments use claimed radii, not externally audited ones.","supporting_citations":[],"review_version":1}