{"id":"f8d007e4-176a-41ac-b81e-bf969d6954f9","arxiv_id":"2607.20860","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Random-generation probes plus a pilot-fitted budget let a text-only auditor detect model substitution, estimate the routing dilution fraction, and attribute the served backend across LLM gateways.","lead":"IRIS is a text-only audit tool that asks an LLM gateway to generate random numbers and strings, then fingerprints the served model from the returned text alone, detecting both whole-stream substitution and partial routing dilution and estimating the dilution fraction. It matters because clients of commercial LLM gateways currently cannot easily verify which backend actually serves their requests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-sized budget/per-pair FPR not calibration-robust: App. S5: 0.9 power only 17.8% feasible pairs at prescribed budget, FPR max 0.39 post-hardening; App. S4: 1000-sample calibration overshoots FPR. Abstract 'sizes itself' and 0.017 FPR overstate reliability.","rationale":"I read the paper in good faith: IRIS is a genuinely engineered, unusually self-aware audit system, with complete proofs, verbatim probes, a full feature dictionary, and a claim-by-claim guarantee table. The detection signal — random-generation fingerprints separating models — is strongly supported by the local ladder, 17- and 45-model gateway pools, real provider-pair deviations, and MET corroboration; the theory in Prop. 1, Thm. 1, and Prop. 6 is coherent under its stated assumptions. The load-bearing weakness is not the existence of the signal but the reliability contract attached to the headline numbers. For the self-sized budget claim to hold, the calibration window must be large/representative enough that the frozen thresholds control per-audit type-I and the pre-committed m⋆ delivers the promised power. The paper's own App. S5/S4 provide direct evidence that this condition fails on a nontrivial fraction of pairs: power at the prescribed budget meets the target only 17.8% of the time, and per-pair FPR carries a heavy tail even after hardening. I therefore agree with the reader's weakest_assumption, and the concrete test I propose would distinguish 'finite calibration size' from 'structural failure of the reliability contract.' The verdict should remain CONDITIONAL: the core detection/attribution/estimation results are likely sound and well-evidenced, but the abstract's wording of the self-sizing budget and pooled FPR should be tightened to match the paper's own 'diagnostic, not coverage guarantee' caveat.","tokens_in":68165,"tokens_out":6188,"duration_ms":66170,"concrete_test":"Using the released bundle, re-run App. S5's end-to-end budget sizing on the same 258 feasible pairs (content-only c10,8; α=0.05, δ=0.1, ε=0.3) with the honest calibration window resampled from the existing reference bank at N=1000 and N=3000 per pair instead of ~33, keeping train/calibration/audit splits disjoint and retaining the Bonferroni-split Clopper–Pearson bounds. Record (a) the per-pair FPR distribution and (b) the fraction of pairs whose realized power at the sized m⋆ is ≥0.9. If the FPR tail collapses (max ≤0.05, no substantial exceedance fraction) and the power fraction rises toward 0.8–0.9 at N=3000, the shortfall is finite-calibration size and the current caveat suffices; if the tail persists even at N=3000, per-audit type-I control is structurally unavailable and the headline FPR/self-budget claims must be explicitly downgraded to diagnostics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central novelty claim — 'a query budget it sizes itself' plus the 0.017 false-positive-rate headline — requires the honest reference calibration window (τ, α₁, q, pilot exponent) to be representative enough that frozen thresholds control per-audit type-I error and the pre-committed budget m⋆ meets the power target. The paper's own Appendix S5 shows this condition fails on a substantial minority: each pair calibrates on only ~33 responses, realized power meets the pre-registered 0.9 target on only 17.8% of feasible pairs at the prescribed budget (12.6% at ε=0.2), and per-pair FPR, though pooled mean 0.013, has max 0.475 before the Bonferroni split and 0.39 residual after hardening. Appendix S4 independently shows a 1000-sample calibration overshoots realized total FPR to 0.09–0.14 on c2,16, and the certifiable budget is only m≤16 at N=1000, m≤50 at N=3000 for α=0.05. Thus the 'representative calibration window' assumption is not a technicality: it is the load-bearing condition for the self-sized budget to function as a reliability contract. The paper honestly labels the prescribed budget a 'feasibility diagnostic, not a coverage guarantee' (Table 3) and reports the FPR tail explicitly, so the flaw is disclosed; but the abstract states the self-sizing capability and pooled FPR without that caveat. This is an internal-consistency/scope concern about the strength of the central claim, not about whether the fingerprint signal exists.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents IRIS, an output-only (visible-string) black-box auditor for LLM gateways. It uses random-generation probes, a 179-dimension visible-string feature vector, and a random-forest posterior to detect whole-stream substitution and fractional dilution, attribute the served backend, estimate the routing fraction ε, and pre-commit a query budget estimated from a cheap pilot. The theoretical core (Section 5, App. S28) proves exponential decay of mean-score verification error, a response-level tell-rate characterization of the dilution budget, and a Θ(ε^{-2}) all-test lower bound under a χ²-regularity condition, with an interpolating tail-exponent law (Prop. 6). Experiments on a local Qwen3 ladder, a 17- and 45-endpoint OpenRouter library, and a live cross-provider audit report strong detection, attribution, and ε-recovery numbers, supported by a released code/data bundle. The appendices are unusually candid about finite-sample limitations, including per-pair FPR tails, under-coverage of the δ-method interval, and the fact that the prescribed budget meets the power target on only a minority of pairs at the calibrated sample size.","tokens_in":68465,"tokens_out":6038,"duration_ms":57971,"significance":"If the claims hold, IRIS is a substantial advance: a text-only auditor that detects dilution, names the diluent, estimates ε, and predicts its own query cost before issuing suspect queries would close a real gap in gateway auditing. The paper's strengths are concrete: machine-checkable-style complete proofs of the main propositions in App. S28; a large frozen response dataset and code release; head-to-head comparisons against FLIPS, MET, and RUT on a shared probe; and a clear identification of the 1/ε versus 1/ε² regimes as a measured tail-exponent phenomenon rather than an unconditional law. The paper also deserves credit for explicitly labeling its own reliability limits in App. S5 (Table 3: the prescribed budget is a 'feasibility diagnostic, not a coverage guarantee') and for reporting FPR tails and held-out margin-selection checks rather than only pooled means. However, the central abstract claims—'a query budget it sizes itself' and a '0.017 false-positive rate'—are stronger than what the paper's own finite-calibration results support, and this is load-bearing for the paper's main novelty.","major_comments":[{"comment":"The central reliability claims are not supported as stated. App. S5's end-to-end check at (α=0.05, δ=0.1, ε=0.3) shows realized power averaging 0.73 but meeting the pre-registered 0.9 target on only 17.8% of feasible pairs (12.6% at ε=0.2), and per-pair type-I FPR with mean 0.013 but maximum 0.475 before hardening and 0.39 after exact-binomial hardening. Table 3 of the appendix explicitly labels the prescribed budget a 'feasibility diagnostic, not a coverage guarantee.' The abstract's '0.017 false-positive rate' is the pooled mean, not a per-audit guarantee, and 'a query budget it sizes itself' is not a reliability contract. Revision must either provide a calibration-window size/composition rule that restores the (α, δ) guarantee at the stated level, or re-scope the claim to 'feasibility diagnostic' and report the per-pair FPR distribution rather than the pooled mean.","section":"Abstract; §1 Contributions; App. S5"},{"comment":"The estimate-then-budget pilot is not calibration-robust. App. S4 shows that a 1000-sample reference calibration overshoots the realized total FPR to 0.09–0.14 on c2,16 (≤0.04 on c10,8), and the certifiable budget is only m≤16 at N=1000 and m≤50 at N=3000 for α=0.05. Since the deployed pipeline calibrates on roughly 33 held-out responses per pair (App. S5), the m⋆ produced by Eqs. (7) and (12) does not in general meet the stated (α, δ) targets. The paper should state the honest-window size required to certify a given m⋆ and qualify the 1/ε law as finite-ε and calibration-bounded, exactly as App. S4 already does.","section":"§4.1 Eq. (7), Eq. (12); App. S4"},{"comment":"The headline detection figure is conditioned on margin-qualified pairs in a way not fully transparent in the body. App. S3 reports that over the 218 margin-qualified pairs the per-pair honest FPR has median 0.000, 95th percentile 0.090, maximum 0.557 (0.39 after exact-binomial hardening), with 7.8% of pairs above 0.05; the body reports only the pooled mean 0.017. The held-out margin-selection power drops to 0.66 vs. 0.85 in-sample; the matched control attributes this to bank size, but it still shows sensitivity to the calibration split. The body should present the FPR distribution, the overall reliable-detection rate (≈0.51 over all 272 ordered pairs), and the caveat that a substantial minority of pairs are returned as indeterminate/low-margin rather than flagged.","section":"§6.2; App. S3"}],"minor_comments":[{"comment":"Several works cited in the appendices are missing from the reference list: Bruckner (2026), Richter et al. (2025), Shekhar and Ramdas (2023), Dima et al. (2025), Huber (1964), Donoho and Jin (2004), and Ong et al. (2025). Please complete the bibliography.","section":"App. S15, S26, S25"},{"comment":"The phrase 'first to combine' should be slightly moderated in the main text. The paper itself identifies Bruckner (2026) as concurrent single-token fingerprinting and FLIPS as sharing the random-generation signal; the defensible novelty is the specific combination (reusable probe, budget estimation, dilution ε-estimation, attribution) rather than the underlying signal.","section":"§1, App. S15"},{"comment":"Table 3's 'which statistic backs which claim, and at what guarantee level' is an excellent self-assessment. Consider moving a condensed version of it into the main text so that the abstract's guarantees are immediately qualified.","section":"App. S5, Table 3"},{"comment":"The δ-method interval for ε is mildly anti-conservative (empirical coverage 0.86–0.87 at nominal 90%, 0.92–0.93 at 95%). The paper describes this accurately as 'slightly optimistic,' but the main text should at least footnote this so readers do not treat the 90% interval as exact.","section":"App. S2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is thorough and unusually self-aware; the appendices disclose the calibration fragility that the abstract obscures. The core signal—random-generation fingerprinting, the exponential accumulation proof, and the measured 1/ε vs 1/ε² boundary with tail-exponent diagnostics—is valuable and reproducible, and the release of frozen responses and code is a real asset. My concern is that the abstract and Section 1 sell the self-sized budget and pooled FPR as guarantees when the paper's own data show the budget meets the power target on only a minority of pairs at the calibrated sample size and the per-pair FPR has a heavy tail. This is fixable by re-scoping the claims, reporting per-pair distributions, and adding honest-window size guidance; it is not a reason to reject the underlying method. I also noticed several appendix citations missing from the reference list."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, IRIS genuinely delivers the hard part: it detects fractional model dilution and names the substitute from visible strings alone, on real OpenRouter endpoints, with math that checks out. Second, the paper's headline reliability numbers—the self-sized budget and the 0.017 false-positive rate—are conditional on a calibration window that is often too small to deliver the guarantee. The authors know this; the appendix says so. The abstract does not.\n\nWhat's new: the combination of random-generation fingerprinting (prior art, credited) with an estimate-then-budget loop, dilution tell-count, and epsilon estimation in one text-only audit. The paper is unusually honest: complete proofs in S28, verbatim probes, feature dictionary, and Table 3 mapping each claim to its actual guarantee level. Prop 1, Thm 1, and Prop 6 are rigorous, and the theory correctly describes the 1/epsilon vs 1/epsilon^2 regimes as a measured tail phenomenon.\n\nThe soft spots are all about the reliability contract. On the paper's own numbers, the prescribed budget hits the pre-registered 0.9 power target on only 17.8% of feasible pairs. The per-pair false-positive rate has a heavy tail—max 0.475 before hardening, 0.39 after—so the pooled 0.017 FPR flatters the worst pairs. And the headline detection stats apply to the margin-qualified subset; the all-pairs rate at that power is about 0.51. The root cause is calibration size: ~33 responses per pair, and even a 1000-sample calibration misses the nominal FPR on one probe. None of this undermines the core signal—IRIS almost certainly detects dilution on distinguishable pairs from text alone—but it means the budget is a diagnostic, not a guarantee. The abstract should say that.\n\nThis paper is for ML-security researchers working on model provenance and gateway auditing. It deserves a serious referee; the empirical work is substantial and the self-disclosures are a model for the field. Send it to peer review, with a request to recalibrate the abstract and move the caveats up.","headline":"IRIS nails the text-only dilution-detection task on well-separated pairs, but the abstract oversells the self-sizing budget and the pooled false-positive rate; the authors themselves disclose the gaps in the appendix.","tokens_in":69100,"tokens_out":3041,"would_cite":true,"duration_ms":29625,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A text-only audit can detect when an LLM gateway serves a cheaper model, name the substitute, and estimate the diluted fraction before spending a query budget it sizes itself.","keywords":["LLM gateway auditing","model substitution","routing dilution","black-box fingerprinting","random-generation probes","query budget estimation","visible-string features"],"falsifier":"A concrete observation that would settle the central claim: take a fresh honest endpoint of the same model not used in calibration, run the frozen IRIS plan (same probe, same budget, same thresholds), and measure the realized false-positive rate across many independent audit runs. If the per-pair FPR exceeds the claimed 0.017 (or the pre-registered alpha) on a material fraction of pairs, or if the realized power at the pre-committed budget falls below the target on a substantial minority of margin-qualified pairs, then the calibration-based budget guarantee fails. The paper itself reports such","tokens_in":67850,"feed_emoji":"🔍","tokens_out":2196,"duration_ms":25633,"temperature":0.7,"pith_summary":"The paper introduces IRIS, an audit that uses only the returned text of random-generation probes to check whether an LLM gateway is truly serving the advertised model. It claims to be the first method that combines, in one text-only audit, detection of whole-stream substitution and fractional dilution, attribution of the served backend, estimation of the routing fraction, and a query budget it sizes itself from a cheap pilot. If correct, a client could verify model identity and quantify dilution without privileged signals like log-probabilities or token ranks, and without fixing the query count in advance. The central validation is that on a commercial OpenRouter library, IRIS catches epsilon=0.3 dilution on margin-qualified pairs at 0.85 mean power with a 0.017 false-positive rate, and recovers the dilution rate to within 0.04 for enrolled diluents.","feed_headline":"Text-only audit catches diluted LLM routing at 0.85 power","feed_subtitle":"IRIS fingerprints backends from random strings, names the substitute, and sizes its own query budget before auditing.","key_machinery":"The central object is the random-generation probe family c_{n,L} (e.g., c_{2,16} for 16 random bits, c_{10,8} for 8 random digits) paired with a 179-dimensional visible-string feature vector and a multiclass classifier's negative log-posterior score. The load-bearing mechanism is the 'estimate-then-budget' pilot: a cheap labeled pilot (m <= 4 queries) fits the exponential rank-error decay log(1-AUROC) ~ const - I_auc * m, and the fitted exponent (with a lower-confidence bound) is used to freeze the live-query budget and select the probe before any suspect traffic is queried. For dilution, response-level 'tell' rates (reference-atypical responses) are calibrated on a separate split, and the t","core_discovery":"The paper's central claim is that visible-string biases in responses to random-generation prompts—such as asking for digits, bits, or short sequences—are backend-specific fingerprints that can be used to audit LLM gateways. IRIS turns each response into a 179-dimensional feature vector, trains a multiclass classifier on enrolled endpoints, and scores responses by the reference negative log-posterior. It then calibrates telling thresholds and, critically, estimates the query budget before any suspect query is issued: a pilot fits the exponential decay of rank error (1 - AUROC) and the dilution tell-rate regime, yielding a pre-committed budget. The paper argues that evidence accumulates across","pith_inferences":["If IRIS's calibration guarantees hold in deployment, a practical consequence is that gateway providers can be held accountable for silent downgrades: a customer could routinely audit a fraction of their traffic and obtain evidence that stands up enough to contest billing, even without provider cooperation.","The paper's distinction between a pure temperature retune (second-order, near-invisible) and a base-model swap (first-order, detectable) suggests that the audit's 'blind spot' is intentionally aligned with economic incentives: tempering alone does not create extra gateway profit under token-count billing, so the audit focuses on the substitution threat that matters financially.","A testable extension is to turn the audit into an anytime-valid sequential test using the same tell-count machinery; the paper's own appendix compares a betting e-process and SPRT, so a deployment could stop early on clearly honest or clearly diluted streams while preserving the pre-committed budget contract.","The measured tail exponent kappa is reported as a pair-specific property; if a gateway adaptively switches between multiple diluents or uses sticky routing, the i.i.d. tell-count interval degrades, and the paper's block-bootstrap repair suggests that per-session or longer audit horizons would be needed to maintain calibrated uncertainty on the routing fraction."],"forward_implications":["A client can audit an LLM gateway unilaterally from returned text alone, with no privileged access, and with a budget that is estimated before the audit begins.","The audit can not only decide whether dilution is happening but also name the substitute and estimate the routing fraction, making the result actionable for billing disputes or price renegotiation.","The measured 1/epsilon budget law means that even low-rate dilution (e.g., epsilon=0.1) is detectable within a realistic number of queries when a separating tail exists, and the paper's tail-exponent measurement can decide which pairs fall into the cheap 1/epsilon regime versus the expensive 1/epsilon^2 wall.","Because the probes are reusable and task-agnostic, the same audit pipeline transfers across endpoints without target-specific probe design, unlike prior knowledge-boundary or border-input audits.","A real cross-provider audit of the same nominal model served by different providers found 14 of 15 pairs distinguishable, indicating that the method can flag genuine quantization/kernel deviations in the wild."],"fun_headline_variants":["Random-number prompts fingerprint LLM backends","IRIS: text-only audit detects hidden model swaps","Catch diluted LLM routing with random-string probes","Self-sizing audit names the substitute backend","Text-only fingerprints unmask gateway substitution"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The honest reference calibration window is representative and large enough that the thresholds and budgets frozen on it control type-I error at the stated level and meet the power target on a new audit stream.","fun_headline_variants_meta":{"raw":{"variants":["Random-number prompts fingerprint LLM backends","IRIS: text-only audit detects hidden model swaps","Catch diluted LLM routing with random-string probes","Self-sizing audit names the substitute backend","Text-only fingerprints unmask gateway substitution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1281,"prompt_tokens":887,"completion_tokens":394,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":326}},"tokens_in":631,"tokens_out":394,"duration_ms":10910,"temperature":1.0,"reasoning_tokens":326,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:07:46.295132+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete observation that would settle the central claim: take a fresh honest endpoint of the same model not used in calibration, run the frozen IRIS plan (same probe, same budget, same thresholds), and measure the realized false-positive rate across many independent audit runs. If the per-pair FPR exceeds the claimed 0.017 (or the pre-registered alpha) on a material fraction of pairs, or if the realized power at the pre-committed budget falls below the target on a substantial minority of margin-qualified pairs, then the calibration-based budget guarantee fails. The paper itself reports such","supporting_citations":[],"review_version":1}