{"id":"ded59d9e-c02b-4ee8-8243-a760d7b190f6","arxiv_id":"2511.08307","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"For bounded-variance response distributions, the DKPS embedding error is O_P((n^3/r)^{1/2-delta}) when r grows faster than n^3.","lead":"This paper derives mathematical formulas for how many sample responses are needed to accurately map black-box generative models (like LLMs) into fixed numerical vectors using the DKPS embedding. If valid, it would let practitioners choose replicate counts to compare models with guaranteed accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the reader's Theorem 1 concern rests on a misreading; central rate likely holds.","rationale":"The reader's REJECT is based on a specific claimed gap in Theorem 1. A careful reading of the definitions shows the gap does not exist: the dissimilarity is the square root of a Frobenius norm, so D^2−Δ^2 is a first-order difference of norms, not a difference of squared norms. The reverse-triangle step is valid. The Chebyshev denominator is indeed m^2 rather than m, but because this makes the paper's bound looser (larger failure probability), the theorem statement remains true as an upper bound on failure probability. Since Corollary 1 and Theorem 2 inherit only this looser bound, the main rate is unaffected. I therefore disagree with the reader's assessment. However, the manuscript has several minor typos (m vs m^2, circular references to Corollary 2 in Proposition A.4 and Lemma 2, 'From Theorem 2' instead of the decomposition reference), and the proof of Theorem 2 depends on an external decomposition not fully reproduced. These do not invalidate the central claim but justify a conditional acceptance pending proof-reading and clarification rather than outright rejection.","tokens_in":16144,"tokens_out":39326,"duration_ms":320419,"concrete_test":"Re-derive the proof of Theorem 1 using the paper's definition D_ii' = m^{-1/2}||Xbar_i-Xbar_i'||_F^{1/2}. Verify that |D^2_ii'-Δ^2_ii'| ≤ (1/m)(||Xbar_i-μ_i||_F + ||Xbar_i'-μ_i'||_F) and that Chebyshev gives P((1/m)||Xbar_i-μ_i|| ≥ ε/4) ≤ 16∑_j γ_ij/(r m^2 ε^2). If both hold, the reader's cross-term objection disappears and the m-denominator is a benign typo.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The reader's weakest-assumption critique is not sustained by the manuscript's definitions. The paper defines the dissimilarity as D_ii' = m^{-1/2} ||Xbar_i - Xbar_i'||_F^{1/2}, so D^2_ii' = (1/m)||Xbar_i - Xbar_i'||_F and similarly for Delta^2. Thus E_ii' = D^2_ii' - Delta^2_ii' = (1/m)(||Xbar_i-Xbar_i'||_F - ||mu_i-mu_i'||_F). The reverse triangle inequality gives |E_ii'| ≤ (1/m)(||Xbar_i-mu_i||_F + ||Xbar_i'-mu_i'||_F), which is controlled exactly by the event (1/m)||Xbar_i-mu_i|| < ε/4 for all i. There is no quadratic cross term in this definition. The Chebyshev calculation yields a failure probability 16∑γ/(r m^2 ε^2), not 16∑γ/(r m ε^2); since m^2 ≥ m, the paper's displayed denominator is a conservative typo that only weakens the bound. The downstream steps (Corollary 1, Theorem 2) are unaffected because the larger failure probability still tends to zero under r=ω(n^3). The central rate (n^3/r)^{1/2-δ} is therefore not undermined by the reader's stated gap. I do not identify a load-bearing flaw in the central argument, though the proof has minor algebraic typos and relies on an external decomposition.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the Data Kernel Perspective Space (DKPS) embedding of black-box generative models. For n generative models, m queries, and r iid responses per model-query pair, the authors construct a sample dissimilarity matrix D from the estimated mean responses and apply classical multidimensional scaling to obtain estimated perspectives. The main results are high-probability bounds: Theorem 1 gives an entrywise bound on the doubly centered dissimilarity matrix, Corollary 1 converts this into a spectral-norm bound under r = ω(n^3), and Theorem 2 / Corollary 2 give a uniform (2-to-infinity) embedding error bound after an orthogonal alignment, of order (n^3/r)^{1/2−δ}, under constant-rank and eigenvalue-separation assumptions. The claimed contribution is a finite-sample rule for choosing the number of responses r to achieve a desired embedding accuracy.","tokens_in":16433,"tokens_out":16234,"duration_ms":163873,"significance":"If the main theorem is correct, the paper provides a useful finite-sample guarantee for a practical embedding method, with explicit polynomial constants and a clear message about the n^3/r tradeoff. The proof strategy is sensible: a second-moment bound on the sample means, followed by spectral perturbation tools (Weyl, Davis-Kahan, and the Agterberg et al. MDS decomposition). The paper also goes beyond a purely asymptotic consistency result and connects the theory to simulations and a real LLM experiment with claimed 100% empirical coverage. The main limitation is that the perturbation part of the proof is not fully self-contained, relying on an external decomposition; nevertheless the central rate appears defensible. The paper does not provide code or machine-checked proofs, but the numerical tables are useful sanity checks.","major_comments":[],"minor_comments":[{"comment":"The Chebyshev step as written gives the failure probability 16Σγ/(r m^2 ε^2), not 16Σγ/(r m ε^2). Since m^2 ≥ m, the displayed bound with rmε^2 is actually a valid conservative lower bound, so the theorem statement is not harmed. Still, the intermediate derivation should be corrected to avoid confusion. Relatedly, the cross-term concern raised in the stress-test does not land: because D_{ii'} = m^{−1/2}||Xbar_i−Xbar_i'||_F^{1/2}, we have D_{ii'}^2 = (1/m)||Xbar_i−Xbar_i'||_F, so the reverse triangle inequality controls |E_{ii'}| exactly as the paper claims.","section":"Theorem 1 proof, Section 7.1"},{"comment":"The first line of the proof says 'From Theorem 2, using Triangle Inequality' and then invokes a decomposition into R_1,...,R_6. This should be a reference to Agterberg et al. (2022), and the decomposition should be stated explicitly, with all quantities (U, Ũ, Λ, Λ̂, I_{p,q}, W_*) defined in one place. As written, the main theorem's proof is hard to verify without consulting the cited preprint.","section":"Proof of Theorem 2, Section 7.1"},{"comment":"The statement 'from Corollary 2' is used where the spectral-norm bound on ||B̂−B|| is needed; this is Corollary 1, not Corollary 2. The same citation typo appears in Lemma 4. These are harmless but should be fixed to avoid the appearance of circularity.","section":"Lemmas 2 and 4, Section 7.2"},{"comment":"The final paragraph says the paper bounds error in Frobenius norm and identifies a uniform two-to-infinity bound as future work. This is contradicted by Theorem 2 and Corollary 2, which already prove a ||·||_{2,∞} bound. The discussion appears to be from an earlier draft and should be rewritten.","section":"Section 6 Discussion"},{"comment":"The text says m=2 is kept constant in the simulations, but Table 1 reports m=3. Please make the experimental settings consistent. The text also says the simulated model 'outputs binary responses' while the procedure describes constructing probabilities; clarify the response mechanism.","section":"Section 5.1 and Table 1"},{"comment":"The population-level doubly centered matrix is defined as B = −(1/2)H_n Δ^{◦2} H_n, but the text says B = −(1/2)H_n Δ H_n. The Hadamard-square is missing. Also, the real-data experiment uses an estimated μ_i from R = n^{3.75} responses as the 'true' ψ*, so the reported coverage checks the bound against a bootstrapped empirical target, not the actual population. This should be described as such.","section":"Section 5.2"},{"comment":"Figure 1 shows n=4,6,8,10,12 while Table 2 uses n=10,12,14,16. The sample sizes should be aligned, and the figure caption's description of 'r=n^5' should be checked against the main text.","section":"Figure 1 and Table 2"}],"recommendation":"minor_revision","confidential_remarks":"The reader's rejection is based on a supposed gap in Theorem 1 that, on reading the paper, does not hold: the definition of the dissimilarity makes the cross-term objection immaterial, and the Chebyshev denominator issue is a conservative typo. The central rate is plausible. My recommendation of minor revision is driven by presentation and proof-completeness issues rather than by a load-bearing mathematical error. I would ask the authors to state the Agterberg et al. decomposition explicitly and to correct the several internal inconsistencies before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: I disagree with the reader's verdict. The alleged load-bearing gap in Theorem 1 does not exist. The dissimilarity is defined as D_ii' = m^{-1/2} ||Xbar_i - Xbar_i'||_F^{1/2}, so D^2_ii' = (1/m)||Xbar_i - Xbar_i'||_F. The difference E_ii' is therefore (1/m)(||Xbar_i - Xbar_i'||_F - ||mu_i - mu_i'||_F), and the reverse triangle inequality bounds it by (1/m)(||Xbar_i - mu_i||_F + ||Xbar_i' - mu_i'||_F). The event in the proof controls exactly those two terms. There is no quadratic cross term. The reader's Chebyshev denominator concern is also a typo: the derivation gives 16 sum gamma / (r m^2 epsilon^2), while the paper displays r m epsilon^2; since m^2 >= m, the displayed bound is conservative, not optimistic. So Theorem 1 stands as stated.\n\nWhat's new: finite-sample concentration bounds for DKPS embeddings, with an explicit rate (n^3/r)^{1/2 - delta} under r = omega(n^3). That is a real step beyond the consistency results in Helm et al. and Acharyya et al., and the template using the Agterberg decomposition can be reused for classical MDS with noisy dissimilarities. The explicit polynomial coefficients in Theorem 2 are a useful addition.\n\nSoft spots: (1) The proof of Theorem 1 has the denominator typo; a referee should ask for the corrected line. (2) Proposition 1 is stated but not proved — it is a sufficient condition for Assumptions 1-2, and maybe it is meant as a remark, but as written it is an unproved proposition. (3) The proof of Theorem 2 leans heavily on Agterberg et al.'s decomposition, so the paper is not fully self-contained; that is acceptable but should be stated clearly. (4) The empirical bounds are very loose (upper bounds roughly 10-100x the actual error), and the authors acknowledge this; the simulations are illustrative, not a stress test of the constants. (5) Minor typos: 'From Theorem 2' in the proof of Theorem 2 should be 'from the decomposition'.\n\nOverall: the central argument holds, the rate is plausible, and the result is useful for anyone doing inference over collections of generative models. It deserves a serious referee. I would send it to review, asking for a corrected Chebyshev line and a proof or clear downgrade of Proposition 1.","headline":"The reader's rejection rests on a misreading of the dissimilarity definition; the central concentration bound is sound and deserves a serious referee despite minor proof typos.","tokens_in":17016,"tokens_out":5889,"would_cite":true,"duration_ms":53526,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves a finite-sample concentration bound for the response-based vector embedding of generative models: the per-model estimation error is controlled by a polynomial in (n^3/r)^{1/2} once r exceeds n^3.","keywords":["generative models","black-box models","response-based embeddings","Data Kernel Perspective Space","classical multidimensional scaling","concentration inequalities","finite-sample bounds"],"falsifier":"Simulate a small number of generative models with known response covariances and fixed $n,m,r$, compute the empirical frequency of events where $|\\hat{B}_{ii'} - B_{ii'}| > \\epsilon$, and compare it against the claimed bound $16 \\sum \\gamma_{ij}/(r m \\epsilon^2)$; if the observed failure rate exceeds the bound for any combination of parameters — especially when the cross term $2 \\langle \\mu_i - \\mu_{i'}, (\\bar{X}_i - \\bar{X}_{i'}) - (\\mu_i - \\mu_{i'}) \\rangle / m$ is large — the concentration claim is refuted.","tokens_in":15967,"feed_emoji":"🤖","tokens_out":14711,"duration_ms":128331,"temperature":0.7,"texified_at":"2026-08-05T20:38:46.008681+00:00","pith_summary":"The paper aims to give the \"Data Kernel Perspective Space (DKPS)\" embedding — a way of representing each black-box generative model by the vector obtained from classical multidimensional scaling of its average vectorized responses to a set of queries — a finite-sample statistical guarantee. It proves that when the response distributions have uniformly bounded variability and the number of responses per query r grows faster than the cube of the number of models n, the sample embedding is close to the population embedding in a strong uniform sense: the worst-case per-model error, after an optimal rotation, is bounded by a polynomial in $(n^3/r)^{1/2-\\delta}$ with high probability. This is the first concentration-type guarantee for response-based generative model embeddings, and it directly tells a practitioner how many responses to collect to reach a desired level of embedding accuracy. The argument relies on a general perturbation bound for classical multidimensional scaling that applies any time dissimilarities are observed with noise.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":14965,"prompt_tokens":902,"completion_tokens":14063,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":902,"completion_tokens_details":{"reasoning_tokens":13213}},"feed_headline":"n^3 samples per query pin down generative model embeddings","feed_subtitle":"Proved with high-probability bounds; tells you how many responses to collect for a target embedding accuracy.","key_machinery":"The load-bearing object is the sample double-centered dissimilarity matrix $\\hat{B} = -\\frac{1}{2} H_n D^{\\circ 2} H_n^T$, where $D_{ii'} = \\frac{1}{\\sqrt{m}} \\|\\bar{X}_i - \\bar{X}_{i'}\\|_F$ is the empirical squared-distance between the mean response vectors of models $i$ and $i'$. The mechanism is a three-step chain: an entrywise Chebyshev concentration bound on $D^2 - \\Delta^2$ (Theorem 1), a spectral-norm concentration bound on $\\hat{B} - B$ via Corollary 1, and a decomposition of the embedding perturbation $\\hat{\\psi} W_* - \\psi$ into six remainder matrices whose norms are each controlled by powers of $\\|\\hat{B} - B\\|$.","core_discovery":"The central claim is Theorem 2: under Assumption 1 (the population double-centered dissimilarity matrix $B$ has constant rank $d$) and Assumption 2 (its nonzero eigenvalues are bounded away from 0 and from infinity), if the response covariances satisfy $\\sup_{i,j} \\operatorname{trace}(\\Sigma_{ij}) = O(1)$ and $r = \\omega(n^3)$, then for every $\\delta \\in (0,1/2)$ there exists an orthogonal matrix $W_*$ such that $\\|\\hat{\\psi} W_* - \\psi\\|_{2,\\infty} \\le \\mathrm{Poly}_3((n^3/r)^{1/2-\\delta})$ with high probability. Corollary 2 states this as an $O_P((n^3/r)^{1/2-\\delta})$ rate for the worst-case per-model embedding error after rotation. The proof obtains an entrywise concentration bound on the squared sample dissimilarities (Theorem 1), co","pith_inferences":["If the missing concentration step in Theorem 1 is repaired for sub-Gaussian responses, the Chebyshev argument could likely be upgraded to an exponential concentration inequality, giving a much faster rate than the polynomial in n^3/r.","The cubic dependence on n suggests a practical ceiling: for very large model collections the per-query response budget becomes prohibitive, so a more scalable design might share responses across queries or models to break the n^3 barrier.","The proof machinery appears generic enough to handle dissimilarities other than squared Euclidean distances (e.g., kernel or Wasserstein distances), as long as the distance satisfies a Lipschitz condition in the mean vectors; this is a natural next target.","The discrepancy between the simulation errors (order 10^{-4}) and the theoretical bound (order 10^{-1}) indicates the constant in Poly_3 is loose; for real use the bound should be recalibrated empirically, not taken as a tight error forecast."],"forward_implications":["A practitioner can choose the number of responses r so that, with high probability, every estimated perspective is within a target distance of its population value, up to a rotation.","Rotation-invariant inference tasks (for example, testing whether two models share the same perspective) inherit the same finite-sample guarantee: the paper's discussion shows that if r is large enough, an observed difference larger than 2kappa forces a true difference.","The result is the first non-asymptotic concentration guarantee for response-based embeddings of black-box generative models, converting previously known consistency results into explicit sample-size rules.","The same algebraic tools transfer to general classical multidimensional scaling: whenever a dissimilarity matrix is observed with noise satisfying the entrywise control, a uniform bound on the resulting embedding follows.","Smaller response variability gamma_{ij} (for example, lower-temperature generation) improves the Theorem 1 bound linearly, so the method's sample requirement is directly tied to how noisy the responses are."],"fun_headline_variants":["n^3 samples per query tighten embedding bounds","Cubic sample rate proves embedding concentration","For embedding accuracy, collect n^3 responses","n^3 sample bound for black-box model embeddings","High-probability bounds: n^3 samples do the trick"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The rate hinges on the unproved implication in Theorem 1 that when every sample mean vector is close to its population mean, the sample squared-distance matrix is close to the population squared-distance matrix; if that implication fails, the main concentration bound and its sample-size rule do not follow.","fun_headline_variants_meta":{"raw":{"variants":["n^3 samples per query tighten embedding bounds","Cubic sample rate proves embedding concentration","For embedding accuracy, collect n^3 responses","n^3 sample bound for black-box model embeddings","High-probability bounds: n^3 samples do the trick"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000397,"raw_usage":{"total_tokens":1908,"prompt_tokens":726,"completion_tokens":1182,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":1108}},"tokens_in":470,"tokens_out":1182,"duration_ms":11887,"temperature":1.0,"reasoning_tokens":1108,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:53:20.636511+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a small number of generative models with known response covariances and fixed $n,m,r$, compute the empirical frequency of events where $|\\hat{B}_{ii'} - B_{ii'}| > \\epsilon$, and compare it against the claimed bound $16 \\sum \\gamma_{ij}/(r m \\epsilon^2)$; if the observed failure rate exceeds the bound for any combination of parameters — especially when the cross term $2 \\langle \\mu_i - \\mu_{i'}, (\\bar{X}_i - \\bar{X}_{i'}) - (\\mu_i - \\mu_{i'}) \\rangle / m$ is large — the concentration claim is refuted.","supporting_citations":[],"review_version":1}