REVIEW 2 major objections 5 minor 18 references
A posterior over meaningful states can be recovered from grouped language probabilities whenever the induced observation law is injective and stable.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:30 UTC pith:76O47CJZ
load-bearing objection A serious and unusually candid methods paper that builds the right statistical scaffolding for turning LLM phrase probabilities into auditable state posteriors, provided the explicitly flagged sufficiency assumption is accepted. the 2 major comments →
Identification and Learning of Semantic Observation Kernels: Partial Observation, Uniform Recovery, & Minimax Limits
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the observable conditional probabilities a text generator assigns to complete continuations can serve as a measurement channel for a posterior over declared states, provided meanings are first grouped into prespecified semantic cells. The paper formalizes the channel as a finite pushforward of the language kernel through a measurable semantic partition, and proves that if this observation map is injective and continuous on a compact supported domain, the inverse exists and its finite-sample error is bounded by an inverse modulus applied to observation and optimization error. It then establishes that the calibration surface m0(s)=E(Theta|S=s) can be recovered from he
What carries the argument
The central object is the semantic observation law g_u(theta) = (Q_u{phi_u(Y)=j | theta})_j, the finite pushforward of an unrestricted language kernel through a declared semantic partition phi_u, together with the observable experiment E_O = {P^O_theta}. The inverse modulus omega_D(epsilon) = sup{||theta-theta'|| : d_E(P^O_theta,P^O_theta') <= epsilon} is the key quantity: it encodes both point identification (omega(0)=0) and stable conditioning (omega(eps)->0). The calibration map m0(s)=E(Theta|S=s) is the identifiable predictive inverse, and the derivative functional kappa(F0)=sup ||D_z F0||_op is the contraction certificate for repeated use. Each carries a distinct part of the argument: t
Load-bearing premise
The load-bearing premise is that the grouped language probability vector S is a sufficient statistic for the reference posterior—that Theta = m0(S) almost surely on the supported domain—together with injectivity of the induced observation map; without sufficiency, the learned map is a conditional mean, not a state posterior.
What would settle it
Find two calibration scenarios on the same supported domain whose observed grouped language probability vectors are equal (or within sampling error) but whose reference posteriors differ materially; if such collisions occur in held-out data beyond expected noise, then S is not sufficient and the predictive inverse is not a structural state measurement. Similarly, observe a small perturbation of a grouped probability vector that produces a posterior jump larger than the modeled inverse modulus, which would falsify stable recovery.
If this is right
- On a declared state space, evidence class, prompt family, and supported calibration domain, an LLM's grouped phrase probabilities can be converted into a state posterior with quantified sampling error, and the conversion is auditable because it uses only observable probabilities.
- Prompt rewording is harmless only when it leaves the grouped semantic vector unchanged; when it changes that vector, presentation has altered the measurement and must be modeled as observation noise.
- Unreturned continuation probability mass must be represented as an identified set; assigning zero would manufacture certainty, and the sequential width of such sets is bounded by contraction of the state transition plus a persistent term.
- One-shot predictive accuracy is insufficient for recursive deployment: stable repeated updating requires a simultaneous band on the derivative of the update, and no estimator can separate contraction from noncontraction faster than n^{-(s-1)/(2s+d)}.
- A conformal region around the recovered posterior gives distribution-free marginal coverage but cannot establish identification; identification and support must be checked before coverage is interpreted.
Where Pith is reading between the lines
- Extension: the same measurement construction could be applied to any fitted probabilistic generator with a defensible probability record, such as clinical prediction models or speech recognition systems, since the theorems do not depend on a transformer architecture.
- Extension: the structural-inversion assumption can be tested empirically by checking whether held-out reference coordinates still vary after conditioning on the semantic probability summary; the paper's own table marks global injectivity as unverifiable, so a deployment guide would need a threshold for when a conditional-mean output is acceptable as a state measurement.
- Extension: because the prompt is part of the observation design, an active measurement strategy that samples multiple information-equivalent prompts and checks agreement in calibrated outputs could strengthen evidence for identification beyond a single frozen prompt family.
- Extension: if the minimax derivative boundary is intrinsic, high-stakes recursive use should budget sample sizes according to the desired resolution of the contraction coefficient near one, rather than only the predictive accuracy of the update.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a semiparametric inverse-problem framework for converting observable grouped language probabilities from a frozen probabilistic text generator into a posterior over declared states. The main theoretical objects are an inverse modulus separating identification from conditioning (Theorem 2.2), a finite-replication/CLT result (Corollary 2.3), an impossibility bound under observational equivalence (Theorem 2.4), uniform nonparametric recovery rates for the predictive inverse under beta-mixing and estimated covariates (Theorem 3.2), a derivative-based contraction gate for sequential stability with an exact Gaussian-sieve band (Corollary 3.4) and a minimax lower bound (Theorem 3.5), and a partial-identification propagation theorem for truncated candidate probabilities (Theorem 4.1). Simulations are paired one-for-one with the theorems, and two frozen LLM studies on a 520-scenario financial benchmark illustrate held-out recovery and conformal coverage.
Significance. If the framework's assumptions hold, it provides a principled way to use LLM phrase probabilities as auditable state measurements without interpreting internal representations, separating semantic measurement from internal belief. The paper is unusually careful: Table 1 explicitly distinguishes verified-by-design conditions from assumed conditions, the proofs are standard, and the simulation design tests each theorem with a reproducibility supplement and fixed seeds. The main scientific value is the formalization of the inverse-modulus/derivative-gate distinction, which clarifies why predictive calibration cannot by itself certify structural recovery or recursive stability. The central caveat is that the learned object is the conditional mean E(Theta|S=s), so the measurement interpretation depends on a sufficiency/structural-identification condition that is plausible but not empirically verified in the presented archive.
major comments (2)
- [Section 3, Table 1, Section 5.2] The identified object is m0(s)=E(Theta|S=s) (Section 3, 'Predictive versus structural inversion'), and all recovery rates in Theorem 3.2 are rates for this regression surface. The interpretation as recovery of the reference posterior requires the additional structural/sufficiency condition Theta_i=m0(S_i) a.s. Table 1 candidly marks global injectivity as 'not identified' and the calibration map as 'assumed and stress-tested, not globally verifiable.' Consequently, the held-out JS divergences in Table 5 demonstrate predictive calibration, not structural posterior recovery; the 520-scenario archive never varies the semantic partition or evidence set to test whether S is a sufficient coarsening. The abstract and Section 5.2 should either report a direct sufficiency check (e.g., nested semantic partitions or evidence ablations) or explicitly restrict the central claim to predictive calibrati
- [Theorem 3.2 and Appendix A.4] The 'estimated semantic probabilities' extension asserts an additional O_P(C_G sqrt(log n/R_min)) function error and its h_n^{-1} multiple for derivatives, based on perturbing the local-polynomial normal equations. But when the covariate S_i itself is estimated, the regression surface identified from the estimated covariate is generally a convolution of m0 with the measurement-error distribution unless further conditions are imposed; a uniform bound on max_i ||hat S_i - S_i|| does not by itself yield the displayed perturbation argument. A formal generated-covariate expansion in the style of Mammen et al. (2012), with the required conditions stated, is needed before this rate claim can be considered established.
minor comments (5)
- [Section 2, Definition 2.1] The notation L[...] is used without definition; please define it as 'the law of' at first use.
- [Section 2, worked equivalence example] 'Theorem 2' in the worked example should be 'Theorem 2.4.'
- [Section 3, 'Simple rate calculation'] The phrase 'reduces function error by about 32^{-2/5}=1/4' is mathematically misleading: 32^{-2/5}=0.25 is the remaining factor, so the error is reduced to about 1/4, not by 1/4.
- [Table 5] Coverage values .940 and .900 are based on single 100-scenario test partitions; report the calibration and validation sizes and give a binomial interval for coverage so the reader can judge sampling variability.
- [Appendix A.2] The phrase 'absorbing the conventional factor between total variation and L1 distance into the displayed modulus' is vague; state the factor explicitly (e.g., TV = (1/2)||.||_1) and adjust the modulus accordingly.
Circularity Check
No significant circularity: the held-out 'recovery' is an out-of-sample regression check, explicitly distinguished from structural inversion; all theorems follow from stated assumptions.
full rationale
The paper's central object is m0(s)=E(Theta|S=s), which it explicitly labels the 'identified predictive inverse' and states 'it equals a structural inverse only when Theta_i=m0(S_i) almost surely' (Section 3, 'Predictive versus structural inversion'). The held-out empirical study fits this regression on a calibration partition and evaluates it on a disjoint test partition, so the reported JS divergence is an honest out-of-sample prediction error, not a fitted parameter renamed as a discovery. Theorems 2.2, 2.4, 3.2, 3.3, 3.5, and 4.1 are proved from explicit assumptions (compact support, injectivity, Holder smoothness, beta-mixing, Gaussian noise, etc.); none of them reduces to its conclusion by construction. The paper repeatedly discloses that global injectivity is 'not identified' and that the smooth calibration map is 'assumed and stress-tested, not globally verifiable' (Table 1). That is a scope limitation, not circularity. There are no self-citations, no imported uniqueness theorems from the authors, and no ansatz smuggled in via citation. The sufficiency condition Theta=m0(S) is indeed untested, but the paper does not claim to verify it; instead it explicitly conditions the structural interpretation on that assumption. Consequently, no circular step can be quoted from the text.
Axiom & Free-Parameter Ledger
free parameters (3)
- Local-polynomial bandwidth h_n
- Finite-difference step h for the derivative gate
- Prompt-offset correction
axioms (7)
- domain assumption Assumption 3.1: calibration pairs (S_i, Θ_i) are stationary, geometrically beta-mixing, with bounded design density, sub-Gaussian errors, and m0 in a Hölder ball C^s(S0; B).
- domain assumption The observation map θ ↦ P^O_θ is continuous and injective on the supported domain.
- domain assumption Reference posterior π*(E) is specified independently of the fitted language law.
- domain assumption Semantic map φ_u is measurable and prespecified; no post-test ontology selection.
- ad hoc to paper Corollary 3.4: F0 lies in a correctly specified p-dimensional Gaussian linear sieve with fixed full-rank design and normal errors.
- domain assumption For generated regressors, each S_i is estimated from R_min repeats via a locally C_G-Lipschitz inverse.
- domain assumption State transition F_t is L_t-Lipschitz and the set-valued update U_t is Hausdorff nonexpansive.
read the original abstract
Probabilistic text generators supply conditional distributions over tokens and complete verbal continuations, whereas scientific use often requires a posterior over a finite state. Large language models are the leading example: phrase probabilities depend on prompt wording, and model-printed percentages are generated text rather than state posteriors. More generally, we ask when an observable language law can support a reproducible posterior over declared states. A semantic map groups meaning-equivalent continuations; held-out cases with reference posteriors identify a semiparametric inverse from grouped language probabilities to state probabilities. The language law remains nonparametric and no hidden model quantity is used. This is principally a theory and methods paper and makes several contributions. On the theoretical side, we derive conditions for existence, identification, stable recovery, and sequential updating; concentration, asymptotic, and nonparametric rates; identified sets under truncated probabilities; and a minimax boundary for uniform stability. On the empirical side, theorem- directed simulations verify recovery rates, compatible-set coverage, and stability gates, while two frozen language-model studies illustrate held-out recovery and conformal coverage. The results specify when observable language probabilities can provide an auditable state measurement without being interpreted as internal belief.
Figures
Reference graph
Works this paper leans on
-
[17]
doi: 10.18653/v1/2025.findings-acl.1101. Bin Yu. Rates of convergence for empirical processes of stationary mixing sequences.The Annals of Probability, 22(1):94–116,
-
[1953]
Mark Braverman, Xinyi Chen, Sham Kakade, Karthik Narasimhan, Cyril Zhang, and Yi Zhang
doi: 10.1214/aoms/1177729032. Mark Braverman, Xinyi Chen, Sham Kakade, Karthik Narasimhan, Cyril Zhang, and Yi Zhang. Calibration, entropy rates, and memory in language models. InProceedings of the 37th Interna- tional Conference on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pages 1089–1099,
-
[1964]
François Le Gland and Laurent Mevel
doi: 10.1214/aoms/1177700372. François Le Gland and Laurent Mevel. Exponential forgetting and geometric ergodicity in hidden markov models.Mathematics of Control, Signals, and Systems, 13(1):63–93,
-
[1982]
doi: 10.1111/j.2517-6161.1982.tb01195.x. Anastasios N. Angelopoulos and Stephen Bates. Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16(4):494–591,
arXiv 1982
-
[1994]
doi: 10.1214/aop/1176988849. 28
-
[1996]
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal
doi: 10.1007/978-94-009-1740-8. Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy.Nature, 630:625–630,
-
[2003]
doi: 10.1023/A:1023818214614. Heinz W. Engl, Martin Hanke, and Andreas Neubauer.Regularization of Inverse Problems. Kluwer Academic Publishers,
-
[2006]
doi: 10.1201/9781420010138. Laurent Cavalier. Nonparametric statistical inverse problems.Inverse Problems, 24(3):034004,
-
[2008]
Victor Chernozhukov, Han Hong, and Elie Tamer
doi: 10.1088/0266-5611/24/3/034004. Victor Chernozhukov, Han Hong, and Elie Tamer. Estimation and confidence regions for parameter sets in econometric models.Econometrica, 75(5):1243–1284,
-
[2009]
doi: 10.1007/ b13794. Ramon van Handel. Observability and nonlinear filtering.Probability Theory and Related Fields, 145:35–74, 2009a. doi: 10.1007/s00440-008-0161-y. Ramon van Handel. Uniform observability of hidden markov models and filter stability for unstable signals.The Annals of Applied Probability, 19(3):1172–1199, 2009b. doi: 10.1214/08-AAP576. Z...
-
[2010]
doi: 10.1146/annurev.economics.050708.143401. Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. InProceedings of EMNLP, pages 5433–5442,
-
[2011]
doi: 10.3982/ECTA8680. David Blackwell. Equivalent comparisons of experiments.The Annals of Mathematical Statistics, 24 (2):265–272,
-
[2012]
doi: 10.1214/12-AOS995. Charles F. Manski.Partial Identification of Probability Distributions. Springer, New York,
-
[2014]
Juan José Egozcue, Vera Pawlowsky-Glahn, Glòria Mateu-Figueras, and Carles Barceló-Vidal
doi: 10.1214/14-AOS1230. Juan José Egozcue, Vera Pawlowsky-Glahn, Glòria Mateu-Figueras, and Carles Barceló-Vidal. Isometric logratio transformations for compositional data analysis.Mathematical Geology, 35(3): 279–300,
-
[2018]
Enno Mammen, Christoph Rothe, and Melanie Schienle
doi: 10.1080/01621459.2017.1307116. Enno Mammen, Christoph Rothe, and Melanie Schienle. Nonparametric regression with non- parametrically generated covariates.The Annals of Statistics, 40(2):1132–1170,
arXiv 2017
-
[2021]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar
doi: 10.1162/tacl_a_00407. Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. InInternational Conference on Learning Representations,
-
[2023]
doi: 10.18653/v1/2023.emnlp-main.330. Alexandre B. Tsybakov.Introduction to Nonparametric Estimation. Springer,
-
[2025]
Zhiqiu Xia, Jinxuan Xu, Yuqian Zhang, and Hang Liu
URLhttps://proceedings.mlr.press/v258/wang25i.html. Zhiqiu Xia, Jinxuan Xu, Yuqian Zhang, and Hang Liu. A survey of uncertainty estimation methods on large language models. InFindings of the Association for Computational Linguistics: ACL 2025, pages 21381–21396,
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.