REVIEW 2 major objections 4 minor 38 references
A language model's probabilities over phrases can be turned into calibrated, auditable posterior estimates over meaningful states when a semantic map and held-out calibration are fixed in advance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:55 UTC pith:5VISVFQG
load-bearing objection A serious, unusually honest attempt to turn LLM continuation probabilities into calibrated posteriors over declared states, but the theoretical recovery guarantee is not certified by the experiments — the empirical claims rest on a constrained grammar and finite-design diagnostics. the 2 major comments →
Calibrating Semantic Uncertainty from Observable Language-Model Probabilities
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the observable, prompt-dependent distribution over verbal responses can serve as a measurement of a reference posterior over a finite set of declared states. The construction fixes a semantic coarsening φ_u (complete continuations mapped to states or to an unexpressed category), forms the normalized lexical composition p_u = pushforward of Q_u under φ_u, and calibrates an inverse map ψ_u in alr coordinates to obtain bπ_u = alr^{-1}∘ψ_u∘alr∘p_u∘C_u. Under a positive lower-modulus condition the recovery error is bounded by 2δ/c_u (or δ/κ_u in the affine case), so observability of the language channel is exactly what protects against amplified noise. The paper's experi
What carries the argument
The carrying object is the semantic map, defined as a prespecified measurable coarsening φ_u of the continuation space into K declared semantic states plus an unexpressed symbol ⊥, followed by a calibrated inverse ψ_u fitted in additive log-ratio (alr) coordinates. The estimator is the composition bπ_u = alr^{-1}∘ψ_u∘alr∘p_u∘C_u, where p_u is the normalized pushforward of the language law under φ_u. The map's load-bearing property is the lower modulus c_u: the minimal separation the language channel preserves between different reference posterior log-ratios. It converts a bounded observation error δ into a bounded recovery error 2δ/c_u, and its reciprocal controls how strongly inversion ampl
Load-bearing premise
The load-bearing premise is that the forward map from reference posterior log-odds to lexical log-odds is identifiable and stable—has a positive lower modulus and is well approximated by the fitted affine class—on the intended operating domain, and the experiments certify this only on a sampled design for a constrained three-code response grammar rather than for unrestricted free text.
What would settle it
Find a pair of evidence values inside the declared operating domain whose exact reference posteriors differ materially but whose lexical compositions under the same prompt and semantic map are (nearly) identical, so the fitted affine map's smallest singular value is effectively zero on that pair. Alternatively, apply the published semantic map and calibration to a different language-model checkpoint or to unrestricted free-text responses and check whether held-out conformal coverage falls below the nominal level or held-out error exceeds the reported bounds.
If this is right
- Language-model continuation probabilities can be treated as a measurement channel with a declared estimand, making posterior estimates auditable rather than ad hoc.
- Uncertainty sets with nominal coverage can be attached to language-derived state probabilities when calibration and conformal radii are computed on held-out scenarios.
- Information-preserving rewording can be validated as a nuisance factor, while reordering that preserves the facts is shown to be consequential.
- Sequential Bayesian updating becomes defensible when the reference filter is contractive and one-step recursion defects are observable.
- The same declared-map-plus-calibration template extends to auditing classifications and recommendations beyond probability reporting.
Where Pith is reading between the lines
- The constrained three-code response grammar used in the controlled experiment may be much easier to calibrate than unrestricted free text; using the method on open-ended prose will likely require richer semantic partitions and explicit handling of unexpressed mass, and the reported results do not certify that setting.
- Because the forward map is certified only as affine on the sampled design, posterior estimates near the simplex boundary or outside the calibrated region could carry larger inversion error than the average held-out metrics suggest.
- The dependence on a frozen fitted language distribution implies that model updates or prompt-service changes invalidate the calibration; deployed systems would need ongoing recalibration rather than a one-time semantic map.
- The entropy contrast (raw word entropy nearly uncorrelated with posterior entropy, calibrated entropy correlated around 0.74) suggests that using raw token entropy as an uncertainty measure is unsupported; the semantic map plus calibration is the effective bridge.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a 'semantic map': a prespecified statistical bridge from an LLM's observable continuation probabilities over verbal responses to a reference posterior over a finite set of declared states. The construction separates the reference experiment from the language experiment, defines a lexical composition p_u through semantic coarsening φ_u and an unexpressed-mass category ⊥, and proposes held-out calibration of an inverse map ψ_u to estimate the reference posterior. The theoretical part derives an approximation decomposition (Theorem 1), a presentation-stability bound (Theorem 2), an inverse-recovery bound based on a lower-modulus observability condition (Theorem 3), and a sequential-filtering error bound (Theorem 4). The empirical part compares continuation-derived lexical probabilities with numerically elicited probabilities on market text, and evaluates posterior recovery and conformal coverage in a controlled three-state experiment with two fitted LLMs. The paper is unusually careful about prespecification, scenario clustering, and limitations, but the central empirical certification is confined to a constrained enum response grammar and does not certify the key observability condition of Theorem 3.
Significance. If the central claim held as stated, the paper would be a substantial contribution: it provides a declared estimand, a decomposition of semantic, probability-mass, and inverse-recovery error, and a held-out validation protocol for converting LLM token probabilities into auditable state posteriors. The strengths are genuine: the semantic coarsening is prespecified, calibration and testing use disjoint scenarios, uncertainty is clustered by scenario, the experiments are replicated across two fitted models, and the limitation statements are unusually explicit. However, the significance is currently limited by a gap between the advertised theorem-driven recovery guarantee and the empirical evidence, which establishes only finite-design, average-case recovery under a constrained grammar. The contribution is therefore promising but needs reframing before the abstract-level claims are supported.
major comments (2)
- [§3.2, 'Observability interpretation'; §2.2.3, Theorem 3] The central recovery guarantee of Theorem 3 is conditional on a positive lower modulus c_u. The paper's own diagnostic reports that the empirical pairwise lower modulus is 'effectively zero' for GPT-4.1-mini and 'zero at machine precision' for GPT-4o-mini, with affine residual RMSEs of 7.62 and 7.35 in lexical log-ratio units. If the pairwise modulus on the sampled design is zero, there exist reference posteriors that are indistinguishable from the mean lexical measurement, and Theorem 3's bound 2δ/c_u is vacuous. The positive bootstrap lower bounds on the smallest singular value (1.686 and 1.600) are properties of the fitted affine matrix, not of the true forward regression h_u, and the large residuals show the affine approximation is poor. Consequently, the abstract's claim that the method 'recover[s] held-out posteriors with valid uncertainty coverage' as an 'auditable posterior estim
- [§2.1.1 (constrained law Q^G_u); §3.2; Abstract] The general theorems are stated for the unrestricted free-text law Q_u, but the experiments evaluate only the constrained enum grammar Q^G_u with three enum codes plus an explicit residual, under a completeness threshold of 0.999999. The paper acknowledges this in Section 2.1.1 and Section 3.2.1, but the abstract and several conclusions state without qualification that 'language-derived probabilities outperform printed numerical probabilities, recover held-out posteriors with valid uncertainty coverage.' This overstates the scope of the empirical certification. The abstract and conclusion should explicitly state that the recovery results are for the constrained response grammar used in the controlled experiment and do not extend to unrestricted free-text continuations.
minor comments (4)
- [§2.2.3, Theorem 3] The theorem assumes existence of the argmin b̂ℓ but does not state a compactness or attainment condition. The appendix proof mentions compactness of L; this assumption should be stated in the theorem itself.
- [Table 7] The split-conformal coverage values 0.94 and 0.90 are reported on 100 test scenarios without a confidence interval. With 100 scenarios, the binomial standard error is about 0.024–0.030, so nominal coverage is plausible but not tightly certified; this should be noted or an interval provided.
- [§3.2, entropy association] The raw entropy correlation r ≈ −0.001 and calibrated posterior entropy r = 0.742 are reported without confidence intervals. Given that these are descriptive and based on a modest number of scenarios, an interval or at least a statement of uncertainty would avoid overinterpretation.
- [Throughout] The distinction between Q_u and Q^G_u is central, but the notation is introduced only in Section 2.1.1 and then not always reused. For readability, consider using Q^G_u consistently in the experimental sections and in Table 4 to remind readers which law is being certified.
Circularity Check
No significant circularity: held-out calibration breaks the main reduction; theorems are standard conditional bounds with explicit assumptions.
full rationale
The paper's derivation chain is not circular. The reference posterior π⋆ is defined by a declared reference model, the lexical composition p_u is derived from the LLM's observable continuation probabilities through a prespecified semantic map φ_u, and the inverse ψ_u is fitted on calibration scenarios and tested on untouched scenarios. This held-out design prevents the central 'recovery' claim from reducing to a fitted input: the test posteriors are not used to fit the inverse, and the reported held-out Jensen–Shannon errors and split-conformal coverage are genuine out-of-sample evaluations. The theorems are standard deterministic bounds (inverse-problem stability, Birkhoff contraction) with assumptions stated explicitly; they are not imported from the authors' prior work, and the paper contains no self-citations. The semantic grouping is credited to the external literature (Farquhar et al., Kuhn et al.) and is explicitly a prespecified design choice rather than an ansatz smuggled in via citation. The paper's own limitations—'Finitely many observations cannot determine this uniform infimum' and the empirical pairwise lower modulus being effectively zero—are validity caveats about whether Theorem 3's condition is certified, not evidence that a prediction is equivalent to its input by construction. The observability diagnostic is honestly framed as a property of the fitted affine map on the sampled design, not as a uniform identifiability result. Thus no specific circular step can be exhibited.
Axiom & Free-Parameter Ledger
free parameters (4)
- Affine calibration map parameters B, d =
not fully reported; σ_min ≈ 1.89 (GPT-4.1-mini), 1.85 (GPT-4o-mini); affine residual RMS ≈ 7.62 and 7.35
- Conformal error radius =
0.1371 (GPT-4.1-mini), 0.0922 (GPT-4o-mini)
- Semantic partition and candidate phrase set =
none (hand-specified before evaluation)
- Completeness threshold for constrained enum law =
0.999999
axioms (7)
- domain assumption Reference validity: the evidence generator, ontology, prior, likelihoods, and reference-inference procedure are declared and frozen; the posterior is exact conditional on this model (Assumption 1).
- domain assumption Observable measurement: only presented evidence, service-supplied token probabilities, and complete-continuation probabilities enter the estimator; if the service hides needed probabilities, the measurement is marked unsupported (Assumption 2).
- domain assumption Design separation: ontology discovery, semantic validation, calibration, model selection, conformal calibration, and final testing use disjoint scenario partitions (Assumption 3).
- domain assumption Sampling units and exchangeability: scenarios, not repeated responses, are the sampling units; calibration and future test nonconformity scores are exchangeable within strata (Assumption 4).
- ad hoc to paper In the experiments h_u is assumed affine (H_aff) with a fitted positive smallest singular value; uniform nonlinear observability is not certified.
- standard math Markov-kernel measurability and pushforward existence for the semantic coarsening (Lemma 1).
- standard math Filtering theorem assumptions: strictly positive transition, positive likelihoods, and Birkhoff contraction coefficient ρ<1 (Theorem 4).
invented entities (2)
-
Semantic map (φ_u, ψ_u)
independent evidence
-
Unexpressed-mass category ⊥
no independent evidence
read the original abstract
Language models produce probabilities over words, but professional decisions require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. A model's printed numerical confidence does not establish reliability. We introduce a semantic map: a prespecified, testable bridge from probabilities over verbal responses to probabilities over declared states, formulated as semiparametric inference for a finite-valued latent state. A reference model defines the target posterior, the language model supplies an unrestricted conditional distribution over verbal responses, and held-out calibration connects them. We derive posterior-error bounds and conditions for existence, uniqueness, stability and sequential Bayesian updating. Crucially, language probabilities depend on the prompt's lexical form, whereas the target posterior is unchanged by information-equivalent rewording. We test the method on professional market text compiled from Federal Reserve economic and financial series and on controlled simulations with exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical probabilities, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrasing, and respond appropriately to altered evidence. The broader implication is that prompt engineering optimizes a wording-dependent response, whereas scientific and professional use requires validated stability of application-relevant meaning. The semantic map turns this general concern into a testable statistical problem and, when its acceptance conditions hold, yields an auditable posterior estimate. The same principle offers a template for auditing classifications, recommendations and other fluent responses that may conceal semantic instability.
Figures
Reference graph
Works this paper leans on
-
[1]
2021 , doi=
Kallenberg, Olav , title=. 2021 , doi=
2021
-
[2]
The Annals of Mathematical Statistics , volume=
Blackwell, David , title=. The Annals of Mathematical Statistics , volume=. 1953 , doi=
1953
-
[3]
The Annals of Mathematical Statistics , volume=
Le Cam, Lucien , title=. The Annals of Mathematical Statistics , volume=. 1964 , doi=
1964
-
[4]
1991 , doi=
Torgersen, Erik , title=. 1991 , doi=
1991
-
[5]
and Hanke, Martin and Neubauer, Andreas , title=
Engl, Heinz W. and Hanke, Martin and Neubauer, Andreas , title=. 1996 , doi=
1996
-
[6]
Transactions of the American Mathematical Society , volume=
Birkhoff, Garrett , title=. Transactions of the American Mathematical Society , volume=. 1957 , doi=
1957
-
[7]
Bushell, P. J. , title=. Archive for Rational Mechanics and Analysis , volume=. 1973 , doi=
1973
-
[8]
Probability Theory and Related Fields , volume=
van Handel, Ramon , title=. Probability Theory and Related Fields , volume=. 2009 , doi=
2009
-
[9]
The Annals of Applied Probability , volume=
van Handel, Ramon , title=. The Annals of Applied Probability , volume=. 2009 , doi=
2009
-
[10]
and Mendelson, Shahar , title=
Bartlett, Peter L. and Mendelson, Shahar , title=. Journal of Machine Learning Research , volume=. 2002 , url=
2002
-
[11]
Transactions of the Association for Computational Linguistics , volume=
Jiang, Zhengbao and Araki, Jun and Ding, Haibo and Neubig, Graham , title=. Transactions of the Association for Computational Linguistics , volume=. 2021 , doi=
2021
-
[12]
arXiv preprint arXiv:2207.05221 , year=
Kadavath, Saurav and Conerly, Tom and Askell, Amanda and others , title=. arXiv preprint arXiv:2207.05221 , year=
-
[13]
Tian, Katherine and Mitchell, Eric and Zhou, Allan and Sharma, Archit and Rafailov, Rafael and Yao, Huaxiu and Finn, Chelsea and Manning, Christopher D. , title=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=. 2023 , publisher=. doi:10.18653/v1/2023.emnlp-main.330 , url=
-
[14]
International Conference on Learning Representations , year=
Kuhn, Lorenz and Gal, Yarin and Farquhar, Sebastian , title=. International Conference on Learning Representations , year=. doi:10.48550/arXiv.2302.09664 , url=
-
[15]
Nature , volume=
Farquhar, Sebastian and Kossen, Jannik and Kuhn, Lorenz and Gal, Yarin , title=. Nature , volume=. 2024 , doi=
2024
-
[16]
, title=
Gneiting, Tilmann and Raftery, Adrian E. , title=. Journal of the American Statistical Association , volume=. 2007 , doi=
2007
-
[17]
Philip , title=
Dawid, A. Philip , title=. Journal of the American Statistical Association , volume=. 1982 , doi=
1982
-
[18]
Philip , title=
Dawid, A. Philip , title=. Journal of the Royal Statistical Society: Series A , volume=. 1984 , doi=
1984
-
[19]
and Fienberg, Stephen E
DeGroot, Morris H. and Fienberg, Stephen E. , title=. The Statistician , volume=. 1983 , doi=
1983
-
[20]
, title=
Guo, Chuan and Pleiss, Geoff and Sun, Yu and Weinberger, Kilian Q. , title=. Proceedings of the 34th International Conference on Machine Learning , series=. 2017 , publisher=
2017
-
[21]
and Bates, Stephen , title=
Angelopoulos, Anastasios N. and Bates, Stephen , title=. Foundations and Trends in Machine Learning , volume=. 2023 , doi=
2023
-
[22]
2005 , doi=
Vovk, Vladimir and Gammerman, Alex and Shafer, Glenn , title=. 2005 , doi=
2005
-
[23]
and Wasserman, Larry , title=
Lei, Jing and G'Sell, Max and Rinaldo, Alessandro and Tibshirani, Ryan J. and Wasserman, Larry , title=. Journal of the American Statistical Association , volume=. 2018 , doi=
2018
-
[24]
Beyond Quantification: Navigating Uncertainty in Professional AI Systems , journal=
Delacroix, Sylvie and Robinson, Diana and Bhatt, Umang and Domenicucci, Jacopo and Montgomery, Jessica and Varoquaux, Ga. Beyond Quantification: Navigating Uncertainty in Professional AI Systems , journal=. 2025 , doi=
2025
-
[25]
Journal of the Royal Statistical Society: Series B , volume=
Aitchison, John , title=. Journal of the Royal Statistical Society: Series B , volume=. 1982 , doi=
1982
-
[26]
Isometric Logratio Transformations for Compositional Data Analysis , journal=
Egozcue, Juan Jos. Isometric Logratio Transformations for Compositional Data Analysis , journal=. 2003 , doi=
2003
-
[27]
Advances in Applied Probability , volume=
Brandt, Andreas , title=. Advances in Applied Probability , volume=. 1986 , doi=
1986
-
[28]
The Annals of Probability , volume=
Bougerol, Philippe and Picard, Nico , title=. The Annals of Probability , volume=. 1992 , doi=
1992
-
[29]
SIAM Review , volume=
Diaconis, Persi and Freedman, David , title=. SIAM Review , volume=. 1999 , doi=
1999
-
[30]
and Adeli, Ehsan and others , title=
Bommasani, Rishi and Hudson, Drew A. and Adeli, Ehsan and others , title=. arXiv preprint arXiv:2108.07258 , year=
-
[31]
Transactions on Machine Learning Research , year=
Lin, Stephanie and Hilton, Jacob and Evans, Owain , title=. Transactions on Machine Learning Research , year=
-
[32]
Dixon, Matthew , title=
-
[33]
The Annals of Applied Statistics , volume=
Su, Weijie , title=. The Annals of Applied Statistics , volume=. 2026 , doi=
2026
-
[34]
, title=
Wang, Ziyu and Holmes, Christopher C. , title=. Proceedings of the 28th International Conference on Artificial Intelligence and Statistics , series=. 2025 , url=
2025
-
[35]
2025 , type=
Humlum, Anders and Vestergaard, Emilie , title=. 2025 , type=
2025
-
[36]
Journal of the Royal Statistical Society: Series A (Statistics in Society) , volume=
Chatfield, Chris , title=. Journal of the Royal Statistical Society: Series A (Statistics in Society) , volume=. 1995 , doi=
1995
-
[37]
arXiv preprint arXiv:2404.03163 , year=
Huang, Xinmeng and Li, Shuo and Yu, Mengxin and Sesia, Matteo and Hassani, Hamed and Lee, Insup and Bastani, Osbert and Dobriban, Edgar , title=. arXiv preprint arXiv:2404.03163 , year=
-
[38]
Grewal, Yashvir S. and Bonilla, Edwin V. and Bui, Thang D. , title=. arXiv preprint arXiv:2410.22685 , year=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.