REVIEW 3 major objections 4 minor 1 cited by
Linearity-Inducing Priors for Poisson Parameter Estimation Under $L^{1}$ Loss
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Under Poisson noise, any sufficiently regular increasing function can be realized as the conditional median by a suitably chosen prior, and affine medians arise from priors beyond the gamma family.
desk verdict Novel construction, correct finite-dimensional piece, but Lemma 3's tail bound uses a moment equality that doesn't hold on the support, so the main theorem is unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the moment system $E[W^y 1_{W\le f(y)}]=E[W^y 1_{W> f(y)}]$ for all $y\in\mathbb{N}_0$, together with the exponential reweighting $dP_X(w)=e^w dP_W(w)$ that turns those equalities into the conditional-median condition. The construction solves truncated systems using an invertible matrix whose rows encode the split of the support at $f(y)$, obtains a tail bound of the form $P[W\ge w_{i+1}]\le (f(\lfloor i/\kappa \rfloor)/f(i))^{\lfloor i/\kappa \rfloor}$, and then uses $\ell^1$ compactness to extract a limiting probability vector. The summability condition (9) is exactly what makes the reweighted prior $P_X$ have finite total mass.
What would settle it
Solve the truncated linear system $A_M p = e_1$ for $f(y)=ay+b$ with $a=b=0.3$ and several values of $M$, then compare the exact tail probability $P[W\ge w_{i+1}]$ with the claimed bound $(f(\lfloor i/\kappa \rfloor)/f(i))^{\lfloor i/\kappa \rfloor}$; if the bound fails in exact arithmetic for some $i$, the equality step behind Lemma 3 is false, while if it holds across many $M$ and $i$, the missing atom term is likely controllable.
Extended reading notes
Core claim
The central claim is that prescribing a conditional median under Poisson noise is equivalent to solving the moment system $E[W^y 1_{W\le f(y)}]=E[W^y 1_{W> f(y)}]$ on an auxiliary random variable $W$ supported on $\{w_i=f(i-1)\}$. Once a probability law $P_W$ satisfies these equalities for every $y$, the reweighted measure $dP_X(w)=e^w dP_W(w)$ is a prior whose Poisson posterior median is exactly $f(y)$. The paper constructs such a $P_W$ by solving finite truncated linear systems in the atom probabilities, using an induction lemma for the existence of solutions, bounding the tail mass with a Markov-type inequality, and then taking a subsequential $\ell^1$ limit that preserves all moment equalities. In the linear case the construction gives a non-gamma prior with $\mathrm{med}(X|Y=y)=ay+b$ for $0<a<1/e$ and $b>0$, and finite-$M$ approximations are shown numerically to converge to the target median.
Load-bearing premise
The construction depends on treating 'W is at least f(y)' as interchangeable with 'W is strictly greater than f(y)' in the moment equalities, but the two events differ by the probability mass placed exactly at f(y), and the proof does not show that this mass term is negligible.
Editorial extensions
If this is right
- For any increasing $f$ satisfying condition (9), a Poisson-noise experiment can be given a prior whose $L^1$-optimal estimator is exactly $f$, so estimator design becomes a direct choice of the estimator rather than a search over priors.
- Affine conditional medians with slope $0<a<1/e$ and intercept $b>0$ are attainable by non-gamma priors, giving the first explicit counterpart to the gamma-prior characterization for conditional means.
- The proof gives a concrete numerical recipe: solve the finite linear system $A_M p = e_1$ and reweight by $e^{w_i}$ to approximate the desired prior to any chosen accuracy.
- Because Poisson posterior medians are always nondecreasing, the monotonicity assumption on $f$ is not an artificial restriction, and the theorem covers every estimator shape that the noise model can in principle produce.
Reading between the lines
- A direct numerical test of the tail bound on the exact truncated solutions, say for $f(y)=0.3y+0.3$, would show whether the atom of mass placed exactly at $f(y)$ is negligible; this is the step the proof leaves implicit.
- The same moment-matching construction should transfer to other integer-valued observation models, such as binomial or negative-binomial noise, where conjugate priors also fail to produce affine conditional medians; only the summability condition would need reworking.
- The slope restriction $0<a<1/e$ is presented as sufficient, so the true admissible range may be wider; the finite-truncation solver could be used to probe larger slopes numerically.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper considers Bayesian Poisson parameter estimation under L^1 loss, where the optimal estimator is the conditional median. The authors claim that for any increasing function f satisfying a summability condition, there exists a prior P_X supported on {f(i-1)} such that med(X|Y=y)=f(y) for all y; specializing to f(y)=ay+b with 0<a<1/e yields a family of priors distinct from gamma that achieve an affine conditional median. The proof proceeds by finite-dimensional moment matching (Theorem 3), a tail bound (Lemma 3), and a compactness/limiting argument to pass to infinite support.
Significance. If correct, the main theorem would be a substantive contribution to Bayesian L^1 estimation: it would provide the first explicit non-conjugate family with the affine conditional-median property under Poisson noise and a general realizability result for arbitrary increasing conditional medians. The paper also gives an algorithm for truncated approximation and numerical illustrations. However, the proof as presented contains a central error in Lemma 3 that invalidates the limiting argument; the claimed results are therefore not established by this manuscript.
major comments (3)
- [Section III-B, Lemma 3, Eq. (40)] The equality in Eq. (40) is invalid. The moment condition (24) (or (35)) states E[W^y 1_{W≤f(y)}] = E[W^y 1_{W>f(y)}] for y=⌊i/κ⌋, but step (40) replaces the indicator 1_{W≥f(y)} with 1_{W≤f(y)}. Because the support contains w_{y+1}=f(y), the two events differ by the atom {W=f(y)}. For the affine example f(i)=0.3i+0.3, M=3, y=1, i=3, κ=2, the probability vector from Theorem 3 has p_2=3/13, and the two sides of (40) are approximately 0.356 and 0.240, so the equality is false. Consequently the tail bound (36) is not proven.
- [Section III-C and Section III-D] The unproven tail bound (36) is the load-bearing component of the infinite extension. It is invoked at (42) to make the truncated vectors tight, at (47) and (50) to show the limiting total mass is one, at (58)-(60) to control the residual Δ in the moment transfer, and at (63) to prove that P_X is a proper distribution. Since (36) is not established, the ℓ1 compactness argument and the conclusion of Theorem 2 do not follow; consequently Theorems 1 and Corollary 1 are unsupported.
- [Section III-C, Eq. (49)] With the indexing w_i=f(i-1), the identity sum_{i=1}^{j+1} P_W(w_i) = 1 - P(W≥w_{j+1}) is off by one; the correct relation is sum_{i=1}^{j+1} P_W(w_i) = 1 - P(W≥w_{j+2}). The subsequent bound (50) should use a tail bound for w_{j+2}. This is secondary to the failure of Lemma 3 but suggests that the indexing in the limit arguments needs careful revision.
minor comments (4)
- [Theorem 2] Theorem 2 states f:N0→R0; this should be f:N0→R+.
- [Lemma 1] In Lemma 1, the measure dP_X(w)=e^w dP_W(w) is defined up to normalization; the paper should explicitly state that P_X is the normalized version and that the normalization constant cancels in the posterior ratio.
- [Corollary 1 proof] The proof of Corollary 1 writes the sum in (14) with an inf over κ≥1 without explaining why the inf appears; the convergence argument should be expanded.
- [Figure 3] Figure 3 plots conditional medians for truncated approximations; the caption should clarify that these are finite-M approximations and not the limiting prior distribution.
Circularity Check
No circularity: the prior is constructed by solving moment equations equivalent to the stated median, and self-citations are contextual only.
full rationale
Reviewing the derivation chain, I find no circular step. The target function f enters as a theorem input: the support W is deliberately chosen as {f(i-1)}, but existence of a probability vector solving the moment equations (24)/(35) is a genuine linear-algebraic result (Lemma 2/Theorem 3), not an assumption of the conclusion. Lemma 1 reduces the median condition to equivalent moment equations; the paper then solves those equations and converts back via (16). This is a standard inverse-problem construction, not a fitted input renamed as prediction. The summability condition (9) is exactly what is needed to make the limiting measure and the reweighted PX proper (Eqs. (62)-(63)), so its use is a sufficient-condition check, not circularity. Self-citations [9] and [12]-[14] are contextual (gamma-mean characterization and prior linear-median work) and do not carry the Poisson median construction. The skeptical issue with Lemma 3, Eq. (40), is a potential proof error: condition (24) equates the indicators {W≤f(⌊i/κ⌋)} and {W>f(⌊i/κ⌋)}, not {W≥f(⌊i/κ⌋)}, so the atom at w_{⌊i/κ⌋+1}=f(⌊i/κ⌋) may make the displayed equality false. That is a correctness gap in the submitted argument, not a circularity: the paper does not assume the tail bound, it attempts to derive it from the moment system. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- standard math The conditional median, defined through the quantile function, is the minimizer of expected absolute error (Eq. 1).
- domain assumption The Poisson likelihood with the 0^0=1 convention (Eq. 2) is the noise model.
- domain assumption Bayes estimators under Poisson noise are nondecreasing in y (cited to [17]), which motivates restricting f to increasing functions.
- standard math Sequential compactness and completeness of probability vectors in l1, used to extract the limiting distribution in Section III-C.
Cite this review
Pith. "Pith review of Linearity-Inducing Priors for Poisson Parameter Estimation Under $L^{1}$ Loss." pith.science (2026). https://pith.science/paper/ICIN5PLH
@misc{pith2026250521102,
author = {Pith},
title = {Pith review of: Linearity-Inducing Priors for Poisson Parameter Estimation Under $L^1$ Loss},
year = {2026},
howpublished = {\url{https://pith.science/paper/ICIN5PLH}},
note = {Machine review of arXiv:2505.21102}
}
abstract
We study prior distributions for Poisson parameter estimation under $L^1$ loss. Specifically, we construct a new family of prior distributions whose optimal Bayesian estimators (the conditional medians) can be any prescribed increasing function that satisfies certain regularity conditions. In the case of affine estimators, this family is distinct from the usual conjugate priors, which are gamma distributions. Our prior distributions are constructed through a limiting process that matches certain moment conditions. These results provide the first explicit description of a family of distributions, beyond the conjugate priors, that satisfy the affine conditional median property; and more broadly for the Poisson noise model they can give any arbitrarily prescribed conditional median.
Figures
Forward citations
Cited by 1 Pith paper
-
Functional uniqueness and stability of Gaussian priors in optimal L1 estimation
Near-linearity of the optimal L1 median is claimed to force the prior close to Gaussian in Hermite coefficients, under strong decay assumptions; the L2 stability rate is re-derived.
Reference graph
Works this paper leans on
-
[1]
The practical limits of photon communication,
R. McEliece, E. Rodemich, and A. Rubin, “The practical limits of photon communication,” Jet Propulsion Laboratory Deep Space Network Progress Reports, vol. 42, pp. 63–67, 1979
work page 1979
-
[2]
Capacity of a pulse amplitude modulated direct detection photon channel,
S. Shamai, “Capacity of a pulse amplitude modulated direct detection photon channel,” IEE Proceedings I (Communications, Speech and Vision), vol. 137, no. 6, pp. 424–430, 1990
work page 1990
-
[3]
S. Verdú, “Poisson communication theory,” International Technion Com- munication Day in Honor of Israel Bar-David , vol. 66, 1999
work page 1999
-
[4]
On the capacity of the discrete-time poisson channel,
A. Lapidoth and S. M. Moser, “On the capacity of the discrete-time poisson channel,” IEEE Transactions on Information Theory , vol. 55, no. 1, pp. 303–322, 2008
work page 2008
-
[5]
A. Dytso, L. Barletta, and S. Shamai (Shitz), “Properties of the support of the capacity-achieving distribution of the amplitude-constrained Pois- son noise channel,” IEEE Transactions on Information Theory , vol. 67, no. 11, pp. 7050–7066, 2021
work page 2021
-
[6]
Capacities and optimal input distributions for particle-intensity channels,
N. Farsad, W. Chuang, A. Goldsmith, C. Komninakis, M. Médard, C. Rose, L. Vandenberghe, E. E. Wesel, and R. D. Wesel, “Capacities and optimal input distributions for particle-intensity channels,” IEEE Transactions on Molecular, Biological and Multi-Scale Communications, vol. 6, no. 3, pp. 220–232, 2020
work page 2020
-
[7]
A theoretical analysis of neuronal variability,
R. B. Stein, “A theoretical analysis of neuronal variability,” Biophysical Journal, vol. 5, no. 2, pp. 173–194, 1965
work page 1965
-
[8]
M. N. Shadlen and W. T. Newsome, “The variable discharge of cortical neurons: Implications for connectivity, computation, and information coding,” Journal of Neuroscience, vol. 18, no. 10, pp. 3870–3896, 1998
work page 1998
Show all 17 references
-
[9]
Estimation in Poisson noise: Properties of the conditional mean estimator,
A. Dytso and H. V . Poor, “Estimation in Poisson noise: Properties of the conditional mean estimator,” IEEE Transactions on Information Theory, vol. 66, no. 7, pp. 4304–4323, 2020
2020
-
[10]
Conjugate priors for exponential fami- lies,
P. Diaconis and D. Ylvisaker, “Conjugate priors for exponential fami- lies,” The Annals of Statistics , vol. 7, no. 2, pp. 269–281, 1979
1979
-
[11]
Characterization of conjugate priors for discrete exponential families,
J.-P. Chou, “Characterization of conjugate priors for discrete exponential families,” Statistica Sinica, vol. 11, pp. 409–418, 2001
2001
-
[12]
L1 estimation in Gaussian noise: On the optimality of linear estimators,
L. P. Barnes, A. Dytso, and H. V . Poor, “ L1 estimation in Gaussian noise: On the optimality of linear estimators,” in Proceedings of the 2023 IEEE International Symposium on Information Theory (ISIT), June 2023, pp. 1872–1877
2023
-
[13]
L1 estimation: On the optimality of linear estimators,
L. P. Barnes, A. Dytso, J. Liu, and H. V . Poor, “ L1 estimation: On the optimality of linear estimators,” IEEE Transactions on Information Theory, vol. 70, no. 11, pp. 8026–8039, 2024
2024
-
[14]
Multivariate priors and the linearity of optimal bayesian es- timators under Gaussian noise,
——, “Multivariate priors and the linearity of optimal bayesian es- timators under Gaussian noise,” in Proceedings of the 2024 IEEE International Symposium on Information Theory (ISIT) , July 2024, to appear
2024
-
[15]
On conditions for linearity of optimal estimation,
E. Akyol, K. Viswanatha, and K. Rose, “On conditions for linearity of optimal estimation,” IEEE Transactions on Information Theory , vol. 58, no. 6, pp. 3497–3508, 2012
2012
-
[16]
Bounds for the difference between median and mean of gamma and Poisson distributions,
J. Chen and H. Rubin, “Bounds for the difference between median and mean of gamma and Poisson distributions,” Statistics & Probability letters, vol. 4, no. 6, pp. 281–283, 1986
1986
-
[17]
Monotonicity of Bayes estimators,
P. Nowak, “Monotonicity of Bayes estimators,” Applicationes Mathe- maticae, vol. 4, no. 40, pp. 393–404, 2013
2013
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.