{"id":"115cb471-4933-452e-b6b8-97538b54b1bb","arxiv_id":"2411.15499","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Asymmetric errors are given a consistent treatment by distinguishing pdf-based from likelihood-based uncertainties, and by providing a family of three-parameter models with validated combination rules and open-source software.","lead":"This paper lays out a consistent set of rules and software for handling measurements with different positive and negative uncertainties, a common situation in particle physics. It separates two meanings of an error, the spread of a measurement's probability distribution versus the confidence interval on a parameter, and shows how each should be propagated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model-family dependence is the load-bearing soft spot: the 'consistent handling' claim holds only for near-Gaussian moderate asymmetries, a scope the paper acknowledges but does not quantify.","rationale":"The reader correctly identifies the near-Gaussian three-parameter assumption as the weakest point. I agree: the strongest version of the paper's claim, namely that a three-number summary can be handled consistently once the pdf/likelihood distinction is made, is established only relative to a chosen model family. The paper is transparent about this caveat in Sections 2.2 and 3.1, and the software plus several exact-validation examples provide real support for the practical method. The remaining question is whether, for realistic non-near-Gaussian cases, model dependence is large enough to change conclusions; the proposed benchmark would settle that. Because the limitation is acknowledged rather than hidden, and because the paper's recommendations (report full likelihood for large asymmetries, use more than one model) already mitigate the risk, this concern does not change the ACCEPT verdict.","tokens_in":46274,"tokens_out":16937,"duration_ms":165604,"concrete_test":"Select true likelihoods that all yield the same quoted summary 5 +2.581 -1.916, for example Poisson(5), an appropriately shifted lognormal, a truncated normal, and a two-component mixture matched to the same 16/50/84 percent quantiles. For each, combine two independent copies by exact convolution or exact likelihood product and compare the exact combined summary with the linear-sigma and linear-variance model results. If the spread among exact combined summaries across the matched true distributions is comparable to or larger than the spread between the two recommended models, the model-dependence concern lands; if the recommended models track the exact answers within the 0.1% level seen in Section 4.1.3, the concern is resolved for this central case.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central procedure replaces the unknown pdf or log-likelihood behind a quoted R +sigma+ -sigma- with a three-parameter model, combines via cumulant addition or likelihood products, and reads off the combined three numbers. This is only a consistency guarantee within a chosen model family. The paper's own comparisons show the model dependence becomes material when asymmetries are large: Table 3 gives sigma+ ranging from 1.78 to 2.07 and sigma- from 0.97 to 1.17 for inputs 0.5/1.5 combined with 0.5/1.5, and Section 3.1 notes that with only three numbers 'a wide range of models can potentially give a wide range of outcomes.' The validation examples (Poisson, exponential, Gaussian transforms) are all smooth, single-peaked, near-Gaussian cases; they do not cover heavy-tailed, boundary-truncated, or mixture-like likelihoods that share the same three quoted numbers. In such cases the recommended linear-sigma/linear-variance pair can agree with each other while both being far from the exact combination, so the 'use two models' advice does not bound the error. The paper states this limitation explicitly (Section 2.2: 'for large asymmetries... the accuracy should not be considered as being definite'), so this is not an internal inconsistency, but it means the headline claim of a consistent procedure is really a heuristic with unquantified scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the practical problem of handling measurements quoted as R +σ+ −σ−. Its central contribution is a taxonomy: asymmetric errors can be pdf errors (rms spread of a probability distribution) or likelihood errors (68% central interval from ΔlnL = −1/2), and the operation can be combination of errors or combination of results. For each cell of this taxonomy the paper proposes families of three-parameter near-Gaussian models, derives conversions between quantile, moment, and model parameters, and gives algorithms for convolution of pdfs and for multiplication/profiling of likelihoods. The methods are validated in cases where the exact answer is known (Poisson counts, exponential lifetimes, transformed Gaussians) and are accompanied by open-source implementations in C++/Python and R. The paper also argues that the common recipe of adding positive and negative errors separately in quadrature is inconsistent with the Central Limit Theorem and that a shift in the central value is generally needed when combining asymmetric errors.","tokens_in":41,"tokens_out":12622,"duration_ms":256096,"significance":"If the claims hold, this is a genuinely useful contribution for experimental particle physics and metrology. The explicit separation of pdf errors from likelihood errors, the warning that separate-quadrature combination violates the CLT, and the demonstration that the median shifts under convolution are all valuable and likely to influence practice. The algebraic appendices are careful, the software is a concrete deliverable, and the validation against exact Poisson and exponential cases provides real evidence that the recommended Bartlett models work well for smooth, near-Gaussian, moderate asymmetries. The honest acknowledgement that no model is ‘correct’ and that large asymmetries require caution is a strength, not a defect. The main weakness is that the model-family dependence is acknowledged but not quantified: for large asymmetries the spread across models can be comparable to the quoted errors, and the paper does not give a threshold or a calibration for when the procedure should be trusted.","major_comments":[{"comment":"The displayed expression for w_i in the linear-variance model is algebraically wrong in the Gaussian limit. With V'_i = 0 the log likelihood is −(1/2)a_i^2/V_i, so ∂lnL_i/∂a_i = −a_i/V_i and the definition w_i = −(a_i^{-1}∂lnL_i/∂a_i)^{-1} gives w_i = V_i. The printed formula gives V_i/2. For V'_i ≠ 0 the correct local weight is 2(V_i + a_i V'_i)^2/(2V_i + a_i V'_i), not (V_i + a_i V'_i)^2/(2V_i + a_i V'_i). The missing factor cancels in the purely parabolic case, so the symmetric Gaussian examples are unaffected, but it does not cancel in general and changes the profile path for asymmetric inputs. Please correct the formula and confirm that the numerical results in Section 4.1.1 were produced with the corrected weight.","section":"Section 3.4, Eq. (13)"},{"comment":"The central claim of a ‘consistent procedure’ is model-dependent in an unquantified way. Table 3 shows that for inputs σ_− = 0.5, σ_+ = 1.5 combined with 0.5/1.5, the predicted combined σ_+ ranges roughly from 1.93 to 2.07 (2.42 for the log-normal model) and σ_− from 0.91 to 1.13 across models that can represent the asymmetry. Section 3.1 similarly states that ‘a wide range of models can potentially give a wide range of outcomes.’ Because no calibration or validity domain is given, the recommendation to use two models does not by itself bound the true answer: the models can agree with each other while both being far from the exact combination for skewed, boundary-truncated, or mixture-like likelihoods. The manuscript should either quantify the asymmetry range within which model spread is below a stated tolerance, add a coverage study on non-smooth cases, or explicitly present the method as an approximate heuristic with a defined scope.","section":"Sections 2.2 and 3.1; Table 3"},{"comment":"The Poisson validation is less convincing than the 5+5 row suggests. For two samples split as 8+2, the linear-variance method gives 5.054 +1.856/−1.516 instead of the exact 5.000 +1.752/−1.419, and for 9+1 it gives 5.201 +1.942/−1.605 instead of 5.000 +1.752/−1.419. These deviations are much larger than the 5+5 case and are not negligible for published Poisson measurements, even if a 9+1 split is not the most likely outcome for mean 5. The text dismisses these rows as ‘unlikely experimental circumstances’ with ‘poor goodness of fit,’ but the paper does not provide a threshold for declaring a poor fit or a fallback procedure when the fit is poor. Please state the conditions under which the ‘excellent match’ claim holds and show the corresponding goodness-of-fit values for the rows in Table 10.","section":"Section 4.1.3, Table 10"}],"minor_comments":[{"comment":"Reference [9] contains a typo in the arXiv identifier (‘physis’ instead of ‘physics’) and the citation ‘Merkat A Possolo and O Biodnar’ is inconsistently formatted; please correct the bibliography entries.","section":"References"},{"comment":"The two inputs in each block of Table 4 are assigned the same asymmetric errors even though their central values differ; the caption should explain that this happens because the underlying Gaussian is the same and only the sampled x value changes.","section":"Table 4"},{"comment":"The code example labelled ‘Likelihood and Pdf Errors’ refers to a ‘combined results of Example 5.3’ but the section number corresponds to Section 5.3, and the code block contains a duplicated phi/plo pair and a ‘readline’ prompt that may confuse readers; please check the displayed code against the released package.","section":"Appendix D"},{"comment":"The statement that the models ‘coincide’ for moderate asymmetries in the upper rows of Figures 6 and 7 is only qualitative; a sentence giving a numerical tolerance (e.g., agreement in the 68% interval to two significant figures) would make the claim easier to verify.","section":"Figures 6 and 7"}],"recommendation":"major_revision","confidential_remarks":"This is a long, encyclopaedic paper that will be useful to practitioners, but it needs a correction to Eq. (13) and a more honest statement of its scope. The model-dependence issue is real: Table 3 and Table 10 both show cases where model choice changes the combined error by more than the paper’s narrative implies. I would be comfortable with acceptance after the formula is fixed and the validity domain is quantified or explicitly hedged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious paper and it deserves a serious referee. It is the most complete treatment I know of the x +sigma+ -sigma- problem, and the authors did the work that makes a methods paper usable: a catalogue of model families, explicit combination formulas, the algebra in appendices, and working code in C++/Python and R.\n\nWhat is genuinely new: a systematic taxonomy, pdf errors versus likelihood errors crossed with combination of errors versus combination of results, turned into a decision procedure, plus several new three-parameter families (railway Gaussian, double cubic Gaussian, symmetric beta Gaussian, molded quartic, conservative spline). The pdf/likelihood distinction itself is not new; it appears in the authors' earlier papers and in d'Agostini. The new part is the catalogue, the formulas, and the software. The validation is the strongest section: combining the Poisson result 5 +2.581/-1.916 with itself through the linear variance model gives 5.000 +1.748/-1.415 against the exact 5.000 +1.752/-1.419, and the lifetime example (Section 4.1.2) agrees with the full-information answer at the 0.1% level. Those are independent benchmarks, not circular fits. The argument that separate-in-quadrature addition violates the Central Limit Theorem, with the 'wrong' row in Table 9 as evidence, is convincing.\n\nSoft spots, in proportion. Model-family dependence is the real limitation. For strongly asymmetric inputs, the same three numbers produce materially different combined errors across models, with Table 3 showing sigma+ ranging from 1.78 to 2.07 for the most asymmetric row, and the paper states this explicitly (Section 2.2: 'for large asymmetries... the accuracy should not be considered as being definite'; Section 3.1: 'a wide range of available models can potentially give a wide range of outcomes'). The advice to use two models checks robustness within the chosen families; it does not bound the error against an unknown true shape such as a heavy-tailed, truncated, or mixture likelihood. That makes the headline claim a heuristic with unquantified scope, but it is the paper's own caveat, not a defect the authors hid. The Section 5.3 conclusion about mixing statistical and systematic errors rests on one toy example; the 10^7 coverage study is reassuring, but it is a single case, so that passage should be read as a caution, not a license. The independence assumption (Section 6) is likewise explicit.\n\nWho this is for: particle physicists and metrologists who deal with this problem daily, plus statisticians looking for testable examples. As a formal contribution it is deliberately heuristic, but it is careful, honest, checked against exact cases, and ships with code. That clears the bar for serious refereeing. I would bring it to a reading group and would cite it in my own work.\n\nRecommendation: send it to referees. The main thing I would ask them to push on is quantifying the model-dependence, but the paper is already honest about where it cannot answer.","headline":"A genuinely useful, software-backed framework for asymmetric errors, honest about its own scope; the model-dependence caveat is real but the paper does not oversell what it cannot know.","tokens_in":47123,"tokens_out":5037,"would_cite":true,"duration_ms":40477,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62F10","62P35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Asymmetric measurement errors are four distinct problems, and the usual quadrature fix is wrong.","keywords":["asymmetric errors","pdf errors","likelihood errors","combination of errors","combination of results","dimidiated Gaussian","linear variance model","Central Limit Theorem"],"falsifier":"Take a measurement whose true distribution is strongly skewed or bimodal (for example a chi-squared variable with one degree of freedom, or a mixture of two well-separated Gaussians), quote it in the form $R\\,{+\\sigma_+\\atop-\\sigma_-}$, combine several independent copies with the paper's recommended linear-variance and dimidiated models, and compare with the exact convolution or product likelihood. If the model-based 68% intervals miss the exact intervals by more than the spread between the two recommended models, the paper's claim that two models suffice for a robustness check is refuted; the paper itself passes this test for a Poisson with mean 5.","tokens_in":45983,"feed_emoji":"📊","tokens_out":13527,"duration_ms":108393,"temperature":0.7,"pith_summary":"Results quoted as $R\\,{+\\sigma_+\\atop -\\sigma_-}$ are ambiguous until two questions are answered: does $\\sigma$ describe the rms spread of a probability distribution ('pdf error') or the 68% confidence region of a likelihood ('likelihood error'), and is the task to combine errors from different quantities or to combine independent results for the same quantity? The paper argues that once these two dichotomies are fixed there are consistent recipes: for pdf errors one adds means, variances, and skewnesses under convolution and converts the summed moments back to the quoted form; for likelihood errors one models each log-likelihood with a variable-width Gaussian and sums the log-likelihoods. It further argues that the widespread habit of adding positive errors in quadrature and negative errors in quadrature separately is not merely crude but wrong, because it keeps the distribution's shape fixed while the Central Limit Theorem forces it toward a symmetric Gaussian. A sympathetic reader would care because particle-physics results are routinely published in this notation, and the paper's recipes turn those numbers into reproducible combined values, with an explicit warning that if the asymmetry is large, no three-parameter model is reliable and the full likelihood should be reported.","feed_headline":"Asymmetric errors: four recipes, one common fix is wrong","feed_subtitle":"Whether to add moments or add log-likelihoods depends on what the error means; separate quadrature violates the Central Limit Theorem.","key_machinery":"The load-bearing objects are two families of three-parameter, near-Gaussian models. For pdf errors, the paper uses distributions such as the dimidiated Gaussian (two half-Gaussians from a one-parameter-at-a-time, 'OPAT', systematic variation approximated by two straight lines) and the distorted Gaussian (a parabolic OPAT dependence), whose first three moments — mean, variance, and unnormalised skewness — add exactly under convolution; errors are combined by summing moments and then converting back to the quantile parameters $M\\,{+\\sigma_+\\atop-\\sigma_-}$. For likelihood errors, the machinery is the variable-width Gaussian log-likelihood, in either the linear-$\\sigma$ form $\\ln L(a)=-\\tfrac12[(a-\\hat a)/(\\sigma+\\sigma'(a-\\hat a))]^2$ or the linear-variance form $\\ln L(a)=-\\tfrac12(a-\\hat a)^2/[V+V'(a-\\hat a)]$, with parameters fixed by the three quoted points; results are combined by summing such log-likelihoods, the combined value being found by the iterative weighted equations (11) and (12), and the combined errors by root-finding where $\\Delta\\ln L=-\\tfrac12$.","core_discovery":"The paper's central claim is that a measurement quoted as $R\\,{+\\sigma_+\\atop-\\sigma_-}$ is not a single kind of object, and that the apparent lack of a consistent procedure comes from conflating four cases: pdf versus likelihood errors, and combination of errors versus combination of results. Under pdf errors the quoted $\\sigma_\\pm$ are properties of a probability density, so combining errors means convolving densities and adding moments; under likelihood errors the quoted $\\sigma_\\pm$ are the $\\Delta\\ln L=-\\tfrac12$ points of a log-likelihood, so combining results means multiplying likelihoods and finding the peak of the sum of log-likelihoods. A directly testable embodiment is the Poisson example: combining the measurement $5\\,{+2.581\\atop-1.916}$ with itself using the linear-variance likelihood model yields $5.000\\,{+1.748\\atop-1.415}$, close to the exact combined answer $5.000\\,{+1.752\\atop-1.419}$. The paper also claims the common recipe of adding the $\\sigma_+$ values in quadrature separately from the $\\sigma_-$ values is wrong because it preserves shape under many additions and therefore contradicts the Central Limit Theorem.","pith_inferences":["If the four-way classification is accepted, publication practice should change: a result quoted as $R\\,{+\\sigma_+\\atop-\\sigma_-}$ is under-specified, and collaborations or journals should state whether the errors are pdf or likelihood (ideally supplying the full likelihood), otherwise a meta-analyst must guess.","The same recipes should transfer to any field quoting asymmetric uncertainties — metrology, astronomy, economics — and the coverage of the two-model interval could be tested on known skewed distributions (Poisson, chi-square, log-normal) to calibrate how wide the model family really is.","A practical decision rule suggested by the paper's large-asymmetry warnings, though not stated as such, is to refuse three-number summaries when $\\sigma_+$ and $\\sigma_-$ differ by more than roughly a factor of two or three; in that regime the spread between models is the dominant uncertainty.","The paper's linear-variance success on Poisson examples suggests a testable extension: for counting experiments, replacing each asymmetric result by a generalised-Poisson log-likelihood with the same peak and errors may give near-exact meta-analysis even for very small counts, without needing full likelihoods."],"forward_implications":["Combining systematic uncertainties should be done as pdf errors: convert each $\\sigma_\\pm$ to moments, add the moments, convert back — the central value shifts, and the asymmetry shrinks as more independent errors are combined.","Combining best-measurement results should be done as likelihood errors: model each quoted peak and error with a linear-sigma or linear-variance log-likelihood, sum the log-likelihoods, and read the combined value and the 68% errors from the summed curve.","The usual 'add positive errors in quadrature, add negative errors in quadrature, then quote a split Gaussian' recipe is not a small approximation error; it is structurally wrong and can be seriously discrepant, as in the lifetime combination example.","Whenever the asymmetry is even moderate, using at least two different models (e.g., dimidiated and distorted for pdfs, linear sigma and linear variance for likelihoods) is necessary to know how much the answer can be trusted; for large asymmetries the models disagree and no definite accuracy is claimed.","If both OPAT deviations go the same way ('flipped'), the paper's advice is to replace the flipped result by a moment-matched ordinary dimidiated distribution rather than build a special flipped model, unless the effect is important enough to demand a full reanalysis."],"supporting_citations":[{"why":"Supplies the variable-width Gaussian ('linear sigma') likelihood form used to model asymmetric log-likelihoods.","marker":"[16]"},{"why":"Second source for the same variable-width Gaussian likelihood construction underlying the linear-variance variant.","marker":"[17]"},{"why":"Provides the Delta ln L = -1/2 error definition and the argument for omitting the ln sigma normalization in variable-width likelihood models.","marker":"[13]"},{"why":"Supplies the confidence-belt construction that distinguishes 68% central confidence regions (likelihood errors) from rms spreads (pdf errors).","marker":"[12]"},{"why":"The piecewise-linear-sigma combination recipe the paper treats as the standard to compare against and to correct.","marker":"[28]"},{"why":"Earlier treatment of asymmetric errors with the distorted-Gaussian model and with bias from estimated variances, which the paper extends and tests.","marker":"[6]"},{"why":"General discussion of sources and dangers of asymmetric uncertainties that motivates the paper's two-by-two classification.","marker":"[9]"},{"why":"Documents the bootstrap confidence-interval caveat used to warn against treating pdf spreads as likelihood confidence regions.","marker":"[19]"}],"fun_headline_variants":["Asymmetric errors: don't add in quadrature, use the right recipe","Asymmetric errors: PDF vs likelihood changes how you combine","Four recipes for asymmetric errors, one common fix is wrong"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole machinery rests on the assumption that the true distribution behind any quoted asymmetric error is well approximated by a three-parameter, single-peaked, near-Gaussian family, and that the errors being combined are statistically independent; the paper itself notes that the models are not 'correct' and lose reliability for large asymmetries.","fun_headline_variants_meta":{"raw":{"variants":["Asymmetric errors: don't add in quadrature, use the right recipe","Asymmetric errors: PDF vs likelihood changes how you combine","Four recipes for asymmetric errors, one common fix is wrong"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000377,"raw_usage":{"total_tokens":1963,"prompt_tokens":859,"completion_tokens":1104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":1047}},"tokens_in":475,"tokens_out":1104,"duration_ms":7849,"temperature":1.0,"reasoning_tokens":1047,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:13:30.171376+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a measurement whose true distribution is strongly skewed or bimodal (for example a chi-squared variable with one degree of freedom, or a mixture of two well-separated Gaussians), quote it in the form $R\\,{+\\sigma_+\\atop-\\sigma_-}$, combine several independent copies with the paper's recommended linear-variance and dimidiated models, and compare with the exact convolution or product likelihood. If the model-based 68% intervals miss the exact intervals by more than the spread between the two recommended models, the paper's claim that two models suffice for a robustness check is refuted; the paper itself passes this test for a Poisson with mean 5.","supporting_citations":[{"cited_title":"On the statistical estimation of mean lifetimes","cited_arxiv_id":null,"evidence_quote":"Supplies the variable-width Gaussian ('linear sigma') likelihood form used to model asymmetric log-likelihoods."},{"cited_title":"Estimation of mean lifetimes from multiple plate cloud chamber tracks","cited_arxiv_id":null,"evidence_quote":"Second source for the same variable-width Gaussian likelihood construction underlying the linear-variance variant."},{"cited_title":"A note on Delta ln L = -1/2 Errors","cited_arxiv_id":"physics/0403046","evidence_quote":"Provides the Delta ln L = -1/2 error definition and the argument for omitting the ln sigma normalization in variable-width likelihood models."},{"cited_title":"(Particle Data Group) Review of Particle Properties","cited_arxiv_id":null,"evidence_quote":"Supplies the confidence-belt construction that distinguishes 68% central confidence regions (likelihood errors) from rms spreads (pdf errors)."},{"cited_title":"Private communication, 2004","cited_arxiv_id":null,"evidence_quote":"The piecewise-linear-sigma combination recipe the paper treats as the standard to compare against and to correct."},{"cited_title":"Averaging Measurements with Hidden Correlations and Asymmetric Errors","cited_arxiv_id":"hep-ex/0006004","evidence_quote":"Earlier treatment of asymmetric errors with the distorted-Gaussian model and with bias from estimated variances, which the paper extends and tests."},{"cited_title":"Asymmetric uncertainties: Sources, treatment and potential dangers","cited_arxiv_id":null,"evidence_quote":"General discussion of sources and dangers of asymmetric uncertainties that motivates the paper's two-by-two classification."},{"cited_title":"The Bootstrap and Edgeworth Expansion","cited_arxiv_id":null,"evidence_quote":"Documents the bootstrap confidence-interval caveat used to warn against treating pdf spreads as likelihood confidence regions."}],"review_version":1}