{"id":"a75de154-9b24-4c01-84cc-d58016625787","arxiv_id":"2506.03500","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A gamma-like email-size distribution fits the observed 2015 data better than the previous log-normal-like model, but the gain is small, fitted on the same data, and not statistically confirmed.","lead":"This short paper replaces the log-normal assumption in an existing email-size model with a gamma distribution and reports a modestly better fit to the same 2015 email data. Because the data and code are not shared and the improvement is not tested for statistical significance, the result is plausible but unverified.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 'significantly improved fit' is not backed by any uncertainty estimate or significance test; the raw D-value gap could be sampling noise and the same data are used for fitting and evaluation.","rationale":"The reader's weakest_assumption focuses on the unvalidated gamma distribution for word lengths, which is a legitimate mechanistic concern. However, the more load-bearing issue for the paper's central claim is statistical: the direct and simulation D comparisons have no uncertainty quantification, no significance test, and use the same data for fitting and evaluation. Without a null distribution, the observed D differences cannot be claimed as 'significantly improved.' This does not prove the model is false; it means the central claim is not yet established. A conditional verdict remains appropriate: the authors should provide a significance test, bootstrap confidence intervals, or cross-validated comparison. The reader's rationale already mentions the absence of error bars and the same-data issue, so the disagreement is not about whether these are problems but about which is the single most load-bearing assumption. I would keep the CONDITIONAL verdict unchanged while sharpening the required check.","tokens_in":5482,"tokens_out":9486,"duration_ms":115421,"concrete_test":"Using the stated sample size (191,993 emails), simulate B=1000 synthetic datasets from the Section 4 fitted pLNL model, apply the same logarithmic binning, fit both pLNL and pGL to each simulated dataset by the same least-squares procedure, and record ΔD = D_LNL − D_GL. If the observed ΔD of 18.55 (direct) or 1.20 (generation) lies outside the central 95% of the bootstrap null distribution, the improvement is statistically credible; otherwise it is not. A complementary check is to report bootstrap confidence intervals for the D difference or perform k-fold cross-validation on the original data if access can be arranged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests entirely on Section 4's D values: 70.3002 vs 88.8502 for the direct fit and 13.99222 vs 15.19232 for the generation model. D is defined in Section 4 as a sum of squared differences in log relative frequency, but no standard error, confidence interval, or null-hypothesis test is reported for either D or the difference between the two models. The same binned observed data are used to estimate parameters by least squares and then to compute the winning D, so the comparison has no protection against sampling noise or overfitting. Moreover, pGL and pLNL are non-nested functional forms, so a raw difference in D cannot be interpreted without a null distribution. The abstract's word 'significantly' is therefore unsupported by the evidence presented. The gamma word-length mechanism is also unvalidated, but the quantitative claim of improved fit stands or falls on whether the D gap is statistically meaningful.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a gamma-like probability density pGL(s) for email size fluctuations, replacing the log-normal-like distribution pLNL(s) used in earlier work. The model is intended to improve fits in the small-size range by assuming that the length of a single word follows a gamma distribution rather than a normal distribution, thereby avoiding negative word lengths. The author fits both models to the size-frequency distribution of 'no attachment' emails and to a simulated email size generation model, reporting smaller least-squares discrepancy D for pGL in both cases. The central claim is that pGL provides a significantly improved fit over pLNL.","tokens_in":5680,"tokens_out":3015,"duration_ms":31107,"significance":"If the improved fit is real and statistically validated, the paper would provide a modestly better empirical description of email size fluctuations and a more plausible generative mechanism. However, the manuscript's significance is limited by the absence of any uncertainty quantification, model-selection test, or out-of-sample validation, and the generative mechanism is explicitly not tested against email content. The paper is honest in stating these limitations, which is a strength, but the central quantitative claim currently rests on raw point estimates of D with no error bars.","major_comments":[{"comment":"The claim that pGL fits better than pLNL is based solely on the reported D differences (70.3002 vs. 88.8502 for the direct fit and 13.99222 vs. 15.19232 for the generation model). No standard errors, confidence intervals, or significance tests are provided for D or for the difference between the two models. Because the same binned observed data are used both to estimate parameters by least squares and to compute the winning D, the comparison has no protection against sampling noise or overfitting. Moreover, pGL and pLNL are non-nested functional forms, so a raw D difference cannot be interpreted without a null distribution. The word 'significantly' in the abstract is therefore unsupported.","section":"Section 4, D values"},{"comment":"The simulation based on the generation model is not an independent test of the gamma-like distribution. The parameters used for pGL (k=24.1803, theta=0.0489) are exactly those obtained from the direct fit to the observed distribution, so the simulation reproduces the fitted distribution rather than validating it. Additionally, the pLNL generation parameters (mu=1.259, sigma=0.235) differ from the direct-fit pLNL parameters (mu=0.0489, sigma=0.2461) without any explanation; if the pLNL simulation is not optimized, the reported D gap of 15.19232 vs. 13.99222 may reflect suboptimal pLNL parameters rather than a genuine advantage of pGL.","section":"Section 4, generation model simulation"},{"comment":"The model's generative mechanism rests on the assumption that the length of a single word epsilon_t follows a gamma distribution, introduced to fix the negative-length problem of the normal distribution. No empirical evidence is provided that actual word lengths in the analyzed emails are gamma-distributed; the conclusion explicitly states that email contents were not analyzed. If the word-length distribution is not gamma, the improved fit of pGL is a curve-fitting artifact rather than support for the proposed mechanism. This should be acknowledged more explicitly and ideally tested on auxiliary data or at least framed as an untested assumption.","section":"Section 3, Eq. (3) and the gamma word-length assumption"}],"minor_comments":[{"comment":"The gamma density in Eq. (3) is written as (1/theta^k)(ln ln s)^(k-1) exp(-ln ln s/theta) without the factor 1/Gamma(k); presumably this normalization is absorbed into the constant a, but the expression is not a normalized density as written and this should be clarified.","section":"Section 3, Eq. (3)"},{"comment":"The text refers to 'the size–frequency distribution pLGL' which appears to be a typo for 'pGL’; correct this notation.","section":"Conclusion"},{"comment":"The definition of D is typeset awkwardly (\"D = ∑_s(ln yO(s) − ln y(s))^2\" with a superscript 2 after the sum). Please render the equation clearly and number it explicitly.","section":"Section 4"},{"comment":"Several typos interrupt the reading: 'wh ich' in the abstract, 's ending' in the introduction, and 'modeled using a gamma-like distribution' in the title is fine but the abstract contains 'sending mail' which is acceptable. A light proofreading pass is recommended.","section":"Abstract and Introduction"},{"comment":"The caption refers to 'sGL(s)' but should be 'pGL(s)'.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short empirical note building directly on the author's prior work. The central issue is not the mathematical derivation but the lack of statistical evidence for the headline claim. I believe the paper can be made publishable if the author adds uncertainty quantification (e.g., bootstrap or repeated simulation), performs an out-of-sample or cross-validated comparison, and either validates the gamma word-length assumption or clearly couches the model as a purely empirical curve. The unexplained discrepancy between direct-fit and generation-fit parameters for pLNL should also be resolved. I do not see this as a reject, because the underlying idea is coherent and the limitations are stated, but the current evidence is not sufficient for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, honest paper that replaces a normal distribution with a gamma distribution in the author's own earlier email-size model, and reports a better fit on the same dataset. The improvement is plausible, but the paper doesn't prove it is real: no error bars, no significance test, and the same data are used for fitting and evaluation.\n\nWhat's new: Eq. 3, the gamma-like density pGL(s) with x=ln ln s, is not in the prior papers. It's a standard composition, but applied to email sizes it's new. The motivation is sound: the previous generation model allowed negative word lengths; gamma fixes that. Both the direct fit and the Monte Carlo generation show lower D values (70.30 vs 88.85; 13.99 vs 15.19), so the author's claim is not baseless.\n\nSoft spots: The D statistic is a sum of squared log differences, and the gap between models is presented without any uncertainty estimate or significance test. The same binned data are used to fit parameters and then to compute the winning D, so overfitting is a real risk. The two models are non-nested, so a raw D gap can't be interpreted without a null distribution. The abstract says 'significantly improved fit' but the evidence doesn't support a statistical meaning of 'significantly'. The gamma word-length assumption is also unvalidated; the author admits email contents weren't analyzed. That's an honest admission, but it means the generative mechanism is speculative. Data and code aren't public, though privacy is a fair reason.\n\nCredit where due: the paper is transparent about its limitations and doesn't oversell the mechanism. The numerical comparison is clear and reproducible in principle from the reported parameter values.\n\nWho it's for: someone working in email-size statistics or power-law models of human communication. It's a very niche increment. I'd pass on citing it unless I needed the gamma-like form specifically. But it does deserve a serious referee: the empirical claim is falsifiable and a referee could ask for error bars or a hold-out check. I'd send it to review but expect heavy revision or a downgrade to a shorter note.","headline":"A small, honest incremental paper that swaps a normal for a gamma distribution in an email-size model and reports a better fit, but the improvement is not statistically substantiated.","tokens_in":6233,"tokens_out":2586,"would_cite":false,"duration_ms":25126,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["89.20.Hh","05.40.-a"],"model":"deepseek-v4-flash","headline":"The gamma-like distribution pGL(s) fits observed email size fluctuations better than the earlier log-normal-like model, with lower least-squares errors in both direct fitting and size-generation simulation.","keywords":["email size","gamma-like distribution","log-normal-like distribution","power-law fluctuations","email size generation model","frequency distribution","MIME content types"],"falsifier":"Count single-word lengths in bytes from a privacy-preserving corpus of actual email bodies and fit a gamma distribution to them; if the fitted shape and scale are far from k = 24.1803 and theta = 0.0489, or if a gamma fit is poor, then the generative mechanism behind the improved D values is unsupported.","tokens_in":5226,"feed_emoji":"📧","tokens_out":2897,"duration_ms":32043,"temperature":0.7,"pith_summary":"This paper is trying to establish that replacing the normal distribution for single-word length with a gamma distribution improves the fit of an email-size frequency model. The improved model, called gamma-like, is compared with the earlier log-normal-like model using the same observed email-size data from a university mail system. The gamma-like model yields smaller least-squares fitting errors in both the direct fit to observed frequencies and the fit of sizes generated by the corresponding generation model. If the improvement is real, it offers a modestly more accurate statistical description of email-size fluctuations, which matters for understanding the load that email services must handle.","feed_headline":"Gamma-like model better fits email size fluctuations","feed_subtitle":"Replacing a normal distribution for single-word length with a gamma distribution improves fits at small email sizes.","key_machinery":"The central object is the gamma-like density pGL(s) = 1/a * 1/(s ln s) * 1/$\\theta$^k * (ln ln s)^{k-1} * exp(-ln ln s / $\\theta$), defined for s > e, which is obtained by substituting a gamma distribution for the normal distribution inside the earlier log-normal-like construction. It is paired with the size generation model s_t = b s_{t-1}^c $e^{{epsilon_t}}$, where epsilon_t, the length of a single word, is now drawn from a gamma distribution (k, $\\theta$) instead of a normal distribution N(mu, $sigma^{2}$). The gamma distribution removes the possibility of negative word lengths while preserving the power-law tail of the generated sizes.","core_discovery":"The paper's central claim is that the gamma-like probability density pGL(s), formed by combining a gamma distribution with the transformation x = ln ln s, fits the observed frequency distribution of 'no attachment' email sizes better than the log-normal-like density pLNL(s). For the direct fit to observed data, the least-squares error is D = 70.3002 for pGL versus D = 88.8502 for pLNL. For the size generation model, which produces 191,993 simulated emails matching the observed count, the gamma-based model gives D = 13.99222 versus D = 15.19232 for the log-normal-based model. Both models preserve the same asymptotic power-law behavior for large email sizes, so the improvement is concentrated in the small-size range where the earlier model's normal assumption allowed negative word lengths.","pith_inferences":["If real single-word lengths are indeed gamma distributed, the nested exponentials in the generation model imply that sentence and compound-sentence lengths inherit heavier tails; this could be checked against sentence-length corpora for Japanese and English.","The improvement in D is modest and the paper reports no uncertainty estimates, so a reader should treat the advantage as suggestive until a model-selection criterion such as AIC or BIC, or a holdout-data comparison, confirms it.","The same substitution of a gamma for a normal distribution in the innermost exponential could be applied to other log-normal-like size distributions of text or file sizes, offering a direct way to test whether the improvement generalizes beyond email data.","Since pGL(s) has an extra shape parameter k compared with pLNL(s), part of the improved fit may simply reflect added flexibility rather than a deeper generative mechanism."],"forward_implications":["The gamma-like density pGL(s) is claimed to be a better empirical description of 'no attachment' email-size frequencies than the log-normal-like density, with D = 70.3002 versus 88.8502.","The gamma-based generation model is claimed to reproduce observed email-size frequencies more closely, with D = 13.99222 versus 15.19232 over 191,993 simulated emails.","Because both models have the same asymptotic power-law behavior, the improved fit does not change the operational implication that very large emails are rare but should still be considered when setting email size limits.","The gamma substitution fixes the formal flaw of negative single-word lengths in the earlier generation model, making the generative mechanism more plausible at the level of word structure."],"supporting_citations":[{"why":"Supplies the observed email-size data, the two subdistributions based on MIME content type, and the log-normal-like baseline pLNL(s) that this paper improves upon.","marker":"[13]"},{"why":"Provides the email size generation model s_t = b s_{t-1}^c e^{epsilon_t} with the normal distribution for epsilon_t that this paper replaces with a gamma distribution.","marker":"[15]"},{"why":"Supplies the logarithmic binning method used to produce the frequency distributions that are compared by the least-squares measure D.","marker":"[28]"},{"why":"Establishes the context of power-law correlations in email flow that motivates modeling the fluctuations in email size.","marker":"[12]"}],"fun_headline_variants":["Gamma beats log-normal for small email sizes","Email size fits improve with gamma curve","Gamma model nails small email size stats","Better email size model: gamma wins","Email size fluctuations: gamma fits better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes that the length of a single word in an email follows a gamma distribution, and this assumption is never checked against actual email content, which the paper says it did not analyze.","fun_headline_variants_meta":{"raw":{"variants":["Gamma beats log-normal for small email sizes","Email size fits improve with gamma curve","Gamma model nails small email size stats","Better email size model: gamma wins","Email size fluctuations: gamma fits better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000435,"raw_usage":{"total_tokens":2139,"prompt_tokens":796,"completion_tokens":1343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":1295}},"tokens_in":412,"tokens_out":1343,"duration_ms":12221,"temperature":1.0,"reasoning_tokens":1295,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:01:15.084960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count single-word lengths in bytes from a privacy-preserving corpus of actual email bodies and fit a gamma distribution to them; if the fitted shape and scale are far from k = 24.1803 and theta = 0.0489, or if a gamma fit is poor, then the generative mechanism behind the improved D values is unsupported.","supporting_citations":[{"cited_title":"Matsubara, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the observed email-size data, the two subdistributions based on MIME content type, and the log-normal-like baseline pLNL(s) that this paper improves upon."},{"cited_title":"Fluctuations in the email size modeled by a log-normal-like distribution","cited_arxiv_id":"2501.04042","evidence_quote":"Provides the email size generation model s_t = b s_{t-1}^c e^{epsilon_t} with the normal distribution for epsilon_t that this paper replaces with a gamma distribution."},{"cited_title":"Milojevi´ c, Power law distributions in information science: Mak ing the case for logarithmic binning","cited_arxiv_id":null,"evidence_quote":"Supplies the logarithmic binning method used to produce the frequency distributions that are compared by the least-squares measure D."},{"cited_title":"Matsubara, Y","cited_arxiv_id":null,"evidence_quote":"Establishes the context of power-law correlations in email flow that motivates modeling the fluctuations in email size."}],"review_version":1}