{"id":"bce66163-413a-4d0b-a20f-737722d7e542","arxiv_id":"2412.04102","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"With ideal statistics, the resolved-transition waiting-time estimator is best; the TUR wins near equilibrium, the blurred-transition estimator wins far from equilibrium, and finite statistics can cause overestimation.","lead":"The paper compares three ways of estimating entropy production from noisy, limited observations of Markov networks, using simulations of two model systems. It shows that coarse measurements can sometimes beat fine ones when data are scarce, and that waiting-time estimators can have a self-averaging effect that reduces their scatter.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed self-averaging variance reduction may be an artifact of discarding zero-count bins; the effective number of summands does not scale with 1/Δt under the paper's own selection rule.","rationale":"The reader's weakest assumption concerned the temporal coarse-graining model, which the authors explicitly scope and acknowledge. I identify a different, more central soft spot: the paper's headline self-averaging effect is asserted from simulation curves but is confounded by the documented practice of discarding bins with zero counts in either the forward or reverse waiting-time histogram. The concern is not that the authors are wrong—the effect may well be real—but that the current analysis does not distinguish a genuine variance reduction from a selection artifact. This is load-bearing because the strongest claim in the paper is precisely this self-averaging phenomenon, and the only evidence is the non-monotonic variance behavior in two examples. The proposed pseudocount test would settle the issue by removing the truncation and seeing whether the variance peak survives. I still regard the paper as a solid comparative simulation study; the analytical inequalities and the equality of the resolved-transition estimator for topologically trivial hidden subgraphs are convincing. The reader's CONDITIONAL verdict remains appropriate, now for a more specific reason.","tokens_in":15632,"tokens_out":13303,"duration_ms":153391,"concrete_test":"Re-run the finite-statistics simulations for Figs. 4(d) and 8(d) with a Laplace pseudocount α added to every bin of each empirical waiting-time distribution before computing the log ratios in Eq. (17), for α = 0.01, 0.1, and 1, with otherwise identical parameters. If the variance maximum at intermediate Δt disappears or becomes monotone under α > 0, the reported self-averaging is caused by the zero-bin discarding rule rather than by an intrinsic bin-count averaging. If the intermediate maximum persists for all α, the self-averaging claim is supported. Also record the number of non-discarded bins in each setting to check whether the variance tracks the effective number of summands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most novel claim—that a large number of bins in empirical waiting-time distributions produces a self-averaging reduction of estimator variance (Sec. VI)—is not supported by the evidence as presented. In Sec. IV.A the empirical estimator σ_WTD^{Δt,T} in Eq. (17) is computed only over bins for which both ψ_{ij→kl}^{Δt,T}(t_i) and ψ_{kl→ij}^{Δt,T}(t_i) are nonzero; all other bins are discarded. Consequently, the number of effective summands is not 1/Δt as the heuristic 'number of bins' suggests, but is bounded by the realized number of observed events and, at sufficiently fine resolution, decreases because the probability that both forward and reverse counts fall in the same tiny bin shrinks. The variance reduction seen in Figs. 4(d) and 8(d) at small Δt may therefore reflect selection: at fine resolution the estimator degenerates toward a small set of bins or zero, rather than averaging over many independent contributions. No analytical derivation, and no control experiment that isolates the binning from the zero-count truncation, is provided; the text only asserts the effect. Because this self-averaging effect is the central claimed new phenomenon, the numerical evidence as analyzed is not yet conclusive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how finite temporal resolution, finite spatial resolution, and finite observation statistics affect three entropy estimators for partially accessible Markov networks: the thermodynamic uncertainty relation (TUR), a waiting-time-based estimator for resolved transitions, and a waiting-time-based estimator for blurred transitions. The authors use two paradigmatic systems, a four-state Markov network and an augmented Michaelis-Menten reaction scheme, and they perform Gillespie simulations. The main analytical ingredients are the log-sum inequality for coarse-grained waiting-time distributions (Eq. (10)) and a binomial fluctuation model for histogram counts (Eq. (16)). The numerical results show that the resolved-transition waiting-time estimator performs best under ideal statistics, that the TUR performs better at low affinities while the blurred-transition estimator performs better at high affinities, that lower resolution can be beneficial at finite statistics, and that finite statistics can lead to empirical estimates exceeding the true entropy production. The paper also claims that a large number of bins in empirical waiting-time distributions introduces a self-averaging effect that reduces estimator variance (Sec. VI).","tokens_in":15858,"tokens_out":12195,"duration_ms":124148,"significance":"If the results hold, the paper offers useful practical guidance for experimental entropy estimation: it identifies a trade-off between resolution and statistical convergence and provides a quantitative warning that empirical estimates from finite data can violate the strict lower-bound property. The analytical bounds are correctly derived, and the simulations are extensive and reproducible in structure. The observation about overestimation at finite statistics, documented in Fig. 5, is practically important. However, the most novel qualitative claim, the self-averaging variance reduction, is not supported by the evidence as presented and requires additional analysis or a control experiment before the central message can be accepted.","major_comments":[{"comment":"The central claim of a self-averaging variance reduction for waiting-time estimators is not supported by the presented evidence. The empirical estimator in Eq. (17) is evaluated only over bins for which both forward and reverse empirical waiting-time counts are nonzero; the text in Sec. IV.A states that contributions with a zero count are discarded. Under this selection rule, the effective number of summands at small Δt is not 1/Δt as the heuristic \"number of bins\" suggests, but is controlled by the realized event counts and the probability that forward and reverse events fall in the same bin, which decreases as Δt→0. The variance reduction visible in Figs. 4(d) and 8(d) at fine resolution may therefore be a selection artifact, caused by the estimator degenerating toward a few surviving bins or toward zero, rather than a self-averaging over many independent contributions. To establish the claimed effect, the authors should provide an analytic calculation of the variance of the truncated estimator, or a control experiment that isolates the binning effect from the zero-count truncation (for example, by using a fixed number of bins with a pseudocount or smoothing regularizer that avoids discarding bins). Without such support, the conclusion in Sec. VI that \"a large number of bins ... introduces a self-averaging effect\" is not justified.","section":"Sec. IV.A and Sec. VI"}],"minor_comments":[{"comment":"The definitions of the empirical average ⟨·⟩_N and variance Var_N appear to be missing the factor 1/N; as written, they are the sum and sum of squared deviations, not the mean and variance. Please include the normalization explicitly or clarify the intended definition.","section":"Eq. (18) and Eq. (19)"},{"comment":"The concluding statement that \"the empirical waiting-time distributions have Gaussian fluctuations in general\" is stronger than the evidence presented: Eq. (16) treats the waiting-time events as independent and identically distributed, which ignores correlations in a Markov network, and only a single illustrative comparison is shown in Fig. 3(b). Please either restrict the claim to the cases considered or provide additional justification for the independence approximation.","section":"Sec. III and Sec. VI"},{"comment":"The text says the coarse-graining is illustrated for ψ_{32→32}(t), while the figure caption and the plot refer to ψ_{43→43}(t); one of these is a typo and should be corrected.","section":"Sec. II.C and Fig. 3(a)"},{"comment":"In the sentence \"The network consists of two fundamental cycles C1 and C1 with three edges,\" the second cycle should be C2, not C1.","section":"Sec. IV.A"},{"comment":"The symbol N is used both for the normal distribution in Eq. (16) and for the number of realizations in Eqs. (18) and (19); please disambiguate, for example by using a calligraphic or script letter for one of them.","section":"Notation"},{"comment":"The caption states \"In the limit T → ∞ (black line), the entropy estimator depends on the time resolution and decreases for large Δt\" in a way that appears to apply to both panels (c) and (d), but the text says the resolved-transition estimator in panel (d) is independent of Δt in that limit; please clarify which panel each statement refers to.","section":"Fig. 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the authors are well-known in the field. The numerical study is carefully done and the analytical elements are standard. The main risk is the unsupported self-averaging claim, which is also the paper's most distinctive contribution. I would encourage the editor to request a control experiment or analytic variance calculation for the truncated estimator before accepting. There is no concern about citation practice or novelty disclosure in the manuscript itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a careful comparative simulation study of three entropy estimators under finite resolution and finite statistics. The practical trade-offs it maps out are useful. But the paper's most novel claim—a self-averaging effect that reduces variance when waiting-time histograms have many bins—is not established, because the authors discard every bin with a zero count in either direction. The stress-test note is right: under that selection rule the effective number of summands need not grow with 1/Δt, and at fine resolution the estimator may be dominated by a few bins or collapse toward zero. The variance reduction in Figs. 4(d) and 8(d) could be a selection artifact rather than averaging over many independent contributions.\n\nWhat is good: The analytical part is solid. The log-sum inequality in Eq. (10) is correctly applied to show the coarse-grained waiting-time estimator remains a lower bound at finite temporal resolution. Eq. (16) gives the binomial fluctuation statistics for histogram counts, and the comparison with simulations in Fig. 3(b) is convincing. The two model systems—four-state network and augmented Michaelis-Menten—are well chosen, and the qualitative conclusions about estimator rankings are consistent across both. The warning that finite statistics break the strict lower-bound property is important and clearly demonstrated. The paper also honestly acknowledges the limitation of its temporal coarse-graining model (it assumes all transitions are still registered even when waiting times are shorter than Δt).\n\nSoft spots: besides the self-averaging issue, the paper does not ship code or data, which would help others check the variance claim. The selection bias from discarding empty bins is acknowledged, but its quantitative effect on the variance is not analyzed. There is no analytical derivation for the self-averaging effect, only a heuristic argument that more bins means more averaging. Given that the claimed effect is the main new phenomenon, this is a real gap.\n\nWho is this for: anyone estimating entropy production from single-molecule or enzyme data with limited temporal resolution and short trajectories. The comparison of TUR vs waiting-time estimators at different affinities is genuinely useful. The paper deserves peer review; a referee should ask for a control experiment that separates the binning effect from the zero-count truncation, or an analytical treatment of the estimator's variance under the selection rule.","headline":"Solid comparative simulation study with useful practical trade-offs, but the headline self-averaging claim is undercut by the zero-count bin discarding and needs a control or derivation.","tokens_in":16408,"tokens_out":2450,"would_cite":true,"duration_ms":25154,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["05.40.-a","05.70.Ln"],"model":"deepseek-v4-flash","headline":"This paper claims that finite observation statistics change entropy-production estimators: waiting-time estimates can exceed the true entropy, and many histogram bins lower variance via self-averaging.","keywords":["entropy production estimation","waiting-time distributions","thermodynamic uncertainty relation","partial observation","finite statistics","Markov networks","Michaelis-Menten kinetics","stochastic thermodynamics"],"falsifier":"Repeat the estimator comparison with an alternative coarse-graining in which at most one transition can be registered per time interval $\\Delta t$; if the TUR estimate then exceeds the true entropy production or the ranking of $\\langle\\sigma_{\\mathrm{WTD}}\\rangle$, $\\Sigma$, and $\\langle\\sigma_{\\mathrm{TUR}}\\rangle$ changes, the paper's qualitative conclusions are limited to its specific observation model. For the self-averaging claim, hold the observation time $T$ fixed, vary the bin width $\\Delta t$, and plot the variance of $\\sigma^{\\Delta t,T}_{\\mathrm{WTD}}$ against the number of occupied bins: the claim predicts a monotonic variance decrease as bins multiply, with a minimum at the finest resolution.","tokens_in":15388,"feed_emoji":"⏱️","tokens_out":7767,"duration_ms":67204,"temperature":0.7,"pith_summary":"Estimating entropy production from real trajectories of partially observed Markov networks usually assumes perfect data. This paper asks what happens when the data actually have finite temporal and spatial resolution and finite statistics, and it compares three estimators—the thermodynamic uncertainty relation (TUR), a waiting-time estimator for resolved transitions, and a waiting-time estimator for blurred transitions—on a four-state network and an augmented Michaelis-Menten reaction scheme. Its central message is that the best estimator depends on the regime: resolved waiting-times win with ideal statistics, the TUR is tighter near equilibrium, and blurred-transition waiting-times improve far from equilibrium, while coarser resolution can be beneficial when observation time is short. The paper's key new finding is that finite statistics can push empirical waiting-time estimates above the true entropy production, so they no longer provide a strict lower bound, even though the large number of histogram bins produces a self-averaging effect that reduces estimator variance.","feed_headline":"Finite data breaks entropy lower bounds—and coarse resolution helps","feed_subtitle":"Estimator rankings shift with resolution, and finite trajectories can overshoot the true entropy.","key_machinery":"The central objects are waiting-time distributions $\\psi_{(ij)\\to(kl)}(t)$, the probability density that a registered transition $(kl)$ follows a registered transition $(ij)$ after a lag $t$, and their blurred counterparts $\\Psi_{I\\to J}(t)$ for lumped transition classes $I,J$. From these, Eq. (8) defines the resolved-transition estimator $\\langle\\sigma_{\\mathrm{WTD}}\\rangle$ and Eqs. (12)-(14) define the blurred-transition estimator $\\Sigma$ as the average of a waiting-time term and a transition-count term, both of which are needed because blurred waiting-time distributions alone do not satisfy a lower-bound inequality. At finite temporal resolution the continuous densities are replaced by histograms of bin width $\\Delta t$, and the log-sum inequality guarantees the binned estimator remains a lower bound only in the ideal-statistics limit. At finite statistics, every histogram bin is treated as an independent, approximately Gaussian fluctuating count, and the fluctuation analysis of these counts is what produces the overestimation and self-averaging conclusions.","core_discovery":"At finite observation time $T$ the empirical waiting-time counts $n^{\\Delta t,T}_{(ij)\\to(kl)}(t_i)$ fluctuate between trajectories, and the paper shows these fluctuations are approximately Gaussian with mean and variance given by the binomial form in Eq. (16). Because the estimators $\\sigma^{\\Delta t,T}_{\\mathrm{WTD}}$ and $\\Sigma^{\\Delta t,T}$ are sums over these fluctuating bin counts, their ensemble averages over $N$ trajectories need not respect the ideal-statistics bound $\\langle\\hat\\sigma\\rangle \\le \\langle\\sigma\\rangle$: finite-statistics estimates can overestimate the true entropy production, so they fail to provide a strict lower bound. At the same time, the sum over many independent bins gives a self-averaging effect: the variance of the waiting-time estimators decreases as the number of bins grows, even though the mean estimate at high temporal resolution is worse because most bins are empty and must be discarded. The paper also shows that lower temporal resolution makes histograms converge faster, creating a trade-off between statistical convergence and the quality of the bound that a fully converged estimator would provide.","pith_inferences":["The Gaussian bin-count formula in Eq. (16) could be turned into a one-trajectory error bar: propagate the binomial variances through the estimator to report confidence intervals without repeating simulations, a step the paper motivates but does not implement.","Because estimator rankings depend on affinity, a hybrid rule that estimates the affinity from the data and then chooses between $\\langle\\sigma_{\\mathrm{TUR}}\\rangle$ and $\\Sigma$ might outperform either estimator alone; the paper does not propose such a rule.","The self-averaging effect suggests that over-resolving waiting-time histograms could be used deliberately as a variance-reduction tool in experiments, at the cost of bias, possibly with a debiasing correction.","The overestimation phenomenon implies that published entropy-production lower bounds based on single finite trajectories should be interpreted with caution; the paper notes the need to quantify this overestimation."],"forward_implications":["For perfect measurement statistics, the resolved-transition waiting-time estimator $\\langle\\sigma_{\\mathrm{WTD}}\\rangle$ outperforms both the TUR and the blurred-transition estimator in every scenario studied, and in the four-state network it recovers the full entropy production independent of temporal resolution whenever one edge of each fundamental cycle is visible.","At low driving affinity the thermodynamic uncertainty relation gives a tighter lower bound than the blurred-transition estimator, while at high affinity the blurred estimator is tighter, so estimator choice should be guided by the expected driving regime.","Higher temporal or spatial resolution slows the convergence of measurement statistics, so for short observation times a deliberately coarser resolution can yield a better estimate than a fine one.","Finite statistics break the strict lower-bound property: a single empirical run can overestimate entropy production, and the probability of overestimation grows as the mean estimate approaches the true value while the maximum size of the overestimation shrinks.","Waiting-time estimators show a self-averaging effect: more histogram bins reduce estimator variance at high temporal resolution, and variance is largest at intermediate resolution where neither self-averaging nor statistical convergence dominates."],"supporting_citations":[{"why":"Barato and Seifert's thermodynamic uncertainty relation supplies the $\\langle\\sigma_{\\mathrm{TUR}}\\rangle \\le \\langle\\sigma\\rangle$ bound used as one of the three estimators.","marker":"[37]"},{"why":"van der Meer, Ertel, and Seifert derive the resolved-transition waiting-time estimator and the condition under which it recovers the full entropy production.","marker":"[69]"},{"why":"Harunari, Dutta, Polettini, and Roldán independently derive the waiting-time based lower bound that supports Eq. (8).","marker":"[70]"},{"why":"Ertel and Seifert derive the blurred-transition bound $\\Sigma$ used for spatially coarse observations.","marker":"[85]"},{"why":"Dechant and Sasa show that the TUR can lose its lower-bound property when only one event per time interval is registered, which motivates the paper's assumption about temporal coarse-graining.","marker":"[94]"},{"why":"Cover and Thomas's log-sum inequality provides the step that keeps the binned waiting-time estimator a lower bound in the ideal-statistics limit.","marker":"[95]"},{"why":"Gillespie's algorithm generates the stochastic trajectories used for all finite-statistics simulations.","marker":"[96]"}],"fun_headline_variants":["Finite data breaks entropy bound; coarse resolution helps","Short stats shatter entropy lower bound, low resolution rescues","Coarse time wins when entropy data is finite","Entropy bound fails at finite stats; coarse beats fine","Finite observations flip entropy bound; lower resolution better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that finite temporal resolution never causes transitions to be missed or merged: the observer always records the true sequence of events and only the measured waiting times are imprecise.","fun_headline_variants_meta":{"raw":{"variants":["Finite data breaks entropy bound; coarse resolution helps","Short stats shatter entropy lower bound, low resolution rescues","Coarse time wins when entropy data is finite","Entropy bound fails at finite stats; coarse beats fine","Finite observations flip entropy bound; lower resolution better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1405,"prompt_tokens":955,"completion_tokens":450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":571,"tokens_out":450,"duration_ms":5163,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:46:05.423113+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the estimator comparison with an alternative coarse-graining in which at most one transition can be registered per time interval $\\Delta t$; if the TUR estimate then exceeds the true entropy production or the ranking of $\\langle\\sigma_{\\mathrm{WTD}}\\rangle$, $\\Sigma$, and $\\langle\\sigma_{\\mathrm{TUR}}\\rangle$ changes, the paper's qualitative conclusions are limited to its specific observation model. For the self-averaging claim, hold the observation time $T$ fixed, vary the bin width $\\Delta t$, and plot the variance of $\\sigma^{\\Delta t,T}_{\\mathrm{WTD}}$ against the number of occupied bins: the claim predicts a monotonic variance decrease as bins multiply, with a minimum at the finest resolution.","supporting_citations":[{"cited_title":"van der Meer, B","cited_arxiv_id":null,"evidence_quote":"van der Meer, Ertel, and Seifert derive the resolved-transition waiting-time estimator and the condition under which it recovers the full entropy production."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Harunari, Dutta, Polettini, and Roldán independently derive the waiting-time based lower bound that supports Eq. (8)."},{"cited_title":"Ertel and U","cited_arxiv_id":null,"evidence_quote":"Ertel and Seifert derive the blurred-transition bound $\\Sigma$ used for spatially coarse observations."},{"cited_title":"Dechant and S.-i","cited_arxiv_id":null,"evidence_quote":"Dechant and Sasa show that the TUR can lose its lower-bound property when only one event per time interval is registered, which motivates the paper's assumption about temporal coarse-graining."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cover and Thomas's log-sum inequality provides the step that keeps the binned waiting-time estimator a lower bound in the ideal-statistics limit."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gillespie's algorithm generates the stochastic trajectories used for all finite-statistics simulations."}],"review_version":1}