{"id":"b7a5d9b7-12ae-4f48-9066-f638c13877a4","arxiv_id":"2411.11271","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"A truncation-based estimator achieves the optimal heavy-tailed mean estimation rate in smooth Banach spaces under martingale dependence with time-uniform guarantees.","lead":"This paper proves time-uniform concentration bounds for a simple truncation-based mean estimator in smooth Banach spaces, valid for heavy-tailed data with between one and two moments and under martingale dependence. The guarantee matches the best known heavy-tailed rates, comes with explicit small constants, and holds at all sample sizes simultaneously.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1's proof contains an algebraic error: the variance term is 2^(j/p)/(n-k), not (2^j/(n-k))^(1/p); the printed factorization is not an identity, so the LIL bound requires correction.","rationale":"After checking the proof of Theorem 2.1, the martingale argument via Proposition 3.3 and Lemma 3.4 is coherent for the value rho = beta^(-1) used in the theorem; the 2-smoothness assumption is explicit and standard, so the reader's weakest assumption is a scope limitation rather than an internal inconsistency. The one place where the argument as written breaks down is the stitching proof of the LIL inequality: the displayed factorization of W(n,j) is algebraically invalid. Since the corrected term still yields the claimed rate, the appropriate response is to require the authors to fix this step, along with the mechanical typos already noted by the reader, rather than to reject the paper. The conditional verdict therefore stands unchanged.","tokens_in":920,"tokens_out":1201,"duration_ms":291362,"concrete_test":"Recompute W(n,j) = lambda_j^(p-1) C_p(j) + beta log(4/delta_j) / (lambda_j (n-k(j))) using the lambda_j defined in the proof of Theorem 4.1, for p=1.5, j=10, n-k=2^j, C_p(j)=1, and beta log(4/delta_j)=1. Verify that the manuscript's displayed expression (1/2^j)^((p-1)/p) + (2^j/(n-k))^(1/p) evaluates to about 1.01, whereas the correct value of W(n,j) is about 0.109. Then redo the factorization with the corrected term 2^(j/p)/(n-k) and confirm that the final bound W(n,j) <= C_p(j)^(1/p) ( beta log(4/delta_j)/(n-k(j)) )^((p-1)/p) [ 2^((p-1)/p) + 2^(1/p) ] follows.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the proof of Theorem 4.1, with lambda_j = ( beta log(4/delta_j) / (2^j C_p(j)) )^(1/p), the second term of W(n,j) is beta log(4/delta_j) / (lambda_j (n-k(j))) = C_p(j)^(1/p) ( beta log(4/delta_j) )^((p-1)/p) * 2^(j/p) / (n-k(j)). The manuscript instead writes (2^j/(n-k(j)))^(1/p) and then factors ( beta log / (n-k) )^((p-1)/p) out of both terms, producing the claimed O((log(4h)/n)^((p-1)/p)) bound. That factorization is not an identity; for example, with p=1.5, j=10, n-k=2^j, and unit constants, the printed first-line value is about 1.01 while the correct W(n,j) is about 0.109. The final LIL rate is nevertheless valid if the corrected term 2^(j/p)/(n-k) is used: after multiplying the target inequality by (n-k)^((p-1)/p), the two bracketed factors become ((n-k)/2^j)^((p-1)/p) <= 2^((p-1)/p) and (2^j/(n-k))^(1/p) <= 2^(1/p). Thus Theorem 4.1 is likely true but the proof as printed has a genuine gap that must be fixed before the LIL claim can be accepted as written. The main fixed-sample-size result, Theorem 2.1, is not affected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies estimation of the shared conditional mean of Banach-space-valued random variables under a conditional p-th moment bound (1 < p ≤ 2) and martingale dependence. It analyzes a truncation estimator centered at a preliminary naive mean estimate, proving a time-uniform line-crossing inequality (Theorem 2.1), a fixed-sample optimized bound with rate O(β v^{1/p}(log(1/δ)/n)^{(p-1)/p}) (Corollary 2.2), and an iterated-logarithm-rate bound obtained by stitching (Theorem 4.1). The proof strategy is to compare the truncated process to a martingale, bound the truncation bias via Lemma 3.1, and use a Pinelis-type supermartingale combined with Ville's inequality. Numerical comparisons against geometric median-of-means are reported.","tokens_in":23256,"tokens_out":12229,"duration_ms":105910,"significance":"If the stated results hold, this is a useful contribution: it extends truncation-based mean estimation to 2-smooth Banach spaces, removes finite-variance and i.i.d. assumptions, and provides dimension-free time-uniform bounds with explicit constants. I verified the main technical ingredients—Lemma 3.1, Lemma 3.2, and the variance bound (3.9)—and the overall Pinelis/Ville architecture is sound. The rate matches Minsker's geometric median-of-means rate while permitting martingale dependence. However, two algebraic issues—a constant mismatch in Theorem 2.1 and an incorrect factorization in the proof of Theorem 4.1—need correction before the displayed results can be taken as proven.","major_comments":[{"comment":"The proof of Theorem 2.1 invokes Proposition 3.5 with ρ = β^{-1}. In Eq. (3.11) this gives the coefficient β C_p(B) + β^{-1} K_p 2^{p-1} in front of (v + r(δ_2, k)^p), not β C_p(B) + K_p 2^{p-1} as displayed in Theorem 2.1. The same issue appears in Theorem C.1. Corollary 2.3 already uses the corrected form C_p = C_p(B) + β^{-1} K_p 2^{p-1}, so the main theorem and its corollaries are internally inconsistent. Please revise the theorem constants (and the corollary statements that depend on them) or supply an additional argument that restores the printed constant.","section":"Section 3, Proposition 3.5 and proof of Theorem 2.1"},{"comment":"The displayed computation of W(n, j) contains a false equality: substituting λ_j = (β log(4/δ_j) / (2^j C_p(j)))^{1/p} yields the second summand C_p(j)^{1/p} (β log(4/δ_j))^{(p-1)/p} 2^{j/p} / (n - k(j)), whereas the proof writes C_p(j)^{1/p} (β log(4/δ_j))^{(p-1)/p} (2^j / (n - k(j)))^{1/p}; the two differ by the factor (n - k(j))^{1 - 1/p}. With the corrected term, the next line factors correctly as ((n - k(j))/2^j)^{(p-1)/p} + (2^j/(n - k(j)))^{1/p}, and the final bound follows, so the theorem is salvageable, but the printed proof must be fixed.","section":"Section 4, proof of Theorem 4.1"}],"minor_comments":[{"comment":"The numerator in the integral term is printed as e^{2ρ} - ρ - 1; it should be e^{2ρ} - 2ρ - 1 to match ∫_0^1 (1 - θ) e^{2ρθ} dθ. Since the printed value is larger, this does not invalidate the inequality, but it is a typo.","section":"Proposition 3.3, Eq. (i)"},{"comment":"The statement says 'simultaneously for all n ≥ k', but the right-hand side has n - k in the denominator; it should say n > k (or n ≥ k + 1).","section":"Theorem 2.1"},{"comment":"The definition of r_n uses h(⌊log_2 n⌋)^{-1} δ, while Theorem 4.1 uses h(...)^{-1} δ/2; align the two definitions.","section":"Corollary 4.2"},{"comment":"The text says n = 100,000 samples are used, while the Figure 2 caption says n = 10^6; please reconcile.","section":"Section 5 / Figure 2 caption"},{"comment":"The proof says 'Let bZ_k be the empirical mean on the first k − 1 observations', but the statement uses the first k observations; this indexing should be unified.","section":"Proof of Corollary 2.3"}],"recommendation":"major_revision","confidential_remarks":"The issues are concentrated in one constant inconsistency and one displayed algebra step; I believe a focused revision can address them. I would not require new statistical ideas. The rate claims and the estimator's contributions are credible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Strong paper, with a genuine new result and a mostly clean proof. The main contribution is real: the truncation estimator from Catoni-Giulini is extended to central p-th moments with p<2, martingale dependence, and 2-smooth Banach spaces, with time-uniform bounds and explicit constants. Replacing the PAC-Bayes machinery with a Pinelis-style supermartingale is an improvement; it is what lets them handle martingales and heavy tails at once. The rate matches Minsker's geometric median-of-means but with a computationally trivial estimator that is valid at stopping times. I checked the key steps in Section 3: Lemma 3.1 and the variance bound (3.9) are correct, and Proposition 3.3 follows the Pinelis template. I do not see circularity; the bounds stand on their own.\n\nThe soft spots are real but not load-bearing. The proof of Theorem 4.1 contains an algebraic error in the stitching calculation. With lambda_j = (beta log(4/delta_j) / (2^j C_p(j)))^(1/p), the second term in W(n,j) is C_p(j)^(1/p) (beta log)^((p-1)/p) * 2^(j/p) / (n-k), not (2^j/(n-k))^(1/p) as printed. The claimed factorization is not an identity; with p=1.5 and j=10, the printed expression and the correct one differ by roughly an order of magnitude. The final LIL rate survives: if you multiply the target inequality by (n-k)^((p-1)/p), the two bracketed terms are bounded by 2^((p-1)/p) and 2^(1/p). So the theorem is likely true, but the proof as written has a gap that a referee should require be fixed. There are also minor mechanics: Theorem 2.1 says \"n >= k\" but the estimator is undefined at n=k; Corollary 2.3's statement uses k while the proof uses k-1; and the simulation sample sizes are stated as both 100,000 and 10^6. All three are cosmetic.\n\nThe 2-smoothness assumption is genuinely limiting—it excludes L^alpha for alpha<2—but the paper states this clearly, and it is standard for Pinelis-type tools. Not a flaw, just scope.\n\nBottom line: this paper deserves a serious referee. I would accept it conditionally: fix the LIL proof, harmonize the indexing, and release the simulation code. The main theorem is solid and the contribution is useful for anyone doing sequential heavy-tailed estimation.","headline":"Solid main result; the LIL proof has a fixable algebraic gap that should be corrected before acceptance.","tokens_in":732,"tokens_out":2642,"would_cite":true,"duration_ms":48778,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F10","60G42","60B11"],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage truncated mean estimator attains the optimal heavy-tailed rate in 2-smooth Banach spaces, with time-uniform guarantees that hold under martingale dependence and infinite variance.","keywords":["heavy-tailed mean estimation","truncation estimator","Banach space","martingale dependence","time-uniform concentration","infinite variance","smooth Banach spaces","dimension-free bounds"],"falsifier":"Look for a concrete martingale difference sequence in a 2-smooth Banach space satisfying the moment assumption with p in (1,2) for which the claimed bound (2.6) fails at some time n with probability exceeding delta; the natural place to search is a sequence whose conditional p-th moment saturates the bound v and whose increments are highly non-symmetric, such as heavy-tailed variables with Pareto tail index p, since the proof's margin shrinks there.","tokens_in":22693,"feed_emoji":"📊","tokens_out":11245,"duration_ms":102054,"temperature":0.7,"pith_summary":"This paper claims that a simple truncation-based mean estimator, previously thought to need finite variance and Euclidean space, works in any separable 2-smooth Banach space under a bound on the p-th central moment for p in (1,2] and under martingale dependence. The estimator takes a small pilot sample to form a naive mean, then truncates subsequent observations to a ball centered at that pilot estimate and averages them. If the claims are right, this gives dimension-free, time-uniform high-probability bounds that match the geometric median-of-means rate, with a computationally cheap and online-updatable estimator. The proof works by building a nonnegative supermartingale from the truncated increments and applying the maximal inequality for nonnegative supermartingales, so the bounds hold at every stopping time.","feed_headline":"Simple estimator reaches optimal heavy-tailed mean rate","feed_subtitle":"A single truncation step gives dimension-free, time-uniform bounds under pth-moment assumptions.","key_machinery":"The carrying mechanism is a nonnegative supermartingale built from the truncated, centered increments: M_n = (1/2) exp(rho ||xi_n - mu sum_{m<=n} lambda_m|| - V_n), where V_n = ($rho^{2}$ $beta^{2}$ C_p(B) + rho K_p $2^{{p-1}}$) sum_{m<=n} lambda_m^p (v + ||mu - hatZ_m||^p). The maximal inequality for nonnegative supermartingales turns the fact that M starts at 1 into a time-uniform probability bound. Two ingredients feed the construction: a lemma bounding the bias of truncation using a moment inequality with constant K_p, and a concentration inequality for bounded martingales in 2-smooth spaces, following the approach of [33, 34], which controls the fluctuation term with the smoothness constant $\\beta$. The pilot estimate hatZ_k enters only through an additive error r(delta_2,k)^p, which fades as n grows.","core_discovery":"The central discovery is that the deficiencies of truncation-based estimators are not fundamental. Theorem 2.1 states that if observations are conditionally centered at the unknown mean, have p-th central moments bounded by v for p in (1,2], and live in a separable $\\beta$-smooth Banach space, then the estimator formed by centering on a pilot estimate and truncating to a ball of radius 1/$\\lambda$ satisfies, with probability at least 1-delta, simultaneously for all n at least k: ||bmu_n(k) - mu|| <= $lambda^{{p-1}}$($\\beta$ C_p(B) + K_p $2^{{p-1}}$)(v + r(delta_2,k)^p) + $\\beta$ log(2/delta_1)/($\\lambda$(n-k)). Optimizing $\\lambda$ yields the rate O($\\beta$ $v^{{1/p}}$ (log(1/delta)/n)^{(p-1)/p}), matching the geometric median-of-means rate while allowing infinite variance, martingale dependence, and any dimension. The same machinery yields iterated-logarithm rates that are tight at all times up to a doubly logarithmic factor in n.","pith_inferences":["Because the proof allows predictable sequences of truncation levels and pilot estimates, lambda could in principle be tuned online using earlier data without breaking the supermartingale; the paper does not advertise this generality.","A natural next target is a fully adaptive estimator that replaces the known moment bound v with an online estimate, since the optimal lambda is chosen using v.","The same supermartingale technique may produce Bernstein-style 'separation of rates' in Banach spaces, analogous to decomposing a covariance into trace and operator-norm terms, if a weak-moment quantity can be tracked instead of the full norm.","The pilot sample size k can be as small as log n and still give the optimal rate, so in practice the estimator's overhead is negligible even in infinite-dimensional settings."],"forward_implications":["The estimator is computationally simple and can be updated online, so the time-uniform guarantee makes it usable in sequential decision problems such as bandits with heavy tails.","Because the bounds hold at all stopping times, they can be used to construct anytime-valid confidence sequences for the mean of a heavy-tailed Banach-space-valued process.","In Hilbert spaces and in L^alpha or ell^alpha spaces with alpha at least 2, the smoothness constant beta is finite and explicit, so the bounds are dimension-free and implementable.","The iterated-logarithm version gives a simultaneous-in-time bound of order O(beta v^{1/p} (log log n / n)^{(p-1)/p}), so the estimator adapts to unknown stopping horizons.","If the p-th moment bound holds for p > 2, Jensen's inequality reduces the problem to p = 2 and yields a sqrt(log(1/delta)/n) rate, extending the result beyond its main focus."],"supporting_citations":[{"why":"Provides the original truncation estimator that this paper centers and analyzes.","marker":"[6]"},{"why":"Gives the geometric median-of-means estimator and its rate, which the paper matches and uses as a pilot estimator.","marker":"[30]"},{"why":"Supplies the martingale inequalities for smooth Banach spaces that the paper extends.","marker":"[33]"},{"why":"Provides the Bennett-type concentration bound and Taylor-expansion argument at the core of the key concentration proposition.","marker":"[34]"},{"why":"The maximal inequality for nonnegative supermartingales that turns the e-process into a time-uniform probability bound.","marker":"[37]"},{"why":"Supplies the stitching technique used to derive iterated-logarithm rates from line-crossing inequalities.","marker":"[18]"}],"fun_headline_variants":["Truncation estimator masters heavy-tailed martingales","Optimal mean rates despite infinite variance and dependence","Simple truncation achieves tight time-uniform heavy-tail bounds","Heavy-tailed mean? One truncation step suffices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument assumes the Banach space is separable and 2-smooth, meaning its norm satisfies the squared-smoothness inequality in Equation (1.1), so the proof's martingale concentration bound has no basis in spaces like L^$\\alpha$ or ell^$\\alpha$ with $\\alpha$ < 2.","fun_headline_variants_meta":{"raw":{"variants":["Truncation estimator masters heavy-tailed martingales","Optimal mean rates despite infinite variance and dependence","Simple truncation achieves tight time-uniform heavy-tail bounds","Heavy-tailed mean? One truncation step suffices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1400,"prompt_tokens":971,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":366}},"tokens_in":587,"tokens_out":429,"duration_ms":4525,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:43:55.481632+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look for a concrete martingale difference sequence in a 2-smooth Banach space satisfying the moment assumption with p in (1,2) for which the claimed bound (2.6) fails at some time n with probability exceeding delta; the natural place to search is a sequence whose conditional p-th moment saturates the bound v and whose increments are highly non-symmetric, such as heavy-tailed variables with Pareto tail index p, since the proof's margin shrinks there.","supporting_citations":[{"cited_title":"Geometric median and robust estimation in Banach spaces","cited_arxiv_id":null,"evidence_quote":"Gives the geometric median-of-means estimator and its rate, which the paper matches and uses as a pilot estimator."},{"cited_title":"An approach to inequalities for the distributions of infinite-dimensional martingales","cited_arxiv_id":null,"evidence_quote":"Supplies the martingale inequalities for smooth Banach spaces that the paper extends."},{"cited_title":"Optimum bounds for the distributions of martingales in Banach spaces","cited_arxiv_id":null,"evidence_quote":"Provides the Bennett-type concentration bound and Taylor-expansion argument at the core of the key concentration proposition."},{"cited_title":"Time-uniform, nonparametric, nonasymptotic confidence sequences","cited_arxiv_id":null,"evidence_quote":"Supplies the stitching technique used to derive iterated-logarithm rates from line-crossing inequalities."}],"review_version":1}