{"id":"448f4c80-df00-4161-97d7-4fc695e77276","arxiv_id":"2506.19186","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A minimax analysis identifies the optimal importance-sampling proposal for atomic targets, and an exact uniform ergodicity criterion is proved for importance-tempered random-walk Metropolis on polynomial-tail targets.","lead":"This paper proves when the target distribution itself is the best importance-sampling proposal, and when tempering the target and reweighting makes the MCMC sampler uniformly ergodic. The second result gives an exact temperature range for polynomial-tailed one-dimensional targets.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4's necessity proof rests on a limsup/Fatou step that is invalid as stated; the 'only if' direction is not yet established.","rationale":"The reader's weakest_assumption concerns the transfer from the continuous-time chain to the discrete self-normalized estimator, which is a real practical gap and is partly conceded in Remark 5. I agree with that concern but do not think it is the most load-bearing issue for the central claim. The reader's rationale also flags the invalid Fatou step in Theorem 4, and that is the step I would put first: it blocks the necessity half of the iff theorem. I also share the reader's note that Theorem 1 assumes an atom attaining the essential supremum, which fails for infinite atom sets unless an approximation argument is added. Both issues are specific and plausibly repairable, and the sufficiency half and the minimax construction are otherwise well supported, so the CONDITIONAL verdict is unchanged. My agreement is partial because the reader's formal weakest_assumption and my primary concern differ, even though they overlap in the rationale.","tokens_in":22427,"tokens_out":21303,"duration_ms":239929,"concrete_test":"Repair the limit passage without the invalid limsup inequality: use uniform ergodicity of the T-skeleton to get P_x(τ_D>t)≤Cρ^t uniformly in x, and with κ supported on [-ξ,ξ] show E_x[V(X_t);τ_D>t]→0 for each fixed x (e.g., via Cauchy-Schwarz and the subexponential growth of log(x+ξN_t)). If this succeeds, replace the Fatou step by dominated convergence and verify that the resulting bound E_xτ_D≥[V(x)-log(1+D)]/α contradicts sup_x E_xτ_D<∞. If no such bound can be established, exhibit initial states for which the missing term does not vanish; that would confirm the current proof is irreparable as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the proof of Theorem 4 (Section 3.3), after Dynkin's formula the paper asserts limsup_t E_x[V(X_{t∧τ_D})] ≤ E_x[limsup_t V(X_{t∧τ_D})] = E_x[V(X_{τ_D})]. This direction of Fatou's lemma requires uniform integrability or a domination argument, and none is supplied. The uncontrolled term is E_x[V(X_t); τ_D>t]: it must be shown to vanish as t→∞ before one may conclude log(1+D) ≥ E_x[V(X_{τ_D})] and hence contradict sup_x E_xτ_D<∞. As written, the inequalities only give a lower bound on liminf E_x[V(X_{t∧τ_D})], which is compatible with the desired conclusion and yields no contradiction. Since Theorem 4's necessity is the sole support for the 'only if' direction of the main ergodicity claim, this is a load-bearing correctness gap. Uniform ergodicity may well supply the missing control through an exponential tail for τ_D and the bounded-jump structure of the chain, but that argument is not present in the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies two related problems. First, for self-normalized importance sampling, it proves a minimax characterization: the target distribution is the minimax optimal trial distribution if and only if no atom carries more than half the total mass; when a large atom exists, the optimal trial downweights that atom to probability 1/2. The paper also gives a continuous-space version for targets concentrated on a small set, and a negative result for multiple importance sampling. Second, it analyzes 'importance-tempered' MCMC, where the chain is run with stationary density proportional to pi(x)^beta and the bias is corrected by importance weights. The continuous-time version of the chain is shown, under a drift condition with bounded concave Lyapunov functions, to be uniformly ergodic for super-exponential targets (any beta in (0,1)) and for polynomial-tailed targets pi(x) proportional to (1+|x|)^{-gamma} when 1/gamma < beta < (gamma-2)/gamma. The main iff result is Theorem 4, which claims necessity of the upper bound beta < (gamma-2)/gamma. Numerical experiments illustrate the absence of burn-in and variance reduction.","tokens_in":22583,"tokens_out":17450,"duration_ms":178933,"significance":"If fully established, the results would be a valuable contribution to both importance sampling theory and MCMC convergence analysis. The minimax characterization in Theorems 1-2 is clean and apparently novel, and the large-atom result gives a concrete, non-obvious prescription for trial distributions. The continuous-time drift analysis in Theorem 3 is a useful technique: using bounded concave Lyapunov functions to exploit the importance-weight factor in the generator is an original idea, and Proposition 4's sufficient condition for uniform ergodicity of heavy-tailed targets is a strong result. The paper is also well structured: the main derivations are from stated assumptions with no fitted constants, and the numerical studies are consistent with the theory. However, the necessity proof of Theorem 4 has a load-bearing gap (an unjustified Fatou/limsup step), and the proof of Theorem 1 has a technical gap for countably infinite atom spaces. Both are plausibly repairable, but the manuscript is not yet fully rigorous as written.","major_comments":[{"comment":"The proof uses the step limsup_t E_x[V(X_{t∧τ_D})] ≤ E_x[limsup_t V(X_{t∧τ_D})] = E_x[V(X_{τ_D})], citing Fatou's lemma. For nonnegative random variables, Fatou's lemma gives E[liminf] ≤ liminf E, not the limsup inequality used here. The limsup version requires uniform integrability or a domination argument, and none is supplied. In particular, the regime τ_D>t contributes E_x[V(X_t); τ_D>t], which need not vanish as t→∞ without additional control. As written, the argument only yields a lower bound on liminf E_x[V(X_{t∧τ_D})], which is compatible with the desired contradiction and does not establish log(1+D) ≥ E_x[V(X_{τ_D})]. Since this is the sole support for the 'only if' direction of Theorem 4, it is a load-bearing correctness gap. A possible repair is to use uniform ergodicity to obtain an exponentially decaying bound on P_x(τ_D>t) and the bounded-jump structure to control V(X_t) on the survival event, or to replace the necessity argument with a direct drift-based proof.","section":"Section 3.3, proof of Theorem 4, after Eq. (29)"},{"comment":"The proof asserts that, when the essential supremum of w on the non-atomic part is below 1, there exists an atom x* with w(x*) = ess sup_{x∈X} w(x) ≥ 1. For a countably infinite set of atoms this is not guaranteed: the weights on atoms can approach a supremum without attaining it (e.g., w(a_n) = 2 − 1/n). The subsequent argument, including the inequality Π|_{E^c}(w) ≤ w(x*) and the decomposition before Eq. (7), requires that w(x*) be the global maximum of w. The gap is fixable by taking a sequence of atoms with weights approaching the essential supremum and passing to the limit, but as written the proof of Theorem 1 is incomplete. Since Theorem 1 is a central result of Section 2, this requires repair; note also that the proof of Proposition 1 in Section 2.3 repeats the same attainment argument and inherits the issue.","section":"Section 2.2, proof of Theorem 1"},{"comment":"The ergodicity theorems (Theorem 4 and Propositions 3-4) are proved only for proposal densities with compact support, specifically the truncated normal in Eq. (25) and Condition (ii) of Theorem 3. The introduction and abstract, however, present the result for 'the Metropolis--Hastings algorithm' generally, and the numerical experiments in Section 4 use an untruncated N(0, 3^2) proposal. Remark 3 states that the truncation is a technical convenience and that the untruncated case can be handled with additional constraints on D, but no proof or precise statement is provided. This is a scope gap between the theoretical claims and the motivating application. I ask the authors to either supply the extension to unbounded symmetric proposals or explicitly state that the theorem is for bounded-support proposals and adjust the abstract and claims accordingly.","section":"Section 3.3 and Remark 3"}],"minor_comments":[{"comment":"The definitions 'wX = π/qY and wY = π/qY' should be 'w_X = π/q_X' and 'w_Y = π/q_Y'.","section":"Section 2.3, paragraph before Eq. (11)"},{"comment":"'In Proposition 2, we allow polynomially decaying tails' should refer to Proposition 4, not Proposition 2.","section":"Section 3.3, paragraph before Proposition 3"},{"comment":"The drift condition is stated for '∀ x∈(−∞,D]∪[D,∞)', which is all of R; the intended statement is clearly '∀ x∈(−∞,−D]∪[D,∞)'.","section":"Lemma 4 statement"},{"comment":"The citation for the stereographic projection sampler appears to be incorrect: reference [38] is Yang, Wainwright and Jordan (2016), while the stereographic sampler is discussed in [37] (Yang, Latuszyński and Roberts, 2024). Please clarify the citation.","section":"Remark 5"},{"comment":"The argument that uniform ergodicity of the T-skeleton chain implies sup_x E_x[τ_D] < ∞ for the continuous-time chain is plausible but needs a brief justification: the skeleton chain may jump over the compact set [-D,D] between skeleton times, so one should consider hitting a slightly enlarged set such as [-D-ξ, D+ξ].","section":"Proof of Theorem 4, application of Meyn--Tweedie Theorem 16.2.2"},{"comment":"The abstract's claim that importance tempering can 'essentially eliminate the need for burn-in' should be read in light of Remark 5, which correctly notes that uniform ergodicity of the continuous-time chain Y_t does not directly imply a uniform finite-sample bound for the discrete estimator eΠ_{β,n}(f) when initialization is poor. Consider adding a qualifier in the abstract to avoid overstating the practical conclusion.","section":"Abstract and Remark 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the main ideas are likely correct, but the two proof gaps identified above are in central results and should be repaired before publication. The Theorem 4 necessity proof is the more serious issue; if the authors can supply the missing uniform-integrability argument, the paper would be a solid contribution. I would not recommend rejection, as the gaps appear fixable without changing the main claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2506.19186. First, the minimax importance-sampling result for atomic targets (Theorem 2) is correct and clean: the construction that downweights a large atom to 1/2 and the matching lower bound via a two-point test function are elementary and convincing. Second, the uniform-ergodicity threshold for polynomial tails (Theorem 4) is a real advance over Livingstone et al. and stereographic MCMC, but the necessity proof has a gap. After Dynkin's formula, the paper applies a limsup/Fatou step to the unbounded function V = log(1+|x|) with no domination argument, so the displayed inequality does not follow. As written, the 'only if' direction is not established. That is load-bearing, not cosmetic.\n\nWhat the paper does well: the sufficiency direction via drift conditions with bounded concave V is clever and, as far as I can tell, correct. Theorem 3's state-dependent drift rate proportional to |V''|/w is a nice idea, and the proofs of Propositions 3 and 4 look sound. The minimax analysis for continuous targets concentrated on a small set (Proposition 2) is a useful extension. The paper is also honest about the gap between the continuous-time chain and the discrete estimator (Remark 5) and about the one-dimensional scope.\n\nSoft spots beyond the Fatou issue. Theorem 1 assumes the essential supremum of the importance weight is attained at an atom; with countably many atoms that need not happen. This looks repairable by a limiting argument using Lemma 3 along a sequence of atoms, but the current text is wrong. The abstract's burn-in claim overstates what is transferred to the discrete self-normalized estimator; Remark 5 essentially concedes this, so the abstract should be softened. The single self-citation for the variance comparison is peripheral and not a problem.\n\nWho is this for? Monte Carlo methodologists and MCMC theory people. The minimax part is citable, and the sufficiency part of Theorem 4 is a genuine contribution. But the main iff claim is not yet proved. I would send this to peer review—the flaws are specific and likely fixable—but a referee should demand a repaired necessity proof, either via a domination argument using the exponential hitting-time bound or a different contradiction. If the author can fix that, this is a solid journal contribution.","headline":"Two solid contributions, but the 'only if' direction of the main ergodicity theorem rests on a fixable Fatou gap.","tokens_in":23124,"tokens_out":4914,"would_cite":true,"duration_ms":46801,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60J25","65C05","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that the target is the minimax importance-sampling proposal unless an atom exceeds 1/2, and that importance-tempered MCMC is uniformly ergodic for polynomial-tail targets exactly when 1/γ<β<(γ−2)/γ.","keywords":["importance sampling","minimax optimality","Metropolis-Hastings","importance tempering","uniform ergodicity","heavy-tailed distributions","continuous-time Markov chain","burn-in"],"falsifier":"Run the continuous-time chain $(Y_t)$ for the target $\\pi(x)\\propto(1+|x|)^{-5}$ with a random-walk proposal and $\\beta=0.55$ versus $\\beta=0.65$; the theorem says the first is uniformly ergodic and the second is not, so the total-variation distance from stationarity should decay at a rate independent of $Y_0$ for $\\beta=0.55$ but not for $\\beta=0.65$. For the independent-sampling half, compute the worst-case asymptotic variance over zero-mean unit-variance $f$ when one atom has mass $0.8$: the predicted minimum, $4(0.8)(0.2)=0.64$, is attained exactly by the trial with atom mass $1/2$.","tokens_in":22172,"feed_emoji":"🎯","tokens_out":13277,"duration_ms":126188,"temperature":0.7,"pith_summary":"The paper answers two questions that share one intuition: downweight the regions where the target is too concentrated. For independent importance sampling, it proves that the target itself is the minimax trial distribution if and only if it has no atom with probability greater than $1/2$; a heavier atom should be cut down to mass $1/2$ in the proposal. For Markov chain sampling, it proves that running Metropolis--Hastings on the tempered density $\\pi(x)^\\beta$ and correcting with importance weights $\\pi(x)^{1-\\beta}$ yields a chain that is uniformly ergodic for polynomial-tail targets $\\pi(x)\\propto (1+|x|)^{-\\gamma}$ exactly when $1/\\gamma<\\beta<(\\gamma-2)/\\gamma$, which is possible precisely for $\\gamma>3$. If correct, this makes burn-in unnecessary for such targets and can reduce the variance of time-average estimators substantially. The practical force of the MCMC claim depends on how faithfully the continuous-time chain's uniform ergodicity transfers to the discrete self-normalized estimator actually used, a limitation the paper flags.","feed_headline":"Tempered MCMC can erase burn-in for heavy-tailed targets","feed_subtitle":"A narrow temperature window makes the chain forget its start, so time averages can use the whole run.","key_machinery":"Two objects carry the argument. The first is the worst-case risk functional $R(Q,L^2(\\Pi))=\\sup_f \\Pi(f^2w)$ over functions with $\\Pi(f)=0$ and $\\Pi(f^2)=1$, where $w=\\pi/q$ is the importance weight; on atomless spaces this reduces to the essential supremum of $w$, while on finite spaces it is either the largest weight or the unique root of the equation $\\sum_i \\pi_i/(w_i-\\lambda)=0$, and minimizing it yields the 'cut the heavy atom to $1/2$' rule. The second is the continuous-time chain $(Y_t)$ whose embedded chain is the tempered Metropolis--Hastings chain and whose holding time at $x$ is exponential with mean $w(x)\\propto \\pi(x)^{1-\\beta}$; its generator is $(Ag)(x)=w(x)^{-1}\\int[g(y)-g(x)]T(x,dy)$. Uniform ergodicity is proved through drift conditions $(AV)(x)\\le -\\alpha V(x)$ outside a compact set, with $V$ bounded, nondecreasing, and strictly concave on each tail: Theorem 3 shows the drift rate is controlled by $|V''(x)|/w(x)$, so the importance weight drives the chain out of the tails. The necessity of $\\beta<(\\gamma-2)/\\gamma$ is proved with the unbounded drift $V(x)=\\log(1+|x|)$, showing that no bounded drift can work at or beyond the threshold.","core_discovery":"The paper establishes two characterizations. First, for independent sampling, the worst-case asymptotic variance of the self-normalized importance sampling estimator over all zero-mean, unit-variance functions is minimized by using the target as the trial distribution if and only if no atom carries probability greater than $1/2$; when an atom has mass $p>1/2$, the minimax trial puts probability $1/2$ on that atom and rescales the target density on the remaining space, lowering the worst-case risk to $4p(1-p)$. An analogous near-optimal construction handles continuous targets concentrated on a small set. Second, for importance-tempered MCMC, the paper shows that the continuous-time chain built from a random-walk Metropolis--Hastings chain with stationary density $\\pi(x)^\\beta$ is uniformly ergodic for $\\pi(x)\\propto(1+|x|)^{-\\gamma}$ if and only if $1/\\gamma<\\beta<(\\gamma-2)/\\gamma$. Such a $\\beta$ exists exactly when $\\gamma>3$, and uniform ergodicity means the chain forgets its starting point at a rate independent of the start, so time averages can be computed without discarding burn-in.","pith_inferences":["The same minimax logic suggests a practical recipe for Bayesian posteriors: estimate a high-probability credible set $A$ and use a trial density that puts about half its mass on $A$, reweighting the complement; the paper proves near-optimality for functions controlled on $A$, but choosing the set and the exact split is left to the user.","The sharp window $1/\\gamma<\\beta<(\\gamma-2)/\\gamma$ suggests a testable multivariate extension for spherically symmetric heavy-tailed targets, where the upper threshold should shift with dimension; the paper does not treat $d>1$.","Uniform ergodicity of $Y_t$ does not guarantee fast variance decay for the discrete estimator from a cold start, since the number of jumps before entering the central region can still be unbounded; a safe implementation would monitor accumulated importance weight early in the run.","For polynomial tails with $\\gamma>4$, the value $\\beta=1/2$ lies inside the ergodicity window, so the theory points to a simple default temperature in that regime; the paper does not draw this recommendation."],"forward_implications":["For any target, a practitioner who cares about worst-case error over all square-integrable functions can safely use the target itself as the trial unless an atom exceeds $1/2$; for a heavy atom of mass $p$, the minimax proposal is explicit: put probability $1/2$ on that atom and renormalize the target density elsewhere.","For a continuous posterior concentrated in a set $A$ with $\\Pi(A)$ close to one, a trial density that puts roughly half its mass on $A$ has worst-case asymptotic variance of order $1-\\Pi(A)$ for functions with small oscillation on $A$, compared with unit variance for direct sampling.","For one-dimensional polynomial-tail targets with $\\gamma>3$, choosing $\\beta\\in(1/\\gamma,(\\gamma-2)/\\gamma)$ makes the continuous-time importance-tempered chain uniformly ergodic, so its long-run time averages are insensitive to the initial state and burn-in can be discarded.","Within that window the importance weight $\\pi(x)^{1-\\beta}$ makes the generator's drift rate large in the tails, so the continuous-time chain spends bounded expected time outside any fixed central interval and the whole trajectory can be used.","Mixing two self-normalized estimators built from different trial distributions cannot beat the minimax trial's worst-case risk, so the explicit optimal proposal is not improved by averaging trial distributions."],"supporting_citations":[{"why":"Supplies the classical division of an atomless probability space into two sets of equal measure, used in the proof that the target is minimax when there is no large atom.","marker":"[32]"},{"why":"Provides the delta-method formula for the asymptotic variance of self-normalized importance sampling and the standard benchmark of sampling directly from the target.","marker":"[26]"},{"why":"Introduces the continuous-time embedding of a Markov chain with importance-weight holding times that Section 3.1 uses as the proxy for estimator efficiency.","marker":"[39]"},{"why":"Supplies the comparison lemma showing the time-average estimator's asymptotic variance is at least that of the self-normalized estimator, justifying the continuous-time proxy.","marker":"[40]"},{"why":"Provides the general theory converting the drift condition $(AV)\\le-\\alpha V$ outside a compact set into uniform ergodicity and exponential hitting-time bounds.","marker":"[6]"},{"why":"Establishes that random-walk Metropolis--Hastings on the real line cannot be uniformly ergodic, the baseline that makes the tempered chain's uniform ergodicity meaningful.","marker":"[23]"},{"why":"Quantifies the polynomial total-variation convergence rate of random-walk Metropolis--Hastings for heavy-tailed targets, the comparison behind the conclusion for $\\gamma>3$.","marker":"[10]"},{"why":"Gives the theorem that uniform ergodicity implies bounded mean hitting times, used in the proof that $\\beta\\ge(\\gamma-2)/\\gamma$ rules out uniform ergodicity.","marker":"[24]"}],"fun_headline_variants":["Importance-tempered MCMC erases burn-in for heavy tails","Tempered MCMC: no burn-in for heavy-tailed targets","Forget burn-in: importance tempering for heavy tails","Tempered MCMC forgets its start: burn-in gone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that uniform ergodicity of the continuous-time chain, whose holding times are random importance weights, reliably describes the discrete self-normalized estimator used in practice; the paper treats the continuous-time convergence rate as a proxy and concedes that with a poor initialization the discrete estimator's variance can still decay slowly in the number of jumps.","fun_headline_variants_meta":{"raw":{"variants":["Importance-tempered MCMC erases burn-in for heavy tails","Tempered MCMC: no burn-in for heavy-tailed targets","Forget burn-in: importance tempering for heavy tails","Tempered MCMC forgets its start: burn-in gone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00113,"raw_usage":{"total_tokens":4741,"prompt_tokens":1034,"completion_tokens":3707,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":3633}},"tokens_in":650,"tokens_out":3707,"duration_ms":27570,"temperature":1.0,"reasoning_tokens":3633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:37:09.138285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the continuous-time chain $(Y_t)$ for the target $\\pi(x)\\propto(1+|x|)^{-5}$ with a random-walk proposal and $\\beta=0.55$ versus $\\beta=0.65$; the theorem says the first is uniformly ergodic and the second is not, so the total-variation distance from stationarity should decay at a rate independent of $Y_0$ for $\\beta=0.55$ but not for $\\beta=0.65$. For the independent-sampling half, compute the worst-case asymptotic variance over zero-mean unit-variance $f$ when one atom has mass $0.8$: the predicted minimum, $4(0.8)(0.2)=0.64$, is attained exactly by the trial with atom mass $1/2$.","supporting_citations":[{"cited_title":"Sur les fonctions d’ensemble additives et continues","cited_arxiv_id":null,"evidence_quote":"Supplies the classical division of an atomless probability space into two sets of equal measure, used in the proof that the target is minimax when there is no large atom."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the delta-method formula for the asymptotic variance of self-normalized importance sampling and the standard benchmark of sampling directly from the target."},{"cited_title":"Scalable importance tempering and Bayesian variable selection","cited_arxiv_id":null,"evidence_quote":"Introduces the continuous-time embedding of a Markov chain with importance-weight holding times that Section 3.1 uses as the proxy for estimator efficiency."},{"cited_title":"Rapid convergence of informed importance tempering","cited_arxiv_id":null,"evidence_quote":"Supplies the comparison lemma showing the time-average estimator's asymptotic variance is at least that of the self-normalized estimator, justifying the continuous-time proxy."},{"cited_title":"Exponential and uniform ergodicity of Markov processes","cited_arxiv_id":null,"evidence_quote":"Provides the general theory converting the drift condition $(AV)\\le-\\alpha V$ outside a compact set into uniform ergodicity and exponential hitting-time bounds."},{"cited_title":"Rates of convergence of the Hastings and Metropolis algorithms","cited_arxiv_id":null,"evidence_quote":"Establishes that random-walk Metropolis--Hastings on the real line cannot be uniformly ergodic, the baseline that makes the tempered chain's uniform ergodicity meaningful."},{"cited_title":"Convergence of heavy-tailed Monte carlo Markov chain algorithms","cited_arxiv_id":null,"evidence_quote":"Quantifies the polynomial total-variation convergence rate of random-walk Metropolis--Hastings for heavy-tailed targets, the comparison behind the conclusion for $\\gamma>3$."}],"review_version":1}