{"id":"e1814c52-0b97-44bd-9f3b-b70249dd1276","arxiv_id":"2502.04836","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In urban large eddy simulations with a turning wind direction, plain time averaging corrupts variance statistics when the averaging time approaches the turning timescale, while 10 to 50 ensemble members with short time averaging restore accuracy.","lead":"This study compares two ways of averaging wind simulations in cities when the wind direction changes over time: time averaging versus ensemble averaging over many runs. It finds that long time averaging badly distorts wind variances, and that 10 to 50 runs with short time averaging give a practical compromise.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10-50 member recommendation rests on treating 648 spatially correlated cube-array units as independent ensemble members; if spatial correlation is non-negligible, the effective ensemble size and convergence estimates are optimistic.","rationale":"The reader's weakest assumption identifies the most load-bearing point. The central qualitative finding — that time averaging over a window comparable to the wind-turning timescale corrupts variance statistics, especially the weaker horizontal component — is robust and visible in both the cube array and the urban case (Figs. 4, 7, 10, 14). It does not depend on the pseudo-ensemble alone: even the 50-member urban ensemble, which is a true ensemble of simulations, shows the same qualitative behavior. However, the paper's practical headline — 'ensembles of 10 to 50 members with 0.131TΩ averaging' — is quantitatively anchored in the cube-array pseudo-ensemble. Since only five independent simulations contribute, the 3,240-member reference is a spatial-temporal mixture whose effective sample size is unknown. The authors state the expectation that time separation mitigates spatial correlation, but they offer no diagnostic. If the 648 repeating units are correlated through large-scale motions, the convergence curve and the 10-50 member recommendation are optimistic; the issue is a statistical-validity problem, not a disagreement with community consensus. I would keep the reader's conditional verdict: the qualitative claim stands, but the quantitative recommendation needs either a cluster-bootstrap reanalysis or a direct spatial-correlation assessment before it can be relied upon. No new concern beyond the reader's is needed, so agreement is complete.","tokens_in":34619,"tokens_out":5667,"duration_ms":64929,"concrete_test":"Recompute the ensemble-size analysis (Fig. 8 and lower half of Table 1) using a cluster bootstrap that resamples whole simulations — draw five simulations with replacement and, within each drawn simulation, sample 648 units — instead of resampling all 3,240 units independently. Compare the interquartile ranges for 10, 25, 50, and 100 members to those in Table 1. Additionally, report the mean spatial correlation of u and v fluctuations between repeating units in the spinup at t/TΩ = 0.526; if the cluster-bootstrap IQRs for 10-50 members materially exceed the original IQRs, the 10-50 recommendation is not supported by the current pseudo-ensemble.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative core of the paper — the recommendation that 10-50 members suffice when 0.131TΩ time averaging is used — is derived from the staggered-cube ensemble of 3,240 'members' that are not independent realizations. Section 2.4 states that five simulations branched from a spinup are each expanded into 648 members by treating each repeating spatial unit as a separate ensemble member, while acknowledging that 'this approach is expected to introduce some spatial correlation among the ensemble members. However, we expect that the time separation between initial conditions mitigates the effect.' That expectation is not demonstrated, and it addresses the wrong axis: time separation decorrelates the five simulations, not the 648 units within a single simulation. In a periodic cube array with coherent large-scale structures (the authors themselves introduced shifted periodic boundaries to weaken large streamwise structures in §2.3.1), the 648 unit-level samples are positively correlated. The bootstrap convergence curve, the Taylor-diagram reference ('true' ensemble statistics), and the resampling experiment in Fig. 8 all treat the 3,240 units as if they carried 3,240 independent pieces of information. If the effective number of independent samples is much smaller, then (i) the reference ensemble variance may understate the true between-realization spread, and (ii) the 10-50 member ensembles constructed by random draws from the pseudo-ensemble inherit a falsely small sampling error, making the recommendation optimistic. The realistic-urban analysis uses a reference built from the same 0.131TΩ averaging that the cube-array study recommends, so it cannot independently validate the compromise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how plain time averaging distorts ensemble statistics in large-eddy simulations of urban flows when the wind direction changes. The authors simulate a staggered cube array with a temporally turning pressure gradient, constructing a 3,240-member ensemble by combining five branched simulations with 648 repeating spatial units, and also simulate a realistic urban area around Turku with a 50-member ensemble. Using vertical profiles, Taylor diagrams, RMSE, and fractional bias, they show that time averaging over intervals comparable to the turning time scale T_Omega severely degrades the variance, especially of the weaker horizontal component, and that shorter averaging (chosen as 0.131 T_Omega) can be combined with modest ensemble sizes of 10-50 members. They further identify building wakes as the regions most affected by long time averaging.","tokens_in":34849,"tokens_out":3727,"duration_ms":42176,"significance":"If the quantitative recommendation holds, the paper provides useful practical guidance for a growing class of nonstationary urban LES studies, and the openly available data set with 50 realistic-urban ensemble members is a valuable resource. The qualitative finding that long time averaging contaminates variances, not only means, is convincingly supported by consistent signals in both the cube-array and Turku cases. The central quantitative claim of the paper, however, rests on two fragile pillars: the treatment of 648 spatially correlated cube-array units as independent ensemble members, and the use of the recommended 0.131 T_Omega averaging window as the reference in the realistic-urban evaluation. Both are acknowledged in the manuscript, but neither is resolved, so the 10-50 member recommendation should be treated as conditional until the independence and circularity concerns are addressed.","major_comments":[{"comment":"The 3,240-member 'ensemble' is not a set of 3,240 independent realizations: it consists of five turning simulations, each expanded into 648 members by treating the repeating spatial units as separate ensemble members. The statement in Section 2.4 that 'time separation between initial conditions mitigates the effect' addresses decorrelation of the five branched simulations, not of the 648 units within a single simulation; in a periodic cube array with coherent large-scale structures, these units are positively correlated, and the shifted periodic boundaries in Section 2.3.1 were introduced precisely to weaken such structures. The bootstrap convergence statement, the reference ensemble statistics used throughout Section 3.1, and the resampling experiment in Fig. 8 and Table 1 all treat the 3,240 units as carrying 3,240 independent pieces of information. If the effective number of independent samples is much smaller, the reference ensemble variance may understate the true between-realization spread and the reported convergence of 10-50 member ensembles is optimistic. Please quantify the spatial correlation among repeating-unit fluctuations (for example, via correlation functions or by repeating the convergence analysis using only the five independent simulations) and, if the correlation is non-negligible, revise the quantitative recommendation accordingly.","section":"Sections 2.4 and 3.1 (Figs. 7-8, Table 1)"},{"comment":"The realistic-urban evaluation is circular to a degree that affects the transferability of the central recommendation. The reference against which all time-averaging intervals are scored is the ensemble mean computed with 0.131 T_Omega time averaging, which is exactly the averaging window recommended in Section 3.1; the text itself acknowledges that this 'can be expected to result in improved performance for at least the 0.131T_Omega averaging time.' The urban case therefore independently demonstrates only that long averaging (0.657-0.920 T_Omega) performs worse than shorter averaging, not that 0.131 T_Omega is the correct threshold or that 10-50 members suffice in a realistic geometry. Please provide an independent reference, for example the instantaneous 50-member ensemble with its sampling noise explicitly characterized, or a subset of members averaged over a window different from the one used in the reference, and state clearly which aspects of the cube-array recommendation the urban case can actually confirm.","section":"Section 3.2 (Fig. 14, Table 2)"}],"minor_comments":[{"comment":"The correlation coefficient formula has an unmatched parenthesis in the numerator: it reads (Co - <Co>)(Cp - <Cp>> and should be (Co - <Co>)(Cp - <Cp>).","section":"Section 2.5, Eq. (11)"},{"comment":"The text discussing Table 2 refers to 'the Taylor diagram in Fig. 8,' but the relevant figure for the realistic urban case is Fig. 14; please correct the cross-reference.","section":"Section 3.2 and Fig. 14 caption"},{"comment":"Panel d) in the caption is labeled 'd) b)' and should be 'd)'; several other typos appear in the text, including 'avaraging', 'chaning', 'waske', 'beheviour', and 'fractioanl'.","section":"Figure 4 caption"},{"comment":"The statement that averaging up to 0.219 T_Omega improves results is not uniformly true across quantities; for example, the mean u component still improves at 0.394 T_Omega in Fig. 7, and the degradation onset differs between means and variances. Please phrase the threshold as quantity-dependent or provide a more granular summary.","section":"Section 3.1, Fig. 7 discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the dataset is a genuine contribution. The qualitative conclusion is likely robust, but the quantitative 10-50 member recommendation is currently supported mainly by a pseudo-ensemble of correlated repeating units and a circular reference in the urban case. I would not reject, because the authors openly acknowledge both limitations and the issues are addressable by additional analysis or by appropriately weakening the quantitative claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should read this paper if you care about how to extract statistics from urban LES when the forcing is nonstationary. The headline result is that plain time averaging severely corrupts the variance of the weaker lateral wind component when the averaging window is comparable to the turning timescale of the pressure gradient, and that a short window (about 0.131 TΩ) combined with an ensemble of 10–50 members is a workable compromise. The qualitative part is convincing.\n\nThe cube-array experiment is the paper's real asset: five simulations are expanded into a 3,240-member 'ensemble' by treating each of the 648 repeating spatial units as an independent member. That is clever but also the main soft spot. The independence assumption is not demonstrated, and the authors' justification (time separation between initial conditions) addresses the five parent simulations, not the 648 simultaneous units. If those units are correlated, the bootstrap convergence curves and the error measures overstate the effective sample size, making the 10–50 recommendation optimistic. The paper acknowledges the concern in Sec. 2.4 but does not quantify it. The good news is that the central qualitative claim also shows up in the realistic urban case with 50 genuine members, so the failure mode is real.\n\nSecond soft spot: the 0.131 TΩ window is chosen after seeing the results, and in the urban case it serves as both the recommended method and the reference against which all other averages are scored. The authors note this in Sec. 3.2, but it means the urban part cannot independently validate the compromise. Also, only one turning rate (Ro≈210) is tested, so we don't know how the recommended window scales with the nondimensional turning rate.\n\nNone of this breaks the paper. It is honestly written, the data are published, and the practical guidance is useful even if the precise member count is not rock solid. For peer review, I would send it out with a request to quantify the spatial correlation (e.g., by subsampling units and checking effective N) and to test at least one other turning rate. That is a reasonable path to a solid contribution.","headline":"The paper delivers a useful practical message about ensemble design for urban LES with changing wind direction, but the specific 10–50 member recommendation rests on pseudo-replicated members and a self-referential reference, so the quantitative core needs revision.","tokens_in":35474,"tokens_out":3863,"would_cite":true,"duration_ms":36595,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When the wind direction turns during a simulation, plain time averaging of urban large-eddy simulations can severely distort both the mean wind and its variance, so ensemble averaging with a short time window is needed.","keywords":["large eddy simulation","ensemble averaging","time averaging","nonstationary urban flow","turning wind direction","staggered cube array","building wakes","Taylor diagram"],"falsifier":"Run the same turning-pressure-gradient cube-array case twice: once with the repeating-unit construction and once with roughly 20 fully independent simulations started from uncorrelated turbulent fields, and compare member-to-member variance and ensemble means at $t = 0.526\\,T_\\Omega$. Substantially larger spread or a shifted mean in the independent ensemble would show that the 3,240-member reference is not unbiased, and the recommended ensemble sizes would need revision.","tokens_in":34354,"feed_emoji":"🌬️","tokens_out":8864,"duration_ms":87502,"temperature":0.7,"pith_summary":"Urban wind simulations at metre-scale resolution are increasingly used for flows whose forcing changes over time, and when the wind direction turns, a simple time average is not a valid substitute for an average over many independent realizations. This paper establishes, from large-eddy simulations of a staggered cube array and of a real city district, that plain time averaging can seriously contaminate both the mean wind and its variance once the averaging window is comparable to the turning time scale of the driving pressure gradient. The errors are largest for the weaker horizontal velocity component and inside building wakes, which are precisely the regions used for pollutant-dispersion and pedestrian-comfort assessments. The paper shows that a practical remedy is to combine ensemble averaging with a short time average, about 13% of the turning time scale, so that ensembles of 10–50 members reproduce the statistics of a much larger reference ensemble.","feed_headline":"Time averaging fails when urban wind direction turns","feed_subtitle":"Meter-scale city simulations need ~13% turning-time averaging and 10–50 members for trustworthy wind statistics.","key_machinery":"The load-bearing construction is a set of large-eddy simulations branched off a fully developed, constant-direction flow at intervals larger than the integral time scale, so each member begins from a distinct turbulent state. For the cube array, the 648 repeating spatial units of each of five runs are counted as additional members, giving 3,240 in total; for the real city, 50 runs provide one member each. The driving force is a pressure gradient of constant magnitude rotating at $\\Omega = 15^\\circ\\,\\mathrm{h}^{-1}$, defining $T_\\Omega = 1/\\Omega \\approx 230$ minutes and a modified Rossby number of about 210 (the ratio of the turning time scale to the bulk-flow time scale). Agreement between time-averaged and ensemble-averaged fields is quantified with Taylor diagrams, normalized standard deviation, correlation, normalized RMSE, and fractional bias, and bootstrap resampling is used to test convergence with ensemble size. The decisive comparison is among averaging windows from $0.0438\\,T_\\Omega$ to $0.920\\,T_\\Omega$; the $0.131\\,T_\\Omega$ window is adopted for the final recommendation.","core_discovery":"The paper's central claim is that with a pressure gradient rotating at rate $\\Omega$, the characteristic turning time $T_\\Omega = 1/\\Omega$ controls how long a time average can be before it corrupts the statistics. Plain time averaging over windows of order $T_\\Omega$ folds the changing wind direction into the mean and, more severely, into the variance, inflating the variance of the weaker lateral component and displacing mean winds; the damage is concentrated in building wakes. Against a reference ensemble of 3,240 members for the cube array and a 50-member ensemble for a real urban district, the paper shows that time averaging over roughly $0.13\\,T_\\Omega$ improves agreement with the ensemble statistics, while longer windows degrade it, and that 10–50 ensemble members then suffice in the roughness sublayer. The conclusion is that plain time averaging should be avoided for nonstationary urban LES, and that short-time-averaged ensembles offer an accurate, affordable compromise.","pith_inferences":["A testable extension is to check whether the same ratio of averaging window to forcing time scale governs other nonstationary forcings, such as a changing pressure-gradient magnitude; the paper does not simulate that case.","The repeating-unit ensemble's independence assumption could be validated directly by comparing its member variance with fully independent realizations; if spatial correlation is sizable, 10–50 members may be optimistic.","Since errors concentrate in building wakes, pollutant-concentration statistics from a single time-averaged run would be biased in exactly the locations where exposure estimates matter, so dispersion studies should report ensemble spread rather than time-mean fields.","An implicit engineering rule follows: record the characteristic forcing time scale, keep averaging windows an order of magnitude below it, and size the ensemble by bootstrap convergence rather than defaulting to large member counts."],"forward_implications":["Plain time averages over long windows should be treated as unreliable for nonstationary urban LES, particularly for variances of the weaker horizontal velocity component.","Combining ensemble averaging with a time window of about $0.13\\,T_\\Omega$ makes 10–50 ensemble members sufficient to approximate the statistics of far larger ensembles.","Building wakes are the regions where time-averaging errors concentrate, so wake-sensitive applications such as pollutant dispersion and pedestrian-level wind studies should use ensemble statistics.","Above the roughness sublayer, horizontal spatial averaging can replace ensemble and time averaging when the flow is horizontally homogeneous.","The turning time scale $T_\\Omega = 1/\\Omega$ provides a practical rule for choosing the averaging window before a nonstationary urban LES campaign begins."],"supporting_citations":[{"why":"Supplies the definition and the premise that time averages are not invariant under time shifts in nonstationary turbulence.","marker":"Pope (2000)"},{"why":"Prior urban LES ensemble with pulsatile forcing and 15,360–192,000 members; provides the comparison case for unsteady canopy flow.","marker":"Li and Giometto (2023)"},{"why":"60-member urban LES dispersion ensemble that estimated roughly 200 members may be needed; motivates the ensemble-size question.","marker":"Harms et al. (2011)"},{"why":"Forest-edge LES showing a ten-member ensemble with 15-minute time averages is sufficient; the precedent for the time-average/ensemble compromise.","marker":"Kanani et al. (2014)"},{"why":"Defines the statistical performance measures (NSD, R, NRMSE, fractional bias) used to compare time-averaged to ensemble-averaged fields.","marker":"Chang and Hanna (2004)"},{"why":"Introduces the Taylor diagram used to display normalized standard deviation, correlation, and normalized RMSE.","marker":"Taylor (2001)"},{"why":"Provides the bootstrap approach used to assess statistical convergence of the 3,240-member ensemble.","marker":"Efron and Tibshirani (1993)"},{"why":"Describes the CMIP6 ensemble construction by branching from a control run, the approach used to generate the LES ensembles.","marker":"Eyring et al. (2016)"}],"fun_headline_variants":["Time averaging fails for turning urban winds","Turning winds break time-averaged city simulations","Ensemble wins when urban wind direction shifts","Short averages fix wind stats in urban LES"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reference 'true' ensemble for the cube array assumes the 648 repeating spatial units of each of five simulations behave as independent members, and the paper's expectation that temporal separation of initial conditions keeps spatial correlation negligible is the load-bearing premise; if that correlation is not small, the quoted errors and the 10–50-member recommendation are biased toward spatially homogeneous statistics.","fun_headline_variants_meta":{"raw":{"variants":["Time averaging fails for turning urban winds","Turning winds break time-averaged city simulations","Ensemble wins when urban wind direction shifts","Short averages fix wind stats in urban LES"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1194,"prompt_tokens":920,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":219}},"tokens_in":536,"tokens_out":274,"duration_ms":3579,"temperature":1.0,"reasoning_tokens":219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T21:18:29.673667+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same turning-pressure-gradient cube-array case twice: once with the repeating-unit construction and once with roughly 20 fully independent simulations started from uncorrelated turbulent fields, and compare member-to-member variance and ensemble means at $t = 0.526\\,T_\\Omega$. Substantially larger spread or a shifted mean in the independent ensemble would show that the 3,240-member reference is not unbiased, and the recommended ensemble sizes would need revision.","supporting_citations":[{"cited_title":"Cambridge University Press","cited_arxiv_id":null,"evidence_quote":"Supplies the definition and the premise that time averages are not invariant under time shifts in nonstationary turbulence."},{"cited_title":"Journal of Fluid Mechanics 974:A33, doi:10.1017/jfm.2023.801","cited_arxiv_id":null,"evidence_quote":"Prior urban LES ensemble with pulsatile forcing and 15,360–192,000 members; provides the comparison case for unsteady canopy flow."},{"cited_title":"Journal of Wind Engineering and Industrial Aerodynamics 99(4):289--295, doi:10.1016/j.jweia.2011.01.007","cited_arxiv_id":null,"evidence_quote":"60-member urban LES dispersion ensemble that estimated roughly 200 members may be needed; motivates the ensemble-size question."},{"cited_title":"Meteorologische Zeitschrift pp 33--49, doi:10.1127/0941-2948/2014/0542","cited_arxiv_id":null,"evidence_quote":"Forest-edge LES showing a ten-member ensemble with 15-minute time averages is sufficient; the precedent for the time-average/ensemble compromise."},{"cited_title":"Meteorology and Atmospheric Physics 87(1):167--196, doi:10.1007/s00703-003-0070-7","cited_arxiv_id":null,"evidence_quote":"Defines the statistical performance measures (NSD, R, NRMSE, fractional bias) used to compare time-averaged to ensemble-averaged fields."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the bootstrap approach used to assess statistical convergence of the 3,240-member ensemble."},{"cited_title":"Geoscientific Model Development 9:1937--1958, doi:10.5194/gmdd-8-10539-2015","cited_arxiv_id":null,"evidence_quote":"Describes the CMIP6 ensemble construction by branching from a control run, the approach used to generate the LES ensembles."}],"review_version":1}