{"id":"f3241142-be48-4d50-8df4-7b3241708482","arxiv_id":"2504.16696","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In simulations, meta-regressions that jointly control location and time heterogeneity outperform specifications that control only one dimension, but standard study-level models remain the strongest performers.","lead":"This simulation study compares ten meta-regression estimators under joint location and time heterogeneity. It finds that models controlling both dimensions beat models controlling only one, while standard study-level models perform best overall.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract overstates the central claim: the paper's own case-level results show joint controls do not beat non-joint controls under small time effects, large location effects, or more locations.","rationale":"The reader's external-validity concern is real, but the more immediate problem is internal: the paper's own Sections 5.3.2 and 5.3.3 describe conditions under which non-joint specifications perform as well as joint ones. This directly limits the unqualified abstract claim and the resulting decision rule. A secondary compounding issue is that Section 5.1 justifies the design by claiming the random-effects assumption is violated because covariance-matrix elements are nonzero, but the RE orthogonality condition concerns correlation between the random effect and regressors, and the stated DGP has common covariate means across studies, making the justification a non sequitur. Both issues support keeping a conditional verdict rather than full acceptance. On the credit side, the qualitative pattern for large time effects is coherent, and the paper is transparent enough to report the contradicting case results even if the abstract does not carry the caveat. There is no formal verification, and the code link is a placeholder, so a rerun is the natural settlement check.","tokens_in":41409,"tokens_out":16514,"duration_ms":167312,"concrete_test":"Re-run the simulation (or re-analyze the stored output) and, for each of the 12 cases in Table 2, compute the difference in the primary performance criteria (average power and estimator variance) between the best joint specification (FElt or FEl,Trend) and the best non-joint specification (FEl or FEt), with Monte Carlo standard errors. If for cases with small time effects (Cases 1-6) or larger location counts (Cases 5-6, 11-12) the difference is not significantly positive, the abstract's unqualified claim is not supported and must be conditionalized to large time effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3.3 reports: 'Unexpectedly, in the face of more locations non-joint methodology specifications performed just as well on average as joint methodology specifications.' Section 5.3.2 similarly reports that when the location effect is large, 'all methodology specifications performed just as well' (excluding FEt and FEs,Trend). These are conditions inside the paper's own 12-case design in which both location and time heterogeneity are present. The unqualified abstract claim—'jointly modeling heterogeneity when heterogeneity is in both dimensions improves performance compared to modeling only one type'—is therefore not a summary of the simulation as a whole. The case-level evidence supports joint > single mainly when the time effect is large (Section 5.3.1). Because the strongest_claim converts this into a blanket decision rule ('practitioners should prefer study-level models or joint controls over location-only or time-only specifications when both dimensions vary'), that rule would be applied in small-time-effect or large-location settings where the paper's own results show no benefit of joint controls. The load-bearing condition—that the advantage of joint controls is uniform across the simulated heterogeneity space—is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses a Monte Carlo simulation with subject-level data to compare ten meta-regression specifications: study-level random, fixed, and mixed effects; location-level random, fixed, and mixed effects; time-level fixed effects; two-way fixed effects; and study-level and location-level fixed effects with a trend. The simulation embeds location heterogeneity as an additive location-specific intercept and time heterogeneity as a common linear trend, with a perfectly balanced design of one study per location-year. Twelve cases vary the size of the location and time effects and the number of locations. The abstract claims that jointly modeling location and time heterogeneity improves performance relative to modeling only one dimension, and the paper recommends REs, FEs, or the newly proposed FEl,Trend for practitioners. The paper's own results, however, include conditions in which joint and non-joint specifications perform equally, and several load-bearing details (Type I error adjustment, trend rescaling, and numerical summaries) are not specified.","tokens_in":41644,"tokens_out":8938,"duration_ms":81107,"significance":"If the findings were established quantitatively, the paper would address a genuine gap: most meta-regression simulations consider only study-level heterogeneity, while applied meta-analyses often span many countries and years. The subject-level DGP with known true coefficients, the inclusion of a newly proposed location-level fixed-effects-with-trend specification (FEl,Trend), and the step-by-step model-selection guide are useful contributions. However, the strength of the contribution is conditional: the central ranking claim is not uniformly supported by the authors' own case-level results, and the absence of numeric Monte Carlo tables and the unspecified Type I error adjustment make it difficult to assess the magnitude and reliability of the reported differences. With the claims properly qualified and the missing details supplied, the paper could be a useful reference for applied meta-analysts.","major_comments":[{"comment":"The unqualified claim in the Abstract and in Section 1 that 'jointly modeling heterogeneity when heterogeneity is in both dimensions improves performance compared to modeling only one type of heterogeneity' is not a summary of this simulation. Section 5.3.2 reports that when the location effect is large, 'all methodology specifications performed just as well' (excluding FEt and FEs,Trend), and Section 5.3.3 reports that 'in the face of more locations non-joint methodology specifications performed just as well on average as joint methodology specifications.' The evidence presented supports the joint-over-single ranking mainly when the time effect is large (Section 5.3.1). The central claim should be revised to state these conditions explicitly and should not be formulated as a blanket decision rule for practitioners.","section":"Abstract; Sections 5.3.1-5.3.3"},{"comment":"The statement that the random-effects assumption is violated because 'each element of the covariance matrix is within [0.2,0.5], so the random effects assumption of 0 correlation is not designed to hold' is incorrect. The random-effects assumption requires zero correlation between the study-level random effect and the regressors, not zero covariance among the regressors themselves or between Y and X. Moreover, the DGP as described contains no study-level random component; studies differ only through the location and time intercepts. This error matters because the paper uses it to justify preferring fixed-effects over random/mixed specifications and to interpret the good performance of REs. The justification and the conclusions drawn from it need to be corrected.","section":"Section 5.1, Eq. (8)"},{"comment":"Statistical power is the first criterion in the performance hierarchy, yet the Type I error adjustment is not specified. The text says only that 'the Type I error was adjusted when the sample size of the simulated meta-regression was increased by either a larger sample size or more countries.' To make the reported power comparisons reproducible, the authors must state the adjustment formula or critical values used, the alpha levels in each of the twelve cases, and how the adjustment affects comparisons of specifications across cases with different sample sizes.","section":"Section 5.2.1"},{"comment":"The results are presented only through selected figures and qualitative statements such as 'performed the best' and 'close second best'; Tables 5 and 6 summarize directions of change with arrows rather than numerical estimates. Without Monte Carlo tables reporting mean power, bias, variance, MSE, and confidence-interval coverage with simulation standard errors for all ten specifications across the twelve cases, the rankings cannot be verified or compared across conditions. I recommend adding such tables.","section":"Section 5.3 and Appendices D/E"},{"comment":"The rescaled trend term is defined inconsistently. Section 4.4 defines n as the number of time periods, so with T=5 the trend values should always be [-1,-0.5,0,0.5,1], regardless of the number of locations. Section 5.3.1 instead reports different rescaled values for n=9 and n=15 locations, which is algebraically impossible under the stated formula. Because the judgment that FEs,Trend provides biased trend estimates depends on the actual trend coding, the authors should clarify the coding and verify that the results in Figure 9 are not an artifact of the rescaling.","section":"Section 4.4 and Section 5.3.1"}],"minor_comments":[{"comment":"The nine-location small-effect vector is printed as {-2,-1.5,-1,0.5,0,0.5,1,1.5,2}; this contains 0.5 twice and omits -0.5, contradicting the text in Section 5.1. Please correct.","section":"Table 2, Cases 3 and 4"},{"comment":"The text uses n both for sample size (Table 1) and for the number of locations in the trend-rescaling discussion; this is confusing and should be made consistent (for example, using n_s, n_l, and T).","section":"Section 5.3.1"},{"comment":"The practitioners' guide recommends REs as one of three preferred specifications while also stating that random effects should not be used when the random-effects assumption does not hold; since the paper argues this assumption rarely holds, the recommendation needs a condition or clarification.","section":"Section 6"},{"comment":"The manuscript says 'Code availability: Github link' but no URL is given; a working link is necessary to verify the simulation and the trend coding.","section":"Code availability"},{"comment":"Because the number of locations and the number of studies are perfectly confounded in the design (L=5,9,15 gives 25,45,75 studies), the finding that performance improves with more locations cannot be separated from the sample-size effect; the text acknowledges this, but it should be stated as a design limitation rather than an independent result.","section":"Section 5.3.3"}],"recommendation":"major_revision","confidential_remarks":"The authors' own case-level results undercut the abstract's blanket claim, but this is fixable through a major revision that qualifies the conclusions and supplies the missing numerical and procedural details. The editor should also ensure that the code link and the Type I error adjustment are provided in the revised version, as replicability is central for a simulation study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: useful simulation comparison with a new specification, but the abstract claims more than the paper's own results show. A conditional accept is right; I would send it to referees, but it needs revision before the decision rule is usable.\n\nWhat is actually new: this is the first systematic comparison of ten meta-regression specifications under joint location and time heterogeneity, including the new FEl,Trend (location fixed effects with a trend). The design is straightforward and mostly internally consistent; the DGP, the twelve cases, and the four performance criteria are laid out clearly enough to follow. In the cases with a large time effect, the paper does show that the standard study-level models and the joint controls beat location-only and time-only controls. That result is useful even if not surprising. The authors also deserve credit for reporting the conditions where the advantage disappears.\n\nSoft spots:\n- The abstract's blanket claim that jointly modeling heterogeneity improves performance is not a summary of the full simulation. In Section 5.3.2 the authors say that with a large location effect all joint and non-joint specifications perform just as well, and in Section 5.3.3 they say that with more locations non-joint specifications perform just as well on average. The stress-test note lands. The claim should be conditioned on the time effect being large.\n- Section 5.1 states that covariance entries in [0.2, 0.5] mean the random-effects assumption of zero correlation is violated. That is wrong as written; covariances among covariates are not the random-effects assumption, which is about correlation between the random effects and regressors. The simulation may still violate it, but this sentence is not the reason.\n- The Type I error adjustment in Section 5.2.1 is not specified. Since power is the first performance criterion, that matters.\n- Results are reported qualitatively with no Monte Carlo standard errors or error bars, so small ranking differences are hard to interpret.\n- The code link is a placeholder, and Table 2 has a typo in the nine-location values.\n\nNone of these destroy the core design. The question being asked—which controls are worth adding when studies vary in both location and time—is reasonable, and the paper gives a partial, usable answer. Bottom line: worth a serious referee. I would ask for a rewritten abstract and conclusion, a specified Type I error adjustment, numerical results with Monte Carlo uncertainty, and working code. For applied researchers doing multi-country, multi-year meta-regressions, the revised version would be citable.","headline":"Useful simulation comparison with a new location-time trend specification, but the abstract overstates the results; worth peer review after revision.","tokens_in":42125,"tokens_out":4300,"would_cite":false,"duration_ms":40493,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When meta-regression data vary across both location and time, jointly modeling both dimensions outperforms modeling only one dimension.","keywords":["meta-regression","simulation study","location heterogeneity","time heterogeneity","joint heterogeneity","fixed effects","random effects","model selection"],"falsifier":"Re-run the simulation with one location following a different time trend (or with randomly missing location-year cells) and compare the joint specifications against location-only and time-only controls; if the one-dimensional controls match or beat the joint specifications on bias and power, the paper's central decision rule fails.","tokens_in":41179,"feed_emoji":"📊","tokens_out":5640,"duration_ms":51100,"temperature":0.7,"pith_summary":"This paper asks whether meta-regressions that pool studies from many countries and years should control for location and time heterogeneity explicitly, rather than only grouping by study. The authors simulate subject-level meta-regression data with location-specific intercepts and a common time trend, and compare the standard study-level random, fixed, and mixed effects models with seven specifications that add location and/or time controls. They find that when heterogeneity exists in both dimensions, jointly modeling it improves power, bias, and variance compared with modeling only one dimension; the standard methodology performs best overall, joint specifications finish a close second, and location-only or time-only specifications perform worst. The paper also introduces location-level fixed effects with a trend, which estimates the overall time trend accurately, unlike study-level fixed effects with a trend. If the finding holds outside the simulated design, practitioners should test for both dimensions and choose study-level or joint controls accordingly.","feed_headline":"Joint location-time controls beat single-dimension meta-regressions","feed_subtitle":"Ten-method simulation shows why meta-analyses spanning countries and decades need both controls.","key_machinery":"The mechanism is a Monte Carlo simulation (10,000 iterations) whose data-generating process embeds location and time heterogeneity directly in the mean of the dependent variable: each location is assigned its own $\\mu_Y$ intercept and each year adds 0.1 or 0.5 to $\\mu_Y$, producing twelve cases that cross small/large location effects with small/large time effects and 5, 9, or 15 locations. Ten meta-regression specifications are estimated by inverse-variance weighted least squares, and performance is ranked by a fixed hierarchy: power first, then estimator bias, estimator variance, and confidence-interval precision. The newly introduced specification, location-level fixed effects with a trend (FEl,Trend), adds a rescaled trend polynomial $t_0 = \\frac{2t - n - 1}{n - 1}$ to location dummies, which lets the model both control and measure a shared time trend; this is the object that carries the paper's positive recommendation.","core_discovery":"The paper's central claim is that conventional meta-regression model selection is incomplete when studies differ in both place and time: the standard practice of grouping at the study level (random, fixed, or mixed effects) performs best overall, and methods that explicitly control both location and time perform a close second, while methods that control only location or only time perform worst under joint heterogeneity. In the simulation, location heterogeneity enters as fixed additive location-specific intercepts and time heterogeneity as a common linear trend scaled by 0.1 or 0.5 per year; all ten specifications estimate the covariate coefficients accurately, so the ranking is driven by statistical power, estimator variance, and trend accuracy. Study-level fixed effects with a trend are singled out as providing biased trend estimates, time-level fixed effects alone suffer from poor power, and the newly proposed location-level fixed effects with a trend gives near-best performance and accurate trend estimates. The authors therefore recommend REs, FEs, or the new FEl,Trend depending on whether the random effects assumption holds and whether the overall time effect is of interest.","pith_inferences":["The paper only simulates joint heterogeneity in both dimensions; its results do not show that location-level fixed effects fail when only location heterogeneity is present, and the authors list that as future work.","Because the simulated time trend is identical across locations, the recommended single global trend may misstate settings where countries follow different time paths; extending the simulation to location-specific trends would test that boundary.","A natural extension is to turn Section 7's parameter extraction into a routine check: for any real meta-regression with location-year cells, fit the ten specifications and confirm the ranking on synthetic data generated from that study's own means and covariances."],"forward_implications":["Practitioners pooling studies across countries and decades should test for location and time heterogeneity and prefer study-level models or joint controls (two-way fixed effects or location fixed effects with trend) over location-only or time-only controls.","Study-level fixed effects with a trend should not be used when the trend estimate matters, because it produced biased trend estimates in all simulated cases.","Time-level fixed effects alone are a poor default under joint heterogeneity, since they lost power even when time heterogeneity was large.","The new location-level fixed effects with trend specification is a viable alternative when the goal is to estimate the overall time effect while controlling for country differences.","A decision rule emerges: if the random effects assumption fails, use fixed effects; if an overall time effect is wanted, use location fixed effects with a trend; otherwise standard study-level models are safe."],"supporting_citations":[{"why":"Supplies the study-level random and mixed effects simulation framework and the finding that high study-level heterogeneity degrades standard methods.","marker":"Bakbergenuly & Kulinskaya, 2018"},{"why":"Introduces Continuous Time Meta-Analysis and the study-level fixed-effects-with-trend approach that the simulation tests.","marker":"Dormann et al., 2020"},{"why":"The closest prior time-varying meta-analysis methodology, from which the study-level fixed effects with a trend specification is taken.","marker":"LeBlanc & Banks, 2022"},{"why":"An empirical experimental-economics meta-analysis whose location-level fixed effects approach the paper tests.","marker":"Oosterbeek et al., 2004"},{"why":"Another trust-game meta-analysis using location controls, motivating the location-level fixed and random effects specifications.","marker":"N. D. Johnson & Mislin, 2011"},{"why":"Provides the weighted least squares estimation strategy that all ten specifications share.","marker":"Stanley & Doucouliagos, 2017"},{"why":"Grounds the claim that the random effects assumption rarely holds, which determines which specifications are appropriate.","marker":"Wooldridge, 2010"}],"fun_headline_variants":["Simulation: model location AND time for better meta-regression","Both controls beat one in meta-regression simulations","Joint place-time controls win in meta-regression study","Meta-analysis needs both place and time heterogeneity controls","New method shines when studies vary in space and time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that real meta-regression data looks like the simulated data: location differences are fixed intercepts, time differences are one shared linear trend, and every location contributes one study in every year with no missing cells; if real data have unbalanced panels or location-specific trends, the recommended ranking of methods could change.","fun_headline_variants_meta":{"raw":{"variants":["Simulation: model location AND time for better meta-regression","Both controls beat one in meta-regression simulations","Joint place-time controls win in meta-regression study","Meta-analysis needs both place and time heterogeneity controls","New method shines when studies vary in space and time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1171,"prompt_tokens":872,"completion_tokens":299,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":223}},"tokens_in":488,"tokens_out":299,"duration_ms":3773,"temperature":1.0,"reasoning_tokens":223,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:57:29.379178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the simulation with one location following a different time trend (or with randomly missing location-year cells) and compare the joint specifications against location-only and time-only controls; if the one-dimensional controls match or beat the joint specifications on bias and power, the paper's central decision rule fails.","supporting_citations":[{"cited_title":"\\ Kulinskaya, E","cited_arxiv_id":null,"evidence_quote":"Supplies the study-level random and mixed effects simulation framework and the finding that high study-level heterogeneity degrades standard methods."},{"cited_title":", Guthier, C","cited_arxiv_id":null,"evidence_quote":"Introduces Continuous Time Meta-Analysis and the study-level fixed-effects-with-trend approach that the simulation tests."},{"cited_title":"Time-varying Bayesian Network Meta-Analysis","cited_arxiv_id":"2211.08312","evidence_quote":"The closest prior time-varying meta-analysis methodology, from which the study-level fixed effects with a trend specification is taken."},{"cited_title":"\\ Doucouliagos, H","cited_arxiv_id":null,"evidence_quote":"Provides the weighted least squares estimation strategy that all ten specifications share."}],"review_version":1}