{"id":"afdecf58-eb1d-4cf5-8916-0dda25c20787","arxiv_id":"2411.13783","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"When four open-source U.S. power system models are given identical data and settings, they produce nearly the same system costs and similar capacity plans, while configuration choices such as economic retirement drive the largest differences.","lead":"This study compared four open-source U.S. electricity capacity expansion models with identical inputs and settings, finding that they agree closely on costs and overall investments. It also shows that choices about retirement rules, operational detail, and planning horizon matter more for results than which model is used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.2–0.3% cost-convergence claim is measured after iterative model revisions whose stopping rule is the same agreement, so it is not an independent test of convergence.","rationale":"Reader's conditional verdict is appropriate. The paper is valuable and honest: it openly reports the iterative harmonization protocol and the Appendix C error, and the common-evaluator design is a genuine strength. But the headline cost-convergence claim is not yet independently supported, because the iterative protocol's stopping rule and the single GenX evaluator are entangled with the measurement. This does not mean the models are wrong; it means the paper should either report a pre-harmonization baseline or test convergence with an independent evaluator. Given the repository is available, the proposed check is feasible. The configuration-effect findings are less exposed to this concern because they compare within-model differences, so the overall conditional-accept recommendation stands unchanged.","tokens_in":16470,"tokens_out":3803,"duration_ms":35188,"concrete_test":"Re-run the four models from their pre-harmonization state—after only the prospective PowerGenome input alignment, skipping the iterative 'revising models' stages 2 and 3—for the net-zero base configuration, using the code in the project repository. Score the resulting portfolios with the same GenX operational simulation and compute the across-model spread in discounted system cost. If the pre-harmonization spread is still ≤0.3%, the convergence claim is robust; if it is materially larger, the headline agreement is an artifact of the harmonization loop rather than an independent finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that harmonized models find 'nearly equivalent global cost minimums'—rests on the 0.2–0.3% NPV spread reported in Section 3.1. But that spread is generated after the harmonization protocol described in Section 2: models were revised in stages 2 and 3 by comparing results and iterating until residual differences were judged 'not large.' The convergence figure is therefore an output of the stopping rule, not an independent measurement of model proximity. Moreover, the common cost metric is the GenX operational simulation (Section 2, 'Calculating model costs'), which scores every proposed portfolio under GenX's dispatch and cycling assumptions. A small spread under that single evaluator does not establish that the models' native optimization problems share a global minimum, nor that residual capacity differences are 'quasi-random variation.' It is compatible with a different possibility: the iterative revisions suppressed structural diversity, or GenX's operational model compresses cost differences among divergent portfolios. Appendix C's admitted error in the 2027 CO2 target is a separate correctness issue but does not bear directly on the convergence claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a structured intercomparison of four open-source capacity expansion models (GenX, Switch, TEMOA, USENSYS) for the continental US power system, using PowerGenome to harmonize inputs and explicit scenarios/configurations. The authors claim that after iterative harmonization, the models produce very similar capacity portfolios and system costs (0.2–0.3% NPV spread in the base case, <1% for most configurations), interpret residual differences as quasi-random variation, and then use the harmonized models to study policy-relevant scenario and configuration effects (carbon buyout prices, transmission constraints, CCS availability, retirement rules, temporal sampling, foresight). The paper also reflects on lessons for future intercomparison efforts.","tokens_in":16718,"tokens_out":3817,"duration_ms":37708,"significance":"If the convergence claim is valid, the paper provides a strong practical result: harmonized open-source capacity expansion models can agree on system costs to within a fraction of a percent, making configuration effects the dominant source of divergence. The study is valuable as a community resource: it is transparent about the harmonization protocol, uses a common operational simulation as a cost metric, makes results and data available via GitHub, and candidly documents an input data error in Appendix C. The configuration-effect findings (e.g., economic retirement and unit commitment are influential) are useful and less affected by the circularity concern. The main risk is that the headline convergence result is partly an artifact of the iterative harmonization stopping rule, which weakens the paper's central claim as currently framed.","major_comments":[{"comment":"The 0.2–0.3% cost spread in Section 3.1 is measured after the iterative harmonization loop described in Section 2, whose stopping rule is that residual differences in energy mix, infrastructure builds, and total costs 'did not show large, unexplainable differences across models.' The convergence claim is therefore partly an output of the protocol: models were revised until agreement was judged adequate, and then agreement was reported as evidence that the models find 'nearly equivalent global cost minimums.' To make the claim load-bearing, the paper should either report results from the un-revised models, quantify how much each revision changed system costs and portfolios, or otherwise demonstrate that the stopping rule did not suppress structural diversity.","section":"Section 2 (Harmonizing models for the net-zero scenario and base configuration) and Section 3.1"},{"comment":"The NPV spread is computed by evaluating each model's proposed portfolio in a single operational model (GenX), not by comparing each model's native objective values. This is a sensible common metric, but it supports a narrower claim: the portfolios are nearly cost-equivalent under GenX's dispatch, cycling, and penalty assumptions. It does not establish that the models' native optimization problems share a global minimum, and residual capacity differences could reflect the evaluator compressing cost differences among divergent portfolios. I recommend softening 'global cost minimums' to 'cost-equivalent under a common operational metric' and, if feasible, adding a robustness check with a second evaluator or reporting native objective values.","section":"Section 2 ('Calculating model costs') and Section 3.1"},{"comment":"The 2027 CO2 target used in all net-zero runs is admitted to be the 2025 REPEAT value (847 Mt) rather than the true 2027 value (494 Mt). Because the net-zero scenario's near-term trajectory is central to Section 3.2.1 and Figure 2, and because the trajectory's steepness affects investment timing in Figures 1, 4, and 5, this error is not merely cosmetic. The authors should correct the target and rerun the net-zero scenario (or clearly demonstrate that all qualitative conclusions are unchanged under the corrected 2027 cap), and should mark the corrected table in Appendix C.","section":"Appendix C"}],"minor_comments":[{"comment":"The abstract states 'less than 1% difference in system costs for most configurations,' but Section 3.1 reports 0.2–0.3% for the base configuration only; please specify the exact set of configurations and which cost metric is used for the abstract's claim.","section":"Abstract and Section 3.1"},{"comment":"The legend uses 'Age' and 'Economic' in a way that may be confused with the retirement-type markers; consider a clearer legend that separates retirement type, temporal sampling, and period type.","section":"Figure 6 and Section 3.3"},{"comment":"Table C1 labels the third column 'REPEAT' but the 2027 row says '847 (in 2025)'; once corrected, add a footnote explaining the correction so readers do not use the old trajectory.","section":"Appendix C, Table C1"},{"comment":"The 4.6% cost increase and 55% emissions increase for no transmission expansion are stated consistently, but Figure 3's panel labels could clarify whether 'costs' are annual operational costs or total system costs.","section":"Section 3.2.2 and Discussion item (iv)"},{"comment":"The statement that 'most or all of the models were extended and adapted as needed' is relevant to interpretation of configuration effects; please state explicitly in Table 5 which models could not implement which configurations.","section":"Section 2 (Harmonizing models for additional scenarios and configurations) and Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful and unusually candid intercomparison, and the authors' willingness to document the iterative harmonization process is a strength. My main concern for the editor is that the headline convergence claim may be over-interpreted by readers if the circularity in the harmonization loop is not addressed or at least explicitly bounded. A versioned record of the model versions before and after each harmonization stage would substantially increase the credibility of the central claim. The Appendix C target error should be corrected before publication, not merely footnoted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading. It does something genuinely useful: four open-source capacity expansion models, fed identical PowerGenome inputs and run under matched configurations, and the result that configuration choices (retirement rules, unit commitment, foresight, temporal sampling) often matter more than which model you use. The shared data pipeline, the explicit scenario-versus-configuration framing, and the common GenX operational simulator used to score all portfolios on the same cost metric are real contributions. The configuration findings are the strongest part, especially the economic-retirement results and the 20-week foresight trade-off. The paper is also admirably honest—Appendix C's admission of the erroneous 2027 CO2 target is the kind of transparency you want in an intercomparison.\n\nThe soft spots are real but not fatal. The headline claim that harmonized models find 'nearly equivalent global cost minimums' rests on a 0.2-0.3% NPV spread measured after an iterative harmonization process whose stopping rule was that residual differences were 'not large.' That makes the convergence partly an artifact of the protocol, not an independent measurement. The paper should present the 0.2-0.3% as evidence of successful harmonization, not as a property of the models under natural usage. Relatedly, the common evaluator is GenX itself, so you are scoring all portfolios under one model's dispatch and cycling assumptions; that can compress cost differences among structurally different portfolios. The claim that residual capacity differences are 'quasi-random variation' is asserted more than tested—a formal similarity metric or a set of near-optimal alternatives would stiffen it.\n\nThe Appendix C error is a genuine correctness issue, though it does not change the qualitative conclusions. The 2027 target is off by nearly half, and the authors should correct it or at least re-run the affected scenarios.\n\nWho is this for? Energy systems modelers, people building or using capacity expansion models, and anyone interpreting policy studies from these tools. It will not resolve the deeper question of structural model uncertainty, but it gives a practical template and a clear warning about configuration sensitivity.\n\nMy recommendation: send it to peer review. It is a solid empirical contribution with honest reporting and reproducible infrastructure. The convergence claim needs reframing and the capacity-similarity analysis needs sharpening, but this deserves referee time and likely publication after moderate revision.","headline":"Genuinely useful intercomparison: the configuration findings are robust and practical, but the 0.2-0.3% cost convergence is partly a product of the harmonization loop and needs honest reframing.","tokens_in":17226,"tokens_out":1538,"would_cite":true,"duration_ms":54709,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Four open-source capacity expansion models, given identical inputs and configurations, converge on nearly the same system costs—within 0.2–0.3%—so residual portfolio differences are near-random substitutions, not model bias.","keywords":["capacity expansion models","model intercomparison","harmonization","open-source energy models","electricity system planning","decarbonization scenarios","unit commitment","economic retirement"],"falsifier":"Take the four models out of the harmonization loop: run each with its own native default inputs and configurations on the same net-zero scenario, and compare system costs using a single common operational simulation. If costs diverge by much more than the reported 0.2–0.3% or capacity choices line up with model identity, the convergence is an artifact of the iterative revision process rather than a property of the models. A complementary check is to apply the same harmonized protocol to a new scenario that was not part of the alignment process (for example, high electrification demand or strict zero-emission compliance with no buyout) and see whether agreement stays below 1%.","tokens_in":16302,"feed_emoji":"⚡","tokens_out":10104,"duration_ms":87778,"temperature":0.7,"pith_summary":"This paper asks whether four independent open-source electricity capacity expansion models—computer programs that choose the lowest-cost fleet of power plants and transmission lines for the future US grid—can be made to agree. The authors gave Temoa, Switch, GenX, and USENSYS identical input data through a shared data pipeline and identical scenario and configuration definitions, then compared outputs. They find that harmonized models produce nearly equal total system costs—within 0.2–0.3% in the headline current-policy and net-zero scenarios—and broadly similar capacity portfolios, with the remaining differences behaving like quasi-random rearrangements among near-equally-good plans. The paper then shows that deliberate configuration choices, such as letting old plants retire for economic reasons or adding unit-commitment constraints, affect results more than the choice of model does. For policymakers the upshot is that consensus cost and portfolio findings from harmonized open models can support reliable decisions, while the details of any single plan should be treated with caution.","feed_headline":"Four open electricity models agree on US grid costs within 0.3%","feed_subtitle":"Given identical inputs, the four open models converge; configuration choices matter more than model choice.","key_machinery":"The scenario–configuration distinction is the load-bearing separation: scenarios are the policy worlds being optimized (current policies versus net-zero trajectories with carbon buyout prices, transmission limits, and CCS availability), while configurations are the model setup choices (age-based vs economic retirement, myopic vs foresight, 52 vs 20 sampled weather weeks, unit-commitment on/off). The harmonization protocol itself is the central mechanism: a shared data pipeline supplies identical costs, fuel prices, loads, and resource potentials to all four models, and an iterative consensus process aligns structural assumptions to a common base case. Finally, a single operational simulation—a full-year dispatch run with fixed capacities and unit-commitment constraints—provides a common cost yardstick, so the cost of any model's portfolio is evaluated in the same way. Together these pieces let the paper attribute differences to configuration rather than data or solver noise.","core_discovery":"The central claim is that when four structurally different open-source capacity expansion models are forced onto a common footing—same data, same policy scenarios, same configuration choices, same solver settings—they converge on nearly the same objective value: net present system costs differ by only 0.2–0.3% across models for the current-policy and net-zero scenarios, and by under 1% for most configurations. The authors interpret this as evidence that all four models are finding essentially equivalent global cost minima, so the visible differences in capacity and generation choices are normal substitutions among plans with similar costs, not persistent biases of any one model. They then show that configuration features—most strongly the choice between age-based and economic retirement of existing plants, and secondarily unit-commitment constraints and the number of sampled weather weeks—move system costs and investment patterns more than the identity of the model does. The paper presents this as a demonstration that model intercomparison with strictly harmonized inputs can isolate structural differences, provide reliable policy insight, and identify which model features matter most for a given scenario.","pith_inferences":["The paper's own appendix reports an erroneous 2027 emissions target (847 million tonnes) inherited from the reference trajectory, likely a typo for 494 million; because all four models shared the same target, it does not by itself upset the convergence result, but it shows how a single input error can propagate uniformly through harmonized models.","If the convergence holds beyond the tested scenarios, practical planning insight will come more from improving data quality and feature fidelity (retirement rules, temporal sampling, unit commitment) than from building new models, since model identity is not the main source of variation.","A testable extension would feed the same harmonized inputs to a structurally different model family, such as an equilibrium-based or production-cost model, to see whether the cost-minimum agreement persists; that would separate the influence of shared data from the influence of shared optimization logic.","The study's result that a 20-week foresight model does not beat a 52-week myopic model suggests a caution for common practice; a systematic sweep of sample-week counts within a single model could identify where temporal detail outweighs intertemporal foresight."],"forward_implications":["If harmonized open models converge, consensus findings from multi-model studies can guide policy even when the models differ in internal implementation.","Configuration choices—especially economic retirement and unit-commitment constraints—are more influential than model choice, so model users should document and align these choices carefully.","Under current policies, the models consistently find that emissions do not fall fast enough; only a carbon buyout price of $1,000/tonne yields steady reductions in the net-zero scenario.","Restricting all inter-regional transmission expansion raises 2050 costs by roughly 4.6% and emissions by 55%, showing that transmission is valuable but can be substituted with local renewable build-out plus batteries.","The 52-week myopic base configuration often beats the 20-week foresight configuration on cost, suggesting that temporal detail can matter more than inter-temporal foresight in these models."],"supporting_citations":[{"why":"Supplies the common data pipeline that harmonizes technology costs, fuel prices, loads, and resource potentials across all four models.","marker":"[14]"},{"why":"One of the four compared models; its native optimization formulation is an object of the intercomparison.","marker":"[10]"},{"why":"One of the four compared models; contributes its handling of storage, technology vintages, and foresight to the comparison.","marker":"[11]"},{"why":"One of the four compared models; also provides the operational simulation used as the common cost yardstick.","marker":"[12]"},{"why":"One of the four compared models; supplies its modular energy-system representation and retirement logic.","marker":"[13]"},{"why":"Provides the net-zero CO2 emissions trajectory that defines the study's decarbonization scenario targets.","marker":"[20]"},{"why":"Earlier intercomparison showing untested structural differences across models; this study's configuration design directly addresses that gap.","marker":"[5]"},{"why":"Showed structural features such as storage and dispatch treatment cause discrepancies even under harmonized scenarios, motivating the present protocol.","marker":"[6]"},{"why":"Established the intermodel comparison methodology this study adapts for strictly harmonized open-source models.","marker":"[4]"}],"fun_headline_variants":["Four open grid models agree within 0.3% on US costs","Grid model settings, not models, drive US cost differences","Harmonized inputs align open grid model costs to under 0.5%","Open grid models: configuration matters more than choice","Four grid models, one answer within 0.3%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The models were revised repeatedly during the harmonization process until residual differences looked small, so the headline cost convergence is partly a product of the harmonization protocol rather than an independent measurement of how close the models would naturally be.","fun_headline_variants_meta":{"raw":{"variants":["Four open grid models agree within 0.3% on US costs","Grid model settings, not models, drive US cost differences","Harmonized inputs align open grid model costs to under 0.5%","Open grid models: configuration matters more than choice","Four grid models, one answer within 0.3%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000829,"raw_usage":{"total_tokens":3635,"prompt_tokens":971,"completion_tokens":2664,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2576}},"tokens_in":587,"tokens_out":2664,"duration_ms":17116,"temperature":1.0,"reasoning_tokens":2576,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:53:04.100321+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the four models out of the harmonization loop: run each with its own native default inputs and configurations on the same net-zero scenario, and compare system costs using a single common operational simulation. If costs diverge by much more than the reported 0.2–0.3% or capacity choices line up with model identity, the convergence is an artifact of the iterative revision process rather than a property of the models. A complementary check is to apply the same harmonized protocol to a new scenario that was not part of the alignment process (for example, high electrification demand or strict zero-emission compliance with no buyout) and see whether agreement stays below 1%.","supporting_citations":[{"cited_title":"org/records/11194213","cited_arxiv_id":null,"evidence_quote":"Supplies the common data pipeline that harmonizes technology costs, fuel prices, loads, and resource potentials across all four models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the four compared models; its native optimization formulation is an object of the intercomparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the four compared models; contributes its handling of storage, technology vintages, and foresight to the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the four compared models; also provides the operational simulation used as the common cost yardstick."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the four compared models; supplies its modular energy-system representation and retirement logic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the net-zero CO2 emissions trajectory that defines the study's decarbonization scenario targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier intercomparison showing untested structural differences across models; this study's configuration design directly addresses that gap."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Showed structural features such as storage and dispatch treatment cause discrepancies even under harmonized scenarios, motivating the present protocol."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established the intermodel comparison methodology this study adapts for strictly harmonized open-source models."}],"review_version":1}