{"id":"37940543-3a59-468b-be72-d4f6f4ca4996","arxiv_id":"1908.01923","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Calibrated probabilistic projections place 2018-2100 CO2 emissions most likely between 700 and 1800 GtC, with SSP5-RCP8.5 as a below-1% tail risk under baseline policies.","lead":"Using a simple economic-climate model calibrated to two centuries of data and expert judgements, this paper assigns probabilities to future CO2 emissions scenarios and finds the highest emissions scenarios are extreme tail risks. The results give risk managers quantitative probabilities for IPCC scenarios instead of treating all scenarios as equally likely.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The <1% SSP5-RCP8.5 exceedance probability is driven by an unidentifiable saturating-TFP structure; alternative TFP dynamics could materially change the tail claim.","rationale":"The reader's weakest assumption matches the concern I identify: the saturating TFP structure and its priors drive the low-growth projection that places SSP5-RCP8.5 below the 1% exceedance threshold. I agree with that identification. The paper is transparent about many caveats (no NETs, fixed resource limits, no COVID-type shocks), and the qualitative conclusion that RCP8.5 is not a central baseline is consistent with other literature. However, the specific quantitative claim 'below 1%' is a posterior predictive probability under a model that cannot represent non-saturating or trend-breaking TFP, and the historical record cannot identify A_s. Because the paper tests priors but not model structure, the central claim needs a structural robustness check before the precise probability can be accepted. This does not change the reader's CONDITIONAL verdict; it sharpens the condition. I recommend UNCHANGED.","tokens_in":21188,"tokens_out":6030,"duration_ms":79047,"concrete_test":"Recalibrate the model under two alternative TFP specifications using the same data and VAR(1) likelihood: (i) exponential TFP growth A_t = A_{t-1}(1+alpha) with the same alpha prior, and (ii) TFP with a trend-break or random-walk drift, e.g. A_t = A_{t-1} exp(nu_t) with estimated volatility. For each, recompute the posterior predictive probability that 2018-2100 cumulative CO2 emissions exceed the SSP5-RCP8.5 value (~2,100 GtC, using the same emissions definition). Also run an identifiability check for A_s: widen its prior to, say, uniform(5,50) and report whether the posterior is materially narrower than the prior. If the exceedance probability rises above roughly 5% in either alternative, or if A_s is not updated, the 'exceptionally unlikely' characterization is conditional on an untested structural assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's quantitative anchor is that SSP5-RCP8.5 cumulative emissions (~2,100 GtC) have posterior predictive exceedance probability below 1% and are 'exceptionally unlikely'. This probability is computed from a model whose TFP process is the saturating logistic law (Supplement S1: A_t = A_{t-1} + alpha A_{t-1}[1 - A_{t-1}/A_s]), with A_s assigned a uniform prior 5.3-16.11 (Table S1). Over the 1820-2014 calibration window TFP is far from saturation, so the likelihood is nearly flat in A_s; the paper itself notes that some parameters selected for alternate-prior testing were 'not updated by the Bayesian inversion' (S4). The low-growth projection (median ~1.2% per year per capita), which the paper states is the primary reason SSP5-8.5 falls outside the likely range, is therefore largely carried by the assumed saturation structure and the priors on alpha and A_s rather than by the data. The alternate-prior analysis (Table S3, Figure S1) changes prior shapes but does not test the structural form of the TFP equation. If TFP follows non-saturating or trend-breaking dynamics, which Section 4 explicitly lists as a relevant extension, the upper tail of cumulative emissions and the sub-1% exceedance claim can move materially. The qualitative ordering of SSP scenarios may survive, but the precise probability statement is conditional on an untested structural assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper calibrates a simple, DICE-like integrated assessment model with globally aggregated population, Solow-Swan economic growth, and a four-technology emissions module to historical observations (1820-2014) and two expert assessments, then produces probabilistic baseline CO2 emissions projections through 2100. Four scenarios vary fossil fuel resource caps (3,000/6,000/10,000 GtC) and prior beliefs about the zero-carbon technology half-saturation year. The central findings are that medium-to-high SSP scenarios (SSP3-7.0, SSP4-6.0) fall within the central 90% prediction intervals, that SSP5-RCP8.5 cumulative emissions (~2,100 GtC) have a posterior exceedance probability below 1% and should be interpreted as a tail-risk scenario, that likely cumulative emissions from 2018-2100 are 700-1,800 GtC, and that economic and technology dynamics dominate sensitivity of cumulative emissions, with population dynamics less important.","tokens_in":21486,"tokens_out":5522,"duration_ms":56443,"significance":"If the quantitative claims hold, the paper makes a valuable, decision-relevant contribution by providing probabilistic baselines for 21st-century CO2 emissions and by explicitly situating SSP scenarios within those distributions. The statistical treatment is careful: a VAR(1) residual structure, explicit likelihood derivation, four-chain MCMC with Gelman-Rubin diagnostics, k-fold cross-validation with 93% coverage, and variance-based Sobol sensitivity analysis with bootstrap confidence intervals are all strengths. The paper's simple, transparent model structure and explicit scenario analysis make it a useful reference point for risk assessments. However, the central quantitative claim--that SSP5-RCP8.5 has a below-1% exceedance probability--is conditional on a structural assumption about saturating total factor productivity that is weakly identified by historical data, and the abstract overstates support regarding the low end of the scenario distribution. The paper is therefore significant but requires additional structural robustness analysis before its headline quantitative claim can be accepted as stated.","major_comments":[{"comment":"The below-1% exceedance probability for SSP5-RCP8.5 cumulative emissions is driven primarily by the model's relatively low economic growth projections, which in turn are driven by the saturating total factor productivity (TFP) equation A_t = A_{t-1} + alpha*A_{t-1}*(1 - A_{t-1}/A_s) (Section S1). The paper itself notes in Section S4 that some parameters used in the alternate-prior testing were 'not updated by the Bayesian inversion,' and Table S1 shows wide priors on alpha and A_s. Over the 1820-2014 calibration window the system is far from TFP saturation, so the data are nearly uninformative about A_s, leaving the low-growth projection (median roughly 1.2% per year) heavily dependent on the assumed saturating functional form. The alternate-prior analysis in Table S3 and Figure S1 changes prior shapes but does not test structurally different TFP dynamics, even though Section 4 explicitly lists trend breaks in TFP growth as a relevant extension. A non-saturating or trend-breaking TFP process could materially shift the upper tail of cumulative emissions and hence the stated <1% exceedance probability. The qualitative ordering of SSP scenarios may survive, but the precise quantitative tail-risk claim is conditional on an untested structural assumption. Please demonstrate robustness under structurally alternative TFP specifications or revise the probability claims to reflect this conditioning.","section":"Section 3 and Supplemental Sections S1, S4"},{"comment":"The abstract claims that 'more moderate scenarios used by the Intergovernmental Panel on Climate Change are more likely than the extreme high or low scenarios,' but Section 2 explicitly states that the model cannot compare with SSP1-1.9, SSP2-2.6, and SSP4-3.4 because those scenarios include negative emissions technologies, which are not represented in the model. The analysis therefore provides no quantitative support for statements about the low end of the scenario distribution; it can only support comparisons among the high and medium scenarios that are actually evaluated. The abstract should be revised to align with the scenarios actually analyzed.","section":"Abstract and Section 2"},{"comment":"The text states that SSP5-RCP8.5 'remains exceptionally unlikely ... with an exceedance probability below 1%,' but the numerical exceedance probability is not reported anywhere in the main text or supplement. Figure 2 shows cumulative density functions, but precise values cannot be read from the figure, especially for tails below 1%. Please report the estimated exceedance probabilities (and, if possible, posterior credible intervals for those probabilities) for all four model scenarios in a table so that the central quantitative claim is directly verifiable.","section":"Section 3, Figure 2"}],"minor_comments":[{"comment":"The sentence 'The average cross-validation coverage of the 90% credible intervals for the held-out data are 93' appears truncated; it should read 'are 93%.' Please also report the standard error or range across the fifty hold-out sets.","section":"Section S2"},{"comment":"The phrase 'The marker baseline SSP-RCP emissions scenarios which will be used...' likely should be 'The marked baseline SSP-RCP emissions scenarios...' or 'The markers show baseline SSP-RCP emissions scenarios...' for clarity.","section":"Figure 1 caption"},{"comment":"The sentence 'Depending on the scenario, however, the 90% credible interval of our projections can include anywhere from 78-85% of the SSP5-RCP8.5 cumulative emissions depending on the scenario' is ambiguous: it is unclear whether the 78-85% refers to the share of the scenario's cumulative emissions that falls inside the interval or the probability mass of the interval overlapping the scenario value. Please clarify.","section":"Section 3"},{"comment":"The prior table lists separate normal priors for rho2 and rho3 with the same bounds, but the text in Section S1 states the constraint rho2 >= rho3 is imposed. The table and text should make clear how the constraint is incorporated in the prior specification.","section":"Table S1"},{"comment":"The observation error variance for emissions is labeled 'epsilon1' in the table, duplicating the label for population; it should be labeled 'epsilon3' (or similar) to be consistent with the three-module notation.","section":"Table S2"},{"comment":"No data or code availability statement is provided in the manuscript. Given the paper's emphasis on transparency and reproducibility, a statement indicating where the data and code (if available) can be accessed would strengthen the contribution.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper uses an expert assessment (Ho et al. 2019) that is co-authored by one of the co-authors of this manuscript. The supplemental analysis (Figure S6) examines the impact of excluding this assessment, so this is not a serious concern, but the editor may wish to note the self-citation. The central issue is whether the <1% exceedance claim can be defended given the weakly identified saturating-TFP structure; this is a substantive scientific question rather than a case of internal inconsistency, so I recommend major revision rather than rejection. The paper is within the scope of the journal and would be a useful contribution if the structural robustness concern is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a careful Bayesian calibration of a simple DICE-like IAM that assigns probabilities to SSP-RCP emissions scenarios, and it concludes that SSP5-RCP8.5 is a tail risk with less than 1% exceedance probability for cumulative emissions 2018–2100. Anyone who uses scenarios in climate risk assessment should engage with it. It does something genuinely new: previous critiques of RCP8.5 were qualitative or based on scenario design; here you get calibrated probabilities and a variance decomposition.\n\nWhat the paper does well: the statistical machinery is solid. VAR(1) residuals are justified by autocorrelation (lag-1 0.94 for emissions, etc.). Cross-validation coverage of 93% is reported. Four-chain MCMC with Gelman-Rubin, explicit prior sensitivity checks. The Sobol analysis with bootstrap confidence intervals is thoughtful. The authors are honest about what the model cannot do: no negative emissions, no policy changes, fossil fuel resource limits varied across scenarios. The main qualitative finding—RCP8.5 is not a plausible baseline; SSP3-7.0 and SSP4-6.0 fit better—is robust across their four scenario designs.\n\nNow the soft spot, and it is the one you should care about. The sub-1% exceedance claim is driven by low economic growth projections: median about 1.2% per year per capita. That growth distribution comes largely from a saturating-TFP equation (Supplement S1, A_t = A_{t-1} + alpha A_{t-1}[1 - A_{t-1}/A_s]) with a uniform prior on the saturation level A_s. Over the calibration window, the economy is far from saturation, so the data can barely identify A_s. The paper's own sensitivity check says some parameters were \"not updated by the Bayesian inversion\" (S4). So the low-growth tail is carried by the structural form and priors, not by historical data. The authors test alternative priors but not alternative structural forms. They list trend breaks in TFP as a relevant extension (Section 4), which is good, but that means the quantitative tail claim is conditional on an untested assumption. I think the qualitative ordering would survive alternative TFP dynamics—most plausible worlds are still not RCP8.5—but the \"below 1%\" number could move materially. The paper should state that more explicitly, or better, show how the exceedance probability changes under a non-saturating TFP process.\n\nSecond soft spot: no code or processed data are provided. For a paper that makes precise probability statements, that is a reproducibility gap. The MCMC details are enough to reassure, but not enough to independently verify the 1% claim.\n\nWho this is for: people working on climate risk and scenario interpretation, especially those who need to justify scenario weights in impact assessments. It is also a nice example of Bayesian calibration of a simple IAM for teaching. My recommendation: accept for peer review, with a request for code/data and a sensitivity analysis on TFP structure. The central claim is important and mostly supportable, but the quantitative anchor needs to be firmed up.","headline":"Careful Bayesian calibration that makes a strong, conditional case that SSP5-RCP8.5 is a tail risk; the quantitative anchor is weakened by an untested saturating-TFP structure.","tokens_in":22087,"tokens_out":2935,"would_cite":true,"duration_ms":27769,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that SSP5-RCP8.5 is a below-1% tail risk, not a baseline.","keywords":["CO2 emissions projections","integrated assessment model","Shared Socioeconomic Pathways","RCP8.5","Bayesian calibration","uncertainty quantification","economic growth","climate risk assessment"],"falsifier":"Track realized global per-capita gross world product growth and annual CO2 emissions over 2021–2040 and compare them with the model's 90% credible intervals. If average growth consistently exceeds the projected upper bound (roughly 2% per year) or if cumulative emissions drift above the interval's upper edge, the calibrated TFP-saturation mechanism is contradicted, and the below-1% tail probability for SSP5-RCP8.5 would not hold.","tokens_in":2008,"feed_emoji":"📉","tokens_out":4773,"duration_ms":117768,"temperature":0.7,"pith_summary":"This paper uses a deliberately simple, transparent model of population, economic output, and carbon emissions, calibrated to centuries of historical data and expert judgments, to attach probabilities to baseline CO2 emissions through 2100. Its central claim is that the highest emissions scenarios used by the IPCC, especially SSP5-RCP8.5, are extreme tail risks rather than plausible business-as-usual futures: the cumulative emissions of SSP5-RCP8.5 have an exceedance probability below 1%. Under current policies and without negative-emissions technologies, more moderate scenarios such as SSP3-7.0 and SSP4-6.0 fall inside the model's likely range, while the projected likely range for cumulative 2018–2100 emissions is 700 to 1800 GtC. This matters because treating every scenario as equally likely biases climate risk assessment toward futures the historical record and expert opinion suggest are very unlikely, and it implies far more aggressive mitigation is needed for the 2°C target.","feed_headline":"The hottest CO2 scenario is a below-1% tail risk","feed_subtitle":"A calibrated model puts likely 2018–2100 CO2 emissions at 700–1800 GtC, with moderate scenarios most probable.","key_machinery":"The machinery is a simple, globally aggregated integrated assessment model that couples logistic population growth, a Solow–Swan/Cobb–Douglas production block, and emissions from four competing technologies with logistic penetration curves: a zero-carbon pre-industrial source, a coal-like high-carbon source, an oil-and-gas-like lower-carbon source, and an advanced zero-carbon source. Two structural assumptions carry the conclusions: total factor productivity grows logistically toward a saturation level, which keeps long-run growth near 1.2% per year, and a hard cap on cumulative fossil-fuel emissions (6,000 GtC in the standard case, with 3,000 and 10,000 GtC variants) forces eventual substitution to zero-carbon technology. The model is calibrated by Markov chain Monte Carlo with a vector-autoregressive error structure, and the prior distributions incorporate two expert assessments of long-run growth and 2100 emissions. The same machinery yields both the probabilistic emissions projections and a Sobol' variance decomposition showing that interactions among productivity growth, the labor elasticity, and the carbon intensity of the lower-carbon fossil technology dominate emissions uncertainty.","core_discovery":"The study's central discovery is a probabilistic ranking of SSP-RCP emissions scenarios conditional on current policies and no negative-emissions technologies. In the standard calibration the central 90% interval for annual CO2 emissions in 2100 is 6–28 GtC, and the 34 GtC emissions of SSP5-RCP8.5 sit above that upper limit; its cumulative emissions of about 2,100 GtC are exceeded by fewer than 1% of posterior simulations. The median cumulative projection is roughly 1,200 GtC, and the likely range is 700–1,800 GtC across the standard and high-fossil-fuel cases. The driver of SSP5-RCP8.5's tail status is primarily the model's modest median economic growth of about 1.2% per year, lower than the inverted expert-growth assessment and SSP5's own assumptions. The result is robust to varying fossil-fuel resource caps and decarbonization priors, though delayed zero-carbon penetration thickens the upper tail.","pith_inferences":["If autonomous productivity growth turns out not to saturate, for example because automation keeps shifting the technological frontier, the model's low-growth prior may be too pessimistic, and the probability attached to SSP5-RCP8.5 could rise well above 1%; this is a direct, testable implication of the structural assumption.","Because the model disallows coal from regaining energy share, it may understate how quickly emissions could rise if coal became cheap again; adding an explicit coal-return mechanism would be a robustness test of the tail-risk claim.","The same calibration framework could be rerun as new policies are implemented, so the baseline and the tail status of high scenarios should be treated as an evolving standard rather than a fixed result."],"forward_implications":["If the projections are right, intermediate-high scenarios such as SSP3-7.0 and SSP4-6.0 are better baselines than SSP5-RCP8.5 for climate risk assessment.","Achieving even a 50% chance of meeting the 2°C target is very unlikely under baseline emissions, because the remaining carbon budgets for 1.5°C and 2°C are below the 1st percentile of projected cumulative emissions.","Fossil-fuel resource uncertainty shifts the projections modestly; even the high-resource case does not bring SSP5-RCP8.5 inside the likely range.","Population uncertainty contributes little to cumulative-emissions variability, while economic and technology uncertainties dominate, with strong interactions among them.","More aggressive mitigation than current pledges is required to reliably achieve the 2°C Paris Agreement target."],"supporting_citations":[{"why":"Supplies the SSP emissions scenarios, including SSP5-RCP8.5, that the projections are judged against.","marker":"Riahi et al. (2017)"},{"why":"Establishes the SSP-RCP scenario matrix used in climate assessments and provides the scenario framework for the comparisons.","marker":"O'Neill et al. (2016)"},{"why":"Supplies the DICE model structure and the standard 6,000 GtC fossil-fuel resource assumption adopted for the base case.","marker":"Nordhaus and Sztorc (2013)"},{"why":"Supplies the expert assessment of long-run economic growth used to inform the prior on productivity and growth parameters.","marker":"Christensen et al. (2018)"},{"why":"Supplies the expert assessment of 2100 baseline emissions used as additional information in calibration.","marker":"Ho et al. (2019)"},{"why":"Provides the coal-expansion critique that motivates treating RCP8.5/SSP5 as unlikely rather than a baseline.","marker":"Ritchie and Dowlatabadi (2017a)"},{"why":"Documents how 'business as usual' labels mislead and frames the interpretation of high-end scenarios as extremes.","marker":"Hausfather and Peters (2020)"},{"why":"Provides the 3,000 GtC low fossil-fuel resource scenario used as a deep-uncertainty case.","marker":"McGlade and Ekins (2015)"},{"why":"Provides the 10,000 GtC high fossil-fuel resource scenario used as a deep-uncertainty case.","marker":"Bruckner et al. (2014)"},{"why":"Provides the remaining carbon budgets for 1.5°C and 2°C targets used to show that baseline emissions exceed the budgets.","marker":"Rogelj et al. (2018b)"}],"fun_headline_variants":["SSP5-RCP8.5: a <1% tail risk, not the baseline","CO2 projections relegate RCP8.5 to tail risk","Probabilistic model: high-emission path has <1% odds","RCP8.5 often misused as baseline; model says <1% likely","Business-as-usual emissions: RCP8.5 is a rare tail"],"cache_read_input_tokens":24064,"weakest_assumption_plain":"The load-bearing premise is that long-run innovation slows as total factor productivity approaches a saturation ceiling, so median economic growth stays near 1.2% per year; if innovation does not saturate, the probability assigned to very high emissions could be materially larger.","fun_headline_variants_meta":{"raw":{"variants":["SSP5-RCP8.5: a <1% tail risk, not the baseline","CO2 projections relegate RCP8.5 to tail risk","Probabilistic model: high-emission path has <1% odds","RCP8.5 often misused as baseline; model says <1% likely","Business-as-usual emissions: RCP8.5 is a rare tail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001321,"raw_usage":{"total_tokens":5392,"prompt_tokens":970,"completion_tokens":4422,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":4318}},"tokens_in":586,"tokens_out":4422,"duration_ms":32855,"temperature":1.0,"reasoning_tokens":4318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:00:54.778358+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Track realized global per-capita gross world product growth and annual CO2 emissions over 2021–2040 and compare them with the model's 90% credible intervals. If average growth consistently exceeds the projected upper bound (roughly 2% per year) or if cumulative emissions drift above the interval's upper edge, the calibrated TFP-saturation mechanism is contradicted, and the below-1% tail probability for SSP5-RCP8.5 would not hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DICE model structure and the standard 6,000 GtC fossil-fuel resource assumption adopted for the base case."},{"cited_title":"Proc Natl Acad Sci U S A 115(21):5409--5414, doi:10.1073/pnas.1713628115","cited_arxiv_id":null,"evidence_quote":"Supplies the expert assessment of long-run economic growth used to inform the prior on productivity and growth parameters."},{"cited_title":"Nature 577:618--620","cited_arxiv_id":null,"evidence_quote":"Documents how 'business as usual' labels mislead and frames the interpretation of high-end scenarios as extremes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 10,000 GtC high fossil-fuel resource scenario used as a deep-uncertainty case."}],"review_version":1}