{"id":"ea812101-4a2a-4eb8-91e5-17f8d2e0c1ef","arxiv_id":"2606.31469","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Firm-level regression on 200 European firms shows flagship index membership predicts wider disclosure-performance gaps, with detected greenwashing conditional on the ESG rating used.","lead":"This paper measures the gap between what large European firms voluntarily disclose about their environmental efforts and their actual emissions performance. It finds that membership in flagship stock indices predicts larger gaps, but the result disappears when switching ESG rating providers.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"DPG may confound scale effects with greenwashing via unnormalized emissions","rationale":"The reader's weakest assumption correctly flags DPG validity. The size-normalization issue is a concrete, testable instance of that assumption that directly threatens the index-membership coefficient; it is not addressed in the reported robustness steps. This keeps the verdict CONDITIONAL at low confidence rather than strengthening it.","tokens_in":1870,"tokens_out":312,"duration_ms":50148,"concrete_test":"Recompute DPG after replacing raw emissions with emissions intensity (Scope 1+2 / revenue) or after adding log(total assets) to the candidate pool before the stepwise/AIC stage; if the flagship β falls below |0.4| or loses significance at p<0.05 while renewable-energy coefficient is stable, the original result is sensitive to scale normalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The flagship-index result interprets wider DPG as ceremonial conformity, but DPG is the standardized divergence between CDP disclosure score and realised emissions performance. Emissions are an absolute quantity; larger firms (more likely to be flagship-index members within STOXX 600) will mechanically show worse performance unless emissions are normalized by revenue, assets, or output. The six-stage specification search and reported robustness checks do not indicate that log(assets) or emissions intensity was tested as a control or alternative construction. If the index coefficient is driven by this scale correlation rather than disclosure-performance mismatch, the greenwashing interpretation does not follow.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The paper investigates the Aggregate Confusion hypothesis at the firm level by constructing a Disclosure-Performance Gap (DPG) as the standardized divergence between voluntary environmental disclosure (CDP Climate Score) and realized emissions performance for 200 large European firms from Energy, Materials, Industrials, and Utilities sectors of the STOXX Europe 600 in 2023. Using a six-stage model selection process (candidate assembly, correlation screening, VIF filtering, stepwise forward AIC search, Cook's distance, HC3 re-estimation) on 421 specifications estimated by OLS with HC3 robust errors, it reports flagship index membership as the strongest predictor of wider DPG (β = +0.78, p < 0.01), consistent with ceremonial conformity; renewable energy use and environmental capex narrow the gap, while results are robust to trimming and rank recoding but the index effect vanishes when CDP is replaced by the LSEG Environmental Pillar Score.","tokens_in":2022,"tokens_out":638,"duration_ms":51632,"significance":"If the DPG construction isolates greenwashing, the finding that detected greenwashing is conditional on the rating lens contributes to the Aggregate Confusion literature by documenting firm-level heterogeneity in measurement. The transparency in reporting the vanishing index effect under the alternative rating, the six-stage specification search with multiple robustness checks, and the use of HC3 errors are strengths that enhance credibility of the estimation approach.","major_comments":[{"comment":"DPG construction (abstract and variable definition): The performance component relies on absolute (unnormalized) emissions. Flagship index members are systematically larger firms within the STOXX 600 sample, which can mechanically widen the standardized divergence even absent any disclosure-performance mismatch. The six-stage search and reported robustness checks do not indicate tests of emissions intensity, log(assets), or revenue as controls or alternative DPG definitions, leaving open a scale confound that directly affects the central interpretation of the flagship coefficient as ceremonial conformity.","section":"DPG construction"},{"comment":"Model selection (six-stage process described in abstract): Stepwise forward selection under corrected AIC across 421 candidates is known to inflate Type I error rates and produce overfitted models; while VIF and Cook's distance are applied, this procedure remains load-bearing for the reported flagship index coefficient and its p-value.","section":"Model selection process"},{"comment":"TCFD result (abstract): The positive coefficient (β = +0.86, p < 0.05) is identified off a small non-supporting subgroup, so the estimate cannot support a precise magnitude claim even if the directional sign is retained.","section":"Results on TCFD"}],"minor_comments":[{"comment":"A summary table listing the number of specifications retained after each of the six stages would improve transparency of the selection process.","section":"Methods"},{"comment":"The abstract states the sample comprises firms from four sectors but does not report the exact sector breakdown or any sector fixed effects; adding this detail would clarify generalizability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful comments, which help strengthen the paper. We address each major comment below, indicating where revisions will be made.","responses":[{"response":"We acknowledge this potential scale confound. Although the DPG is constructed as a standardized measure, larger firms may have higher absolute emissions. To address this, we will incorporate log(assets) as an additional control variable in the model selection process and test an alternative DPG definition using emissions intensity (emissions per unit of revenue). These additional specifications will be reported in a revised robustness section.","revision_made":"yes","referee_comment":"[DPG construction] DPG construction (abstract and variable definition): The performance component relies on absolute (unnormalized) emissions. Flagship index members are systematically larger firms within the STOXX 600 sample, which can mechanically widen the standardized divergence even absent any disclosure-performance mismatch. The six-stage search and reported robustness checks do not indicate tests of emissions intensity, log(assets), or revenue as controls or alternative DPG definitions, leaving open a scale confound that directly affects the central interpretation of the flagship coefficient as ceremonial conformity."},{"response":"We recognize the limitations of stepwise selection procedures, including potential inflation of Type I errors. However, our six-stage process includes correlation screening, VIF filtering to address multicollinearity, Cook's distance for influence, and HC3 robust errors. Moreover, the flagship index result is robust to trimming, rank recoding, and alternative rating measures. In revision, we will add a discussion of these limitations and report results from a pre-specified model that includes the key variables of interest without relying solely on stepwise selection.","revision_made":"partial","referee_comment":"[Model selection process] Model selection (six-stage process described in abstract): Stepwise forward selection under corrected AIC across 421 candidates is known to inflate Type I error rates and produce overfitted models; while VIF and Cook's distance are applied, this procedure remains load-bearing for the reported flagship index coefficient and its p-value."},{"response":"We agree with this assessment. The manuscript already notes that the TCFD coefficient is identified off a small group of non-supporting firms and should be interpreted as directional rather than providing a precise magnitude. No further change is required on this point.","revision_made":"no","referee_comment":"[Results on TCFD] TCFD result (abstract): The positive coefficient (β = +0.86, p < 0.05) is identified off a small non-supporting subgroup, so the estimate cannot support a precise magnitude claim even if the directional sign is retained."}],"tokens_in":1683,"tokens_out":576,"duration_ms":44139,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core finding is that whether the same firms look like greenwashers depends on the ESG rating lens, with flagship index membership as the strongest predictor of a wider Disclosure-Performance Gap. That conditional result on rating choice is the most useful part.\n\nThe work is new in moving the aggregate confusion hypothesis to the firm level with an explicit DPG construct on 200 STOXX Europe 600 firms. The six-stage specification search is laid out clearly, and they report that swapping CDP for LSEG eliminates the index effect while renewable energy use survives. That transparency about sensitivity is worth crediting.\n\nThe soft spots are real. Emissions performance is an absolute quantity, so larger firms (more likely in flagship indices) will show mechanically wider gaps unless emissions are scaled by revenue, assets, or output. Nothing in the abstract or described checks indicates they tested log assets or intensity as a control or alternative construction. The TCFD coefficient rests on a small non-supporting group, and stepwise forward selection under AIC can still inflate apparent significance even after VIF and Cook's distance steps. The index result vanishing under the alternative rating already weakens the ceremonial conformity claim.\n\nThis is for people working on ESG data reliability and CSRD implementation. It has enough structure and reported checks to deserve referee time rather than desk rejection, though any review would need to press on the DPG construction and simpler robustness specs.","headline":"The paper extends rating disagreement to firm-level greenwashing via a new DPG measure, but the flagship index result likely reflects unnormalized absolute emissions rather than ceremonial conformity.","tokens_in":2481,"tokens_out":359,"would_cite":false,"duration_ms":28399,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The measured gap between environmental disclosure and performance varies by ESG rating, with flagship index membership predicting larger gaps only under the CDP score.","keywords":["greenwashing","ESG ratings","disclosure performance gap","index membership","environmental disclosure","European firms","aggregate confusion hypothesis"],"falsifier":"Observing no difference in the gap for index members when using a third independent rating source or when emissions data is verified by a single standard would challenge the rating-dependence claim.","tokens_in":2786,"feed_emoji":"📉","tokens_out":613,"duration_ms":30632,"temperature":0.7,"pith_summary":"This paper measures the Disclosure-Performance Gap as the standardized difference between firms' voluntary environmental disclosures and their actual emissions outcomes for 200 large European companies. It tests what predicts this gap and finds that membership in flagship stock indexes is the strongest factor widening it, pointing to ceremonial conformity in reporting. Renewable energy use and environmental spending narrow the gap, while governance variables do not. Replacing the CDP Climate Score with the LSEG Environmental Pillar Score removes the index effect, demonstrating that conclusions about greenwashing depend on the chosen rating lens.","feed_headline":"Index membership widens greenwashing gap under some ESG ratings","feed_subtitle":"For 200 European firms, flagship status predicts larger disclosure-performance gaps with CDP scores but not LSEG, showing rating choice matt","key_machinery":"The Disclosure-Performance Gap (DPG), defined as the standardized divergence between voluntary disclosure and realized emissions performance, compared across CDP and LSEG ratings to identify rating-dependent predictors.","core_discovery":"The paper establishes that the Disclosure-Performance Gap is significantly wider for firms in flagship indexes, with a coefficient of 0.78, consistent with institutional pressures for symbolic compliance rather than substantive performance. This effect disappears when using an alternative rating source, while actual investments in renewables and capital expenditure reduce the gap, supporting signalling over monitoring explanations.","pith_inferences":["Rating agencies' differing methodologies can lead to inconsistent greenwashing identifications across the same firms.","Investors relying on single ratings may misjudge firm environmental commitment.","Future studies could test if similar patterns hold after mandatory reporting under CSRD begins.","The ceremonial conformity effect may apply to other ESG dimensions like social or governance scores."],"forward_implications":["Flagship index membership correlates with larger gaps under CDP ratings.","Renewable energy use and environmental capex narrow the measured gap.","TCFD endorsement shows a positive but directional association with wider gaps.","Governance and monitoring variables have no explanatory power for the gap.","Detected greenwashing is conditional on the ESG rating applied."],"fun_headline_variants":["Same firms get different greenwashing verdicts by ESG rating","Index membership widens talk-walk gaps under CDP ratings","Rating lens flips greenwashing detection for index members","Flagship firms show bigger gaps in CDP but not LSEG scores"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The standardized divergence between disclosure and emissions performance isolates intentional greenwashing rather than differences in reporting incentives or measurement errors.","fun_headline_variants_meta":{"raw":{"variants":["Same firms get different greenwashing verdicts by ESG rating","Index membership widens talk-walk gaps under CDP ratings","Rating lens flips greenwashing detection for index members","Flagship firms show bigger gaps in CDP but not LSEG scores"]},"model":"grok-4.3","cost_usd":0.00439,"raw_usage":{"total_tokens":2262,"prompt_tokens":797,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":43899500,"prompt_tokens_details":{"text_tokens":797,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1401,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":797,"tokens_out":64,"duration_ms":19864,"temperature":1.0,"reasoning_tokens":1401,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T02:37:23.630400+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Observing no difference in the gap for index members when using a third independent rating source or when emissions data is verified by a single standard would challenge the rating-dependence claim.","supporting_citations":[],"review_version":1}