{"id":"89255b8e-a619-42db-9ec8-03ab03f9c95b","arxiv_id":"2411.13958","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A human-annotated, economics-specific sentiment lexicon outperforms existing dictionaries in several in-sample economic text applications.","lead":"This paper builds a new economics-specific sentiment dictionary, the Economic Lexicon, with human-annotated scores for more than 6,600 words. It claims this lexicon outperforms existing dictionaries at capturing uncertainty, consumer sentiment, and recession risk in economic news.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EL's superiority claim is not established: evaluation is in-sample on the construction corpus, and the paper's own footnote admits LMD retains unique predictive content.","rationale":"The reader's conditional verdict is appropriate and I do not move it. The paper delivers a useful, reproducible human-annotated lexicon, with a downloadable resource and a replication package, which are real contributions. However, the central superiority claim is not yet supported. The most load-bearing problem is that the evaluation is in-sample: the EL's vocabulary was selected from the same newspaper corpus (and with the same 1980-2020 window) on which the EP regressions are computed. Because the corpus is not a random sample of economic text, and because Section 4.2 itself observes that recession periods generate more economic sentences, the frequency-based word selection can embed a recession-predictive vocabulary into the lexicon. That makes the dominance in Tables 2, 3, and 5 hard to interpret. The footnote in Section 7 is even more direct: it says that when EP_EL is controlled for, only EP_LMD retains significance, implying a competitor has unique predictive content. This contradicts the Introduction's 'makes the other EP measures irrelevant' claim. I also agree with the reader that the AUC test in Figure 9 is misread: a one-sided p-value that fails to reject the null of EL AUC >= alternative does not demonstrate EL superiority. These issues are fixable with a holdout evaluation and a more careful statistical interpretation, so a conditional verdict (revise with out-of-sample evidence) is the right call; the paper should not be rejected outright because the lexicon itself is potentially valuable and the authors are transparent enough to report the conflicting robustness result.","tokens_in":18507,"tokens_out":6465,"duration_ms":59597,"concrete_test":"Rebuild the EL using only pre-2010 data from the same sources, then re-estimate Models (2), (3), and (4) on the held-out 2010-2020 period (or on an unused outlet set, e.g., central bank documents) and compare EL's dominance against LMD, REN, and SSW. If EL's coefficient loses significance or the alternative lexicons become jointly significant, the reported superiority is an in-sample selection artifact. In addition, request the full regression output behind footnote 19: if EP_LMD remains significant when EP_EL is included, the claim that EL makes other EP measures irrelevant is false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that EP_EL is significant and makes other EP measures irrelevant (Section 1) is supported mainly by regressions run on the same six US newspapers (1980-2020) used to build the EL (Sections 3 and 4.2). The lexicon's word list was selected by applying a 65-occurrence frequency threshold to exactly this corpus, and Section 4.2 notes that economic sentences are oversampled during recessions. A lexicon whose vocabulary is selected from the same corpus on which it is evaluated can therefore inherit a pro-cyclical negative-word bias, making its EP spuriously informative for recession forecasting. This is an in-sample artifact, not evidence of general superiority. The paper's own Section 7 footnote 19 compounds the problem: when EP_EL is added as an extra regressor to model (5), 'only EP_LMD,t retains its significance,' indicating that LMD contains terms not in EL that matter for recessions. That statement is in direct tension with the Introduction's claim that once EL is included the other EP measures become irrelevant. The AUC analysis in Figure 9 also does not help: failing to reject a null that EL's AUC is at least as large as the alternatives does not establish that it is larger. The central claim therefore rests on an evaluation loop closed by construction and is partially contradicted by the authors' own robustness exercise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new domain-specific Economic Lexicon (EL) for sentiment analysis in economics, built from a large corpus of news articles and central bank documents. The lexicon is constructed by extracting noun phrases around a list of economic concepts, selecting frequent modifier words, and assigning sentiment scores via human annotators, with ambiguous terms removed. The EL is then compared with three existing lexicons (LMD, REN, SSW) in several applications: explaining movements in the VIX and economic policy uncertainty, correlating with consumer sentiment, and forecasting NBER recessions using an Economic Pessimism (EP) measure. The authors claim that EL is superior because it covers more economically relevant terms and categorizes sentiment more accurately, and that its EP measure is significant and makes alternative EP measures irrelevant when included jointly.","tokens_in":18752,"tokens_out":3441,"duration_ms":34193,"significance":"If the claims were established, the paper would provide a valuable, freely available economic sentiment lexicon with fine-grained human-annotated scores, and a transparent construction pipeline. The authors also make available a replication package, which is a commendable feature. However, the central superiority claim is not established by the current evidence: the evaluation is performed largely on the same corpus used to select the lexicon's vocabulary, the word-level 'accuracy' claim is benchmarked only against the annotators' own scores, and the paper's own footnote 19 contradicts the introduction's dominance claim. The resource itself is a useful contribution, but the empirical evidence for its general superiority needs substantial additional work before the claim can be accepted.","major_comments":[{"comment":"The evaluation is in-sample by construction. The EL vocabulary is selected by applying a frequency threshold of 65 occurrences to the corpus of six US newspapers over 1980-2020 (Section 4.2), and the EP measure used for all comparisons in Section 6 is computed on exactly the same six newspapers over the same period. This creates a selection loop: words are chosen because they are frequent in this corpus and then the EP measure is shown to be informative in this same corpus. The superior performance in Tables 2, 3, and 5 may therefore reflect frequency-based overfitting rather than a general property of the lexicon. The authors should provide an out-of-sample evaluation on text not used in the construction (e.g., UK newspapers, central bank documents, or a temporal holdout such as constructing the lexicon on pre-2000 data and evaluating on post-2000 data). The oversampling of negative terms during recessions noted in Section 4.2 makes this concern more acute, as the in-sample evaluation could inherit a pro-cyclical negative-word bias.","section":"Section 4.2 and Section 6"},{"comment":"There is an internal contradiction between the claim in the Introduction that 'once included, makes the other EP measures irrelevant' and the statement in footnote 19 that when EP_EL is added as an additional regressor to model (5), 'only EP_LMD,t retains its significance.' This implies that LMD contains predictive content not captured by EL, which directly undermines the irrelevance claim. The authors must either reconcile these findings or substantially revise the claim. As written, the paper simultaneously asserts dominance and provides evidence against it.","section":"Section 7, footnote 19"},{"comment":"The claim that EL provides 'more accurate categorization of the word sentiment' is not supported by an external benchmark. The word-level 'accuracy' is assessed against the same human annotator scores that define the lexicon (Section 4.3), and Section 5 explicitly notes that no human-annotated sentence-level scores are available to benchmark the lexicons. The comparison to other lexicons is thus an agreement study, not a validation against a gold standard. A human-annotated benchmark such as the 800 articles used in Shapiro et al. (2022) should be used to ground the accuracy claim; Section 7 acknowledges this as future work, but it is load-bearing for the paper's central claim.","section":"Section 5 and Section 4.3"},{"comment":"The discussion of the AUC results overstates what the test shows. The p-values in Figure 9 are for the null hypothesis that the AUC of the EL-based model is larger than or equal to the alternative's AUC; failing to reject this null does not establish that EL is superior. The text states that 'at most horizons' EL provides 'better performance, at least in terms of larger AUC,' but the only horizon at which the null is rejected is h = 2 for LMD. This should be described as a non-inferiority result, not a superiority result, and the conclusions in Sections 1 and 6.4 should be adjusted accordingly.","section":"Section 6.3, Figure 9"}],"minor_comments":[{"comment":"The caption refers to 'RLM' where the intended dictionary is LMD (Loughran-McDonald); this spelling should be corrected.","section":"Figure 8 caption"},{"comment":"The phrase 'makes the other EP measures irrelevant' is too strong even apart from footnote 19; a more precise statement would say that EP_EL is significant when included jointly with each alternative, as in Tables 2, 3, and 5, while noting that the joint inclusion of more than one alternative is not reported.","section":"Section 1, last paragraph"},{"comment":"The notation Si,l is defined as the sentiment score of word i in lexicon l, but for the fine-grained lexicons the equation uses Si,l directly even though the text says the scores were converted to ±1 for comparability; clarifying this, as is done in footnote 13, would improve readability.","section":"Section 6, Equation (1)"},{"comment":"The data availability statement references the replication package for the related paper by Barbaglia et al. (2024) at openICPSR; it would be helpful to state explicitly whether this package includes the EL lexicon and the code for the current paper's tables and figures.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The central finding of the paper is a superiority claim that is currently supported only by in-sample regressions, a contradictory footnote, and a weakly interpreted AUC test. The construction of the lexicon is sound and the resource is valuable, but the empirical evidence for dominance is not yet at the level required for publication. The paper would benefit from an out-of-sample evaluation, an external human benchmark, and a careful revision of the claims so that they match the reported results. I would not recommend rejection, as the issues are addressable within the scope of the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuine resource paper with a reproducible pipeline, but the paper's central claim of superiority over existing lexicons is not yet established. The in-sample evaluation loop and a misread AUC test are the two places I'd push back.\n\nWhat's new and good: the EL lexicon — 6,670 human-annotated terms with fine-grained scores, built by dependency-parsing around 355 economic entities — is a useful addition. The construction is transparent, the data are available (replication package on openICPSR), and the authors are honest about several design choices, including the frequency threshold and the lack of a sentence-level benchmark. The decomposition exercise in Table 6 is a good idea: it separates the contribution of corrected sentiment scores from expanded coverage, and it actually shows that both matter for forecasting. That part is careful.\n\nWhere it goes soft: the evaluation uses the same six newspapers (1980–2020) from which the vocabulary was selected, so the EP measure is computed on the construction corpus. The 65-occurrence threshold is applied to that same corpus. That doesn't automatically kill the resource, but it does mean the in-sample superiority could be a word-selection artifact, particularly since economic sentences are oversampled during recessions. The bigger logical problem is the AUC test in Figure 9. The null is AUC_EL >= AUC_alt; failing to reject that null merely says you can't rule out that EL is no worse. It does not establish that EL is better. The paper reads it as evidence of superiority, which is wrong. And footnote 19 (Section 6.4) concedes that when EP_EL is added to model (5), EP_LMD retains significance — directly contradicting the Introduction's claim that EL makes the other measures irrelevant. Those two issues together mean the headline superiority claim is not supported.\n\nWho this is for: people who work with economic text and need a domain sentiment lexicon. The resource itself deserves attention; the evaluation needs revision. I'd send this to a serious referee, but I'd expect major revisions: out-of-sample validation on a different corpus, a corrected interpretation of the AUC test, and public admission that LMD retains some unique content. Recommendation: accept for review, ask for these fixes.","headline":"A genuinely useful and reproducible lexicon resource, but the superiority claim over existing dictionaries is not yet established due to in-sample evaluation and a misread AUC test.","tokens_in":19262,"tokens_out":2847,"would_cite":false,"duration_ms":26216,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a sentiment lexicon built from economic text and scored by human annotators yields a pessimism measure that dominates three established dictionaries in explaining uncertainty, consumer sentiment, and recession…","keywords":["economic lexicon","sentiment analysis","textual analysis","economic pessimism","human annotation","dependency parsing","recession forecasting","dictionary comparison"],"falsifier":"Compute the EL-based pessimism index and the three rival indices on an out-of-sample corpus never used in lexicon construction, for example central bank speeches or non-US news from after 2020, and test whether EL still dominates in explaining VIX, EPU, consumer sentiment, and recession forecasts. If the EL advantage shrinks or reverses under this test, the claimed general superiority would be refuted.","tokens_in":18312,"feed_emoji":"📈","tokens_out":6705,"duration_ms":63821,"temperature":0.7,"pith_summary":"This paper builds a sentiment dictionary, the Economic Lexicon (EL), from a large corpus of economic news and central bank documents, and then argues that a simple text-based measure of economic pessimism built from it outperforms measures built from three established lexicons. The EL selects words by parsing sentences that discuss economic concepts and has human annotators score each word on a scale from -1 to 1. The paper shows that the EL-based pessimism index correlates with financial and policy uncertainty and consumer sentiment, and helps predict recessions; in joint regressions it makes the other lexicons' measures statistically irrelevant. The authors attribute this advantage to wider coverage of economics-specific terms and a more accurate assignment of word sentiment than the alternatives. A sympathetic reader would care because the result offers a reusable, human-annotated tool for measuring tone in economic text.","feed_headline":"New lexicon beats established dictionaries at recession forecasting","feed_subtitle":"A human-annotated dictionary of economics-specific terms makes other sentiment measures statistically irrelevant in joint tests.","key_machinery":"The central object is the EL lexicon itself, together with the Economic Pessimism (EP) measure that turns it into a time series. Construction follows four steps: sentences are selected only if they contain one of 355 economic entities drawn from an economic topical taxonomy and central bank documents; dependency parsing extracts adverb, verb, and adjective modifiers of those entities, keeping words that occur at least 65 times; ten human annotators assign each term a sentiment score in [-1,1], and the median score becomes the word's tone; and words flagged as ambiguous are removed, leaving 6,670 terms. The EP measure is the negative of the frequency-weighted sum of term scores, so positive values indicate that negative words outnumber positive words in a month. This pipeline carries the argument because it is designed to include words that co-occur with economic concepts and to avoid finance-only or general-purpose words, and the fine-grained human scores allow both signed categorization and strength-sensitive sentiment measurement.","core_discovery":"The authors' central claim is that the Economic Lexicon, constructed from economic text rather than general-purpose or finance-only text, produces a sentiment measure that is more accurate for economics than the Loughran-McDonald financial dictionary, the Renault model-based lexicon, and the Shapiro-Sudhof-Wilson news lexicon. After extracting noun phrases tied to 355 economic concepts from over 13 million newspaper articles and central bank documents, dependency parsing yields modifier words; frequency filtering and human annotation leave 6,670 terms with median sentiment scores on [-1,1] and ambiguous terms removed. An Economic Pessimism index formed as the frequency-weighted sum of these scores is significantly related to VIX, economic policy uncertainty, and the Michigan Consumer Sentiment Index, and adds predictive power for NBER recessions. Once the EL-based index is included in regressions alongside the other lexicons' indices, the alternatives lose significance. The paper attributes this result to two factors: EL includes more sentiment-bearing terms that actually appear in economic discussion, and its human-annotated scores correct the sign or strength of words that the other lexicons misclassify.","pith_inferences":["My inference: the same-corpus evaluation in the paper is the main threat to generalization; an out-of-sample corpus test would settle whether EL's dominance is a property of the lexicon or of the newspapers used to build it.","My inference: the dependency-parsing-plus-human-annotation pipeline could be transferred to other specialized domains, such as climate finance or health policy, where general sentiment dictionaries are known to misfire.","My inference: the fine-grained human scores might also serve as a benchmark to audit model-based sentiment classifiers, since the paper shows machine and human scores agree in sign on most but not all shared words.","My inference: the paper's evidence that negative-word sentiment is more volatile and procyclical than positive-word sentiment implies that future sentiment measures could weight the two components separately rather than forcing them into a single pessimism index."],"forward_implications":["Using the EL, a researcher can compute economic sentiment from news or central bank text without training a model-based sentiment classifier, because the lexicon ships with human median scores on a continuous scale.","In the paper's regressions, the EL-based pessimism measure absorbs the explanatory power of the LMD, REN, and SSW measures for VIX, EPU, consumer sentiment, and recessions at a three-month horizon, so those rival indexes add no significant information once EL is included.","The finding that positive and negative economic words appear with similar frequency suggests sentiment indexes built only from negative words may miss a comparable share of sentiment variation, and the EL supplies positive terms for that purpose.","Re-scoring disagreeing terms with EL values or adding EL-only terms improves the recession-forecasting performance of the other lexicons, implying that both word coverage and sentiment assignment contribute to EL's edge.","Because EL scores are continuous, the same lexicon can be used for simple signed counts and for strength-weighted sentiment measures such as the EP index."],"supporting_citations":[{"why":"Supplies the LMD financial dictionary that EL is compared against and the human-annotation approach EL follows.","marker":"Loughran and McDonald (2011)"},{"why":"Supplies the REN model-based lexicon used as an alternative in the EP comparisons.","marker":"Renault (2017)"},{"why":"Supplies the SSW lexicon and a sentence-level human benchmark for lexicon accuracy.","marker":"Shapiro et al. (2022)"},{"why":"Provides the EPU uncertainty index against which the EL-based pessimism measure is tested.","marker":"Baker et al. (2016)"},{"why":"Provides the VADER algorithm whose sentence scores underlie SSW's model-based word sentiments.","marker":"Hutto and Gilbert (2014)"},{"why":"Defines the universal dependency relations used to extract modifier words from economic sentences.","marker":"De Marneffe et al. (2014)"},{"why":"Introduces the pessimism measure and the negative-word focus that this paper re-examines.","marker":"Tetlock (2007)"}],"fun_headline_variants":["Economic lexicon outperforms finance dictionaries in recession tests","New economic dictionary improves recession forecasting over rivals","Domain-specific sentiment lexicon wins on economic text","Human-annotated economics terms beat legacy lexicons"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the corpus used to select EL's words is representative of economic text generally; the comparisons are run on the same six US newspapers (1980-2020) from which the vocabulary was extracted, so if those newspapers are not representative, EL's apparent edge could be an artifact of in-sample word selection.","fun_headline_variants_meta":{"raw":{"variants":["Economic lexicon outperforms finance dictionaries in recession tests","New economic dictionary improves recession forecasting over rivals","Domain-specific sentiment lexicon wins on economic text","Human-annotated economics terms beat legacy lexicons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1226,"prompt_tokens":851,"completion_tokens":375,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":326}},"tokens_in":467,"tokens_out":375,"duration_ms":4343,"temperature":1.0,"reasoning_tokens":326,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:41:00.156294+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the EL-based pessimism index and the three rival indices on an out-of-sample corpus never used in lexicon construction, for example central bank speeches or non-US news from after 2020, and test whether EL still dominates in explaining VIX, EPU, consumer sentiment, and recession forecasts. If the EL advantage shrinks or reverses under this test, the claimed general superiority would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the LMD financial dictionary that EL is compared against and the human-annotation approach EL follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the REN model-based lexicon used as an alternative in the EP comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the EPU uncertainty index against which the EL-based pessimism measure is tested."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the VADER algorithm whose sentence scores underlie SSW's model-based word sentiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the pessimism measure and the negative-word focus that this paper re-examines."}],"review_version":1}