{"id":"679e5d62-5d10-4d8a-adcd-a21c36e1d1ca","arxiv_id":"2411.13180","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reviewing 71 published studies, this paper finds sentiment is significant in most models, yet complex sentiment measures do not clearly beat simpler ones, and the BW index appears to be losing explanatory power.","lead":"An author-led review of 71 research papers asks whether investor sentiment improves stock return models. It reports that sentiment usually adds explanatory power, but that the popular BW sentiment index may be losing relevance in recent data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own quantitative analysis contradicts its headline claim: Section 4.2.1 and Table 9 find no statistically significant improvement in adjusted R2 from sentiment augmentation, yet the abstract and Section 4.1.7 assert RH1 is confirmed.","rationale":"The reader correctly identifies the central claim: the abstract's assertion that higher complexity of sentiment measures and models improves R2, and the Section 4.1.7 confirmation of RH1. However, the reader's weakest-assumption focuses on the representativeness of the 71-paper sample and cross-study comparability of adjusted R2. My stress-test finds a more direct, internal problem: the paper's own quantitative section explicitly disclaims the evidence needed for RH1 and reports insignificant differences in R2 means across model classes. The abstract and Section 4.1.7 are not just weakly supported by a non-representative sample; they are contradicted by the paper's own Table 9 and t-tests. This is a correctness risk that does not depend on external assumptions about the literature. It is the most load-bearing concern because it attacks the main contribution as stated. The paper still has value as a structured review and descriptive synthesis, and the qualitative observation that sentiment is often significant is defensible. But the headline claim should be downgraded or qualified. Since the reader's verdict is already CONDITIONAL and calls for reconciling this overclaim, my concern does not move the verdict; it reinforces the conditional and identifies a sharper target for the required revision.","tokens_in":32761,"tokens_out":2556,"duration_ms":25037,"concrete_test":"Recompute the Section 4.2.1 comparison and, more importantly, collect the one quantity RH1 actually requires: the incremental adjusted R2 before vs after adding sentiment in each study. For Table 9, re-run Welch's t-test using the reported n, mean, and SD (e.g., single-factor 0.23±0.29, n=11; multifactor 0.32±0.26, n=13). If the p-value remains >0.10, the quantitative conclusion is unchanged. Then, to test whether the qualitative claim can substitute, extract from Appendix A (or from the 71 full texts) the number of studies reporting any delta-adjusted-R2 from sentiment augmentation. If that subset is empty or small, or if the mean delta is not positive, the abstract's complexity-improves-R2 claim must be withdrawn or explicitly relabeled as a descriptive observation with no statistical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assertion is that augmenting models with sentiment improves the coefficient of determination (RH1), restated in the abstract and concluded in Section 4.1.7. The paper's own quantitative analysis undermines this. Section 4.2.1 states that RH1 'cannot be directly verified quantitatively due to the lack of appropriate data (e.g. incremental R-squared).' The only quantitative evidence, Table 9, reports mean adjusted R2 of 0.23 (single-factor, n=11), 0.20 (medium complex, n=14), and 0.32 (multifactor, n=13). The author's t-tests comparing single-factor vs multifactor and medium-complex vs multifactor both give p-values above 10%, i.e., insignificant differences. The abstract's claim that 'higher complexity of sentiment measures and models improves the coefficient of determination' goes beyond even the insignificant means, because the model classes differ in many ways besides sentiment complexity (factor structure, data frequency, asset universe), and no paper in Appendix A is shown to report an incremental R2 from adding sentiment. The qualitative support cited in Section 4.1.7 is the frequency of significant sentiment coefficients (65/71 studies), but statistical significance of a coefficient is neither necessary nor sufficient for an increase in adjusted R2, particularly when the baseline model varies. The paper itself cautions at the end of Section 4.2.1 that 'one should be careful with interpreting those results,' so the confirmed-hypothesis framing in the abstract is an overstatement. This is an internal inconsistency, not a disagreement about external consensus: the effect may be real, but this review does not establish it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reviews 71 empirical studies published between 2000 and 2021 that include investor sentiment in asset pricing models. It categorizes sentiment measures (direct, indirect, composite, media-based) and model classes (single-factor, medium-complex, multifactor, machine learning), summarizes qualitative findings on coefficient significance, and reports a quantitative comparison of mean adjusted R-squared across model classes. The abstract claims that 'higher complexity of sentiment measures and models improves the coefficient of determination,' while the second hypothesis about predictive power is judged unverifiable. The author concludes that the first hypothesis (RH1) is confirmed.","tokens_in":33040,"tokens_out":4000,"duration_ms":38073,"significance":"If the central claim were supported, the paper would be a useful synthesis of the sentiment-augmented asset pricing literature. Its main asset is the detailed appendix cataloguing 71 studies with asset, period, frequency, sentiment measure, model, sign, and significance, which could serve as a resource for future meta-analyses. However, the central claim is contradicted by the paper's own quantitative evidence: Section 4.2.1 reports t-tests on adjusted R-squared means with p-values above 10%, and the paper explicitly states that RH1 cannot be directly verified quantitatively due to the lack of incremental R-squared data. The qualitative frequency of significant sentiment coefficients (65/71 studies) is not evidence of an improvement in the coefficient of determination. The paper also lacks machine-checked proofs or reproducible code; its contribution is a review, so the internal inconsistency in the main conclusion is the decisive issue. With a corrected framing, the descriptive value of the survey could still justify publication.","major_comments":[{"comment":"The assertion that 'the obtained results confirm the first research hypothesis (RH1)' is not supported by the paper's own quantitative analysis in Section 4.2.1, where the t-tests comparing adjusted R-squared means (Table 9) are insignificant (p>0.10), and where the paper states that RH1 'cannot be directly verified quantitatively due to the lack of appropriate data (e.g. incremental R-squared).' The frequency of statistically significant sentiment coefficients (65/71 studies) does not imply an improvement in the coefficient of determination, and the abstract's stronger claim that 'higher complexity of sentiment measures and models improves the coefficient of determination' is not tested anywhere in the manuscript.","section":"§4.1.7 and Abstract"},{"comment":"The comparison of mean adjusted R-squared across single-factor, medium-complex, and multifactor models does not test RH1, because the model classes differ in factor structure, data frequency, asset universe, and sample period, and no paper in Appendix A is shown to provide the incremental R-squared from adding sentiment to a baseline model. The paper itself cautions that 'one should be careful with interpreting those results,' but this caution is not carried into Section 4.1.7 or the conclusions, where RH1 is treated as confirmed.","section":"§4.2.1, Table 9"},{"comment":"The abstract's claim that 'higher complexity of sentiment measures and models improves the coefficient of determination' conflates two distinct dimensions: the complexity of the sentiment measure and the complexity of the asset pricing model. The analysis provides no decomposition of R-squared gains between these two dimensions, and the direct comparative evidence on sentiment-measure complexity (Section 4.1.4, five of eight comparative studies favoring composite indices) concerns significance or predictive power, not the coefficient of determination.","section":"§1 and Abstract"},{"comment":"The inclusion criterion of a minimum of 60 citations on Web of Science, together with the exclusion of non-English papers and papers not reporting detailed model specifications, means that the 71-paper sample is not representative of the full sentiment-asset pricing literature. This limits the generalizability of prevalence claims such as 'sentiment is almost always an important factor' (Section 4.1.7) and should be stated as a first-order limitation in the conclusions, not only as a procedural choice in Section 3.1.","section":"§3.1"}],"minor_comments":[{"comment":"The phrase 'The research failed to reject (both quantitatively and qualitatively) one out of the two hypotheses (RH1)' is ambiguous and seems to use 'failed to reject' in the opposite sense of the standard statistical meaning; the paper actually claims RH1 is confirmed, so the conclusion should state this directly.","section":"§5"},{"comment":"The citation 'Fang and Taylor (2021)' for the MGMT and PERF factors is missing from the bibliography.","section":"§2.3.3"},{"comment":"There are several spelling and notation errors in Table 8, including 'Amercian' (entries 6, 13, 24, 31, 36, 39), 'balance' for 'imbalance' (entry 9), and 'Sign' entries such as 'Impossible to detect' (entries 28, 43, 49, 56, 57, 62, 69) that are not signs and should be moved to a separate column or explained.","section":"Appendix A, Table 8"},{"comment":"The definitions in Eqs. (1)–(3) present sentiment as a price, return, or characteristic difference from a rational benchmark, but the text does not explain how this conceptual definition relates to the operational measures (e.g., BW index, survey measures, media-based indices) used in the reviewed studies.","section":"§2.1.1"},{"comment":"The text refers to 'medium-factor' models in the first sentence of Section 4.2.1, while elsewhere the paper uses 'medium complex' models; the terminology should be standardized.","section":"§4.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's main problem is an internal inconsistency: the abstract and Section 4.1.7 claim confirmation of RH1, while the paper's own quantitative test in Section 4.2.1 finds insignificant differences in adjusted R-squared and admits that incremental R-squared data are lacking. This is a load-bearing issue, but it can be fixed by carefully rewording the claims and conclusions to match the evidence. The review dataset itself, particularly Appendix A, may be a useful contribution if the central claims are revised. No concerns about citation manipulation or circularity arose in my reading."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know before reading: this is a genuinely useful map of the sentiment-asset pricing literature, but the headline claim is not backed by the paper's own numbers. The abstract says higher complexity of sentiment measures and models improves the coefficient of determination, and Section 4.1.7 says RH1 is confirmed. Yet when the author tries to test this quantitatively in Section 4.2.1, he finds the adjusted R-squared differences between single-factor, medium-complex, and multifactor models are insignificant (p-values above 10%), and he admits that incremental R-squared data are missing. That is an internal contradiction, and it is the main thing to fix.\n\nWhat is genuinely new: this is the most comprehensive systematic review I have seen on investor sentiment proxies in asset pricing, covering 71 papers from 2000 to 2021 with a transparent selection protocol and a detailed appendix table. The qualitative synthesis is useful: sentiment coefficients are significant in at least one specification in 65 of 71 studies, the BW index appears to be losing relevance in recent years, and the effect depends heavily on asset, period, and measure. The breakdown by model class and data frequency is valuable, and the paper makes a fair point that heterogeneous designs make meta-analytic claims hard.\n\nWhere the soft spots are: the R-squared claim is the obvious one. The t-tests in Table 9 are underpowered and the comparison is confounded because model classes differ in factor structure, data frequency, and universe, not just sentiment complexity. The author even warns about this, but then contradicts the warning in the abstract and conclusions. A second, lesser issue is the selection criterion: requiring a minimum citation count on Web of Science biases toward older, influential papers and may overrepresent the BW index; the paper acknowledges this but does not discuss how it might affect the prevalence claims. The quantitative analysis is presented as descriptive, which is fine; it just should not be used to confirm RH1.\n\nOverall: the synthesis is solid and the paper deserves a serious referee, but the abstract and conclusions need to be rewritten so the main claim matches the evidence. The author should clearly say the R-squared finding is descriptive, not statistically significant, and should release the extracted dataset. After those changes, this would be a citable survey.\n\nI'd bring it to a reading group as a starting point for proxy selection, but I'd frame it as 'a useful map with an overambitious headline.'","headline":"A useful but overclaimed systematic review: the abstract asserts RH1 confirmed, while the paper's own tests show insignificant R-squared gains; revision can fix it.","tokens_in":33573,"tokens_out":2151,"would_cite":true,"duration_ms":22335,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review of 71 empirical studies claims that adding investor sentiment proxies to asset pricing models improves the coefficient of determination, while evidence that more complex sentiment measures predict better than simple ones…","keywords":["Investor sentiment","Asset pricing","Multifactor models","Behavioral finance","Coefficient of determination","Sentiment measures","Baker-Wurgler index","Literature review"],"falsifier":"Compute the incremental adjusted R-squared obtained by adding a sentiment proxy to the same base factor model across a complete set of published studies and test whether the average gain is positive after penalizing the extra parameters. If that average is zero or negative, the central claim that sentiment proxies improve the coefficient of determination collapses.","tokens_in":32554,"feed_emoji":"📈","tokens_out":5314,"duration_ms":54066,"temperature":0.7,"pith_summary":"This paper reviews 71 empirical studies published between 2000 and 2021 that add investor sentiment measures to asset pricing models. It claims that sentiment is significant in at least one tested relationship in 65 of the 71 studies, and that adding sentiment proxies improves the coefficient of determination of the models. It also claims that more complex sentiment measures and models are associated with higher adjusted R-squared. The paper does not claim that complex measures forecast better: only nine studies compared measures directly, and the evidence was too thin to confirm that hypothesis. If the review is right, standard rational factor models omit a real behavioral component, and the field needs a new consensus sentiment measure as the Baker-Wurgler index loses relevance.","feed_headline":"Adding sentiment to price models raises fit, 71-study review finds","feed_subtitle":"But the same evidence fails to show that complex sentiment measures beat simple ones at forecasting.","key_machinery":"The central object is the coefficient of determination (adjusted R-squared), used as a common yardstick to compare how much return variation sentiment explains across otherwise irreconcilable studies. The argument runs through a taxonomy that sorts papers into single-factor, medium-complex, multifactor, and machine-learning models, and sorts sentiment measures into simple proxies and complex composites such as the Baker-Wurgler index and media-based text measures. The load-bearing step is a frequency count: sentiment is reported significant in at least one relationship in 65 of 71 studies, which the author treats as confirmation that augmenting models with sentiment improves explanatory power. Mean adjusted R-squared comparisons and t-tests on those means are secondary; the author notes the R-squared differences between model classes are not statistically significant at conventional levels.","core_discovery":"The paper's central claim is that investor sentiment belongs inside asset pricing models: augmenting a single-factor, medium-complex, or multifactor model with a sentiment proxy raises the model's ability to explain returns, measured by the coefficient of determination. The author distinguishes simple sentiment measures (a single indicator such as a survey or Google search volume) from complex ones (composites of several indicators, or media/social-media based measures), and reports that the more complex measures and models show higher adjusted R-squared. The same evidence does not support a second hypothesis: there are only nine studies that compare sentiment measures directly, five favor composite indices and three do not, so the review concludes that complex measures cannot be shown to have better predictive power. The author also finds that the BW index, the field's most common sentiment proxy, is significant less often on recent data and at individual-stock level, and that sentiment's effect frequently reverses after a few periods.","pith_inferences":["Because the review selects papers by a minimum citation count, the 65-in-71 significance rate may overstate the true prevalence of sentiment effects if null results are less cited; a registry of all published tests would be needed to correct for this.","The finding that more complex models have higher R-squared could be mechanical: adjusted R-squared only partially penalizes extra parameters, and studies using richer models tend to use richer data; re-analysis with information criteria would test this.","A concrete extension is a formal meta-analysis of incremental adjusted R-squared from sentiment-augmented versus baseline models, controlling for data frequency, return horizon, and factor set; the review had too few incremental R-squared values to do this itself.","If BW index significance is decaying over time, sentiment effects may be migrating to new channels such as social media and retail trading platforms, so a time-varying sentiment measure could outperform any fixed composite; this follows from the review's temporal pattern but is not tested in the paper."],"forward_implications":["If sentiment proxies reliably raise adjusted R-squared, then rational factor models that omit sentiment are misspecified, and reported alphas in those models may partly be sentiment compensation.","The BW index should not remain the default measure: the review finds it less often significant in recent samples and for individual stocks, so future work needs updated composite or text-based measures.","Researchers should not yet build trading strategies on the claim that complex sentiment measures forecast better; the review finds the comparison evidence inconclusive.","Machine-learning sentiment models look promising for fit, but their predictive advantage over simple proxies remains unproven in the reviewed literature."],"supporting_citations":[{"why":"Supplies the benchmark composite sentiment index (BW) that most reviewed studies add to factor models.","marker":"Baker and Wurgler (2006)"},{"why":"Supplies the PLS-based composite measure that outperforms BW in out-of-sample tests, key for assessing RH2.","marker":"Huang et al. (2015)"},{"why":"Provides the noise-trader theory that motivates treating sentiment as a priced risk factor.","marker":"De Long et al. (1990)"},{"why":"Supplies the baseline three-factor rational model that reviewed studies augment with sentiment.","marker":"Fama and French (1992)"},{"why":"Supplies the baseline four-factor model, the other common benchmark to which sentiment is added.","marker":"Carhart (1997)"},{"why":"Provides the FEARS search-based sentiment measure, a representative complex alternative to the BW index.","marker":"Da et al. (2015)"},{"why":"Shows BW sentiment is significant across many anomalies, load-bearing evidence for the first hypothesis.","marker":"Stambaugh et al. (2012)"}],"fun_headline_variants":["Sentiment complexity improves R2 but not predictive power","71 studies: sentiment improves fit not forecast power","Complex sentiment measures raise fit, not forecast power","Sentiment's pricing impact varies by asset and time period","Review of 71 studies questions BW index as sentiment gauge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the 71 papers it selected from a citation-indexed database with at least 60 citations represent the wider sentiment-asset-pricing literature, and that the adjusted R-squared values reported by different studies can be compared even though those studies use different models, data frequencies, and return horizons.","fun_headline_variants_meta":{"raw":{"variants":["Sentiment complexity improves R2 but not predictive power","71 studies: sentiment improves fit not forecast power","Complex sentiment measures raise fit, not forecast power","Sentiment's pricing impact varies by asset and time period","Review of 71 studies questions BW index as sentiment gauge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001605,"raw_usage":{"total_tokens":6322,"prompt_tokens":806,"completion_tokens":5516,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":5440}},"tokens_in":422,"tokens_out":5516,"duration_ms":41782,"temperature":1.0,"reasoning_tokens":5440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:44:05.670720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the incremental adjusted R-squared obtained by adding a sentiment proxy to the same base factor model across a complete set of published studies and test whether the average gain is positive after penalizing the extra parameters. If that average is zero or negative, the central claim that sentiment proxies improve the coefficient of determination collapses.","supporting_citations":[{"cited_title":"F., Yu, J., & Yuan, Y","cited_arxiv_id":null,"evidence_quote":"Shows BW sentiment is significant across many anomalies, load-bearing evidence for the first hypothesis."}],"review_version":1}