Pith. sign in

REVIEW 4 major objections 5 minor 13 references

Investor Sentiment in Asset Pricing Models: A Review of Empirical Evidence

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This review of 71 empirical studies claims that adding investor sentiment proxies to asset pricing models improves the coefficient of determination, while evidence that more complex sentiment measures predict better than simple ones…

desk verdict A useful but overclaimed systematic review: the abstract asserts RH1 confirmed, while the paper's own tests show insignificant R-squared gains; revision can fix it. read the letter →

arxiv 2411.13180 v1 pith:YANBZEAR submitted 2024-11-20 q-fin.PM

classification q-fin.PM
keywords InvestorsentimentAssetpricingMultifactormodelsBehavioralfinanceCoefficientofdeterminationmeasuresBaker-WurglerindexLiteraturereview
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reviews 71 empirical studies published between 2000 and 2021 that add investor sentiment measures to asset pricing models. It claims that sentiment is significant in at least one tested relationship in 65 of the 71 studies, and that adding sentiment proxies improves the coefficient of determination of the models. It also claims that more complex sentiment measures and models are associated with higher adjusted R-squared. The paper does not claim that complex measures forecast better: only nine studies compared measures directly, and the evidence was too thin to confirm that hypothesis. If the review is right, standard rational factor models omit a real behavioral component, and the field needs a new consensus sentiment measure as the Baker-Wurgler index loses relevance.

What carries the argument

The central object is the coefficient of determination (adjusted R-squared), used as a common yardstick to compare how much return variation sentiment explains across otherwise irreconcilable studies. The argument runs through a taxonomy that sorts papers into single-factor, medium-complex, multifactor, and machine-learning models, and sorts sentiment measures into simple proxies and complex composites such as the Baker-Wurgler index and media-based text measures. The load-bearing step is a frequency count: sentiment is reported significant in at least one relationship in 65 of 71 studies, which the author treats as confirmation that augmenting models with sentiment improves explanatory power. Mean adjusted R-squared comparisons and t-tests on those means are secondary; the author notes the R-squared differences between model classes are not statistically significant at conventional levels.

What would settle it

Compute the incremental adjusted R-squared obtained by adding a sentiment proxy to the same base factor model across a complete set of published studies and test whether the average gain is positive after penalizing the extra parameters. If that average is zero or negative, the central claim that sentiment proxies improve the coefficient of determination collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that investor sentiment belongs inside asset pricing models: augmenting a single-factor, medium-complex, or multifactor model with a sentiment proxy raises the model's ability to explain returns, measured by the coefficient of determination. The author distinguishes simple sentiment measures (a single indicator such as a survey or Google search volume) from complex ones (composites of several indicators, or media/social-media based measures), and reports that the more complex measures and models show higher adjusted R-squared. The same evidence does not support a second hypothesis: there are only nine studies that compare sentiment measures directly, five favor composite indices and three do not, so the review concludes that complex measures cannot be shown to have better predictive power. The author also finds that the BW index, the field's most common sentiment proxy, is significant less often on recent data and at individual-stock level, and that sentiment's effect frequently reverses after a few periods.

Load-bearing premise

The review assumes that the 71 papers it selected from a citation-indexed database with at least 60 citations represent the wider sentiment-asset-pricing literature, and that the adjusted R-squared values reported by different studies can be compared even though those studies use different models, data frequencies, and return horizons.

Editorial extensions

If this is right

  • If sentiment proxies reliably raise adjusted R-squared, then rational factor models that omit sentiment are misspecified, and reported alphas in those models may partly be sentiment compensation.
  • The BW index should not remain the default measure: the review finds it less often significant in recent samples and for individual stocks, so future work needs updated composite or text-based measures.
  • Researchers should not yet build trading strategies on the claim that complex sentiment measures forecast better; the review finds the comparison evidence inconclusive.
  • Machine-learning sentiment models look promising for fit, but their predictive advantage over simple proxies remains unproven in the reviewed literature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the review selects papers by a minimum citation count, the 65-in-71 significance rate may overstate the true prevalence of sentiment effects if null results are less cited; a registry of all published tests would be needed to correct for this.
  • The finding that more complex models have higher R-squared could be mechanical: adjusted R-squared only partially penalizes extra parameters, and studies using richer models tend to use richer data; re-analysis with information criteria would test this.
  • A concrete extension is a formal meta-analysis of incremental adjusted R-squared from sentiment-augmented versus baseline models, controlling for data frequency, return horizon, and factor set; the review had too few incremental R-squared values to do this itself.
  • If BW index significance is decaying over time, sentiment effects may be migrating to new channels such as social media and retail trading platforms, so a time-varying sentiment measure could outperform any fixed composite; this follows from the review's temporal pattern but is not tested in the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reviews 71 empirical studies published between 2000 and 2021 that include investor sentiment in asset pricing models. It categorizes sentiment measures (direct, indirect, composite, media-based) and model classes (single-factor, medium-complex, multifactor, machine learning), summarizes qualitative findings on coefficient significance, and reports a quantitative comparison of mean adjusted R-squared across model classes. The abstract claims that 'higher complexity of sentiment measures and models improves the coefficient of determination,' while the second hypothesis about predictive power is judged unverifiable. The author concludes that the first hypothesis (RH1) is confirmed.

Significance. If the central claim were supported, the paper would be a useful synthesis of the sentiment-augmented asset pricing literature. Its main asset is the detailed appendix cataloguing 71 studies with asset, period, frequency, sentiment measure, model, sign, and significance, which could serve as a resource for future meta-analyses. However, the central claim is contradicted by the paper's own quantitative evidence: Section 4.2.1 reports t-tests on adjusted R-squared means with p-values above 10%, and the paper explicitly states that RH1 cannot be directly verified quantitatively due to the lack of incremental R-squared data. The qualitative frequency of significant sentiment coefficients (65/71 studies) is not evidence of an improvement in the coefficient of determination. The paper also lacks machine-checked proofs or reproducible code; its contribution is a review, so the internal inconsistency in the main conclusion is the decisive issue. With a corrected framing, the descriptive value of the survey could still justify publication.

major comments (4)
  1. [§4.1.7 and Abstract] The assertion that 'the obtained results confirm the first research hypothesis (RH1)' is not supported by the paper's own quantitative analysis in Section 4.2.1, where the t-tests comparing adjusted R-squared means (Table 9) are insignificant (p>0.10), and where the paper states that RH1 'cannot be directly verified quantitatively due to the lack of appropriate data (e.g. incremental R-squared).' The frequency of statistically significant sentiment coefficients (65/71 studies) does not imply an improvement in the coefficient of determination, and the abstract's stronger claim that 'higher complexity of sentiment measures and models improves the coefficient of determination' is not tested anywhere in the manuscript.
  2. [§4.2.1, Table 9] The comparison of mean adjusted R-squared across single-factor, medium-complex, and multifactor models does not test RH1, because the model classes differ in factor structure, data frequency, asset universe, and sample period, and no paper in Appendix A is shown to provide the incremental R-squared from adding sentiment to a baseline model. The paper itself cautions that 'one should be careful with interpreting those results,' but this caution is not carried into Section 4.1.7 or the conclusions, where RH1 is treated as confirmed.
  3. [§1 and Abstract] The abstract's claim that 'higher complexity of sentiment measures and models improves the coefficient of determination' conflates two distinct dimensions: the complexity of the sentiment measure and the complexity of the asset pricing model. The analysis provides no decomposition of R-squared gains between these two dimensions, and the direct comparative evidence on sentiment-measure complexity (Section 4.1.4, five of eight comparative studies favoring composite indices) concerns significance or predictive power, not the coefficient of determination.
  4. [§3.1] The inclusion criterion of a minimum of 60 citations on Web of Science, together with the exclusion of non-English papers and papers not reporting detailed model specifications, means that the 71-paper sample is not representative of the full sentiment-asset pricing literature. This limits the generalizability of prevalence claims such as 'sentiment is almost always an important factor' (Section 4.1.7) and should be stated as a first-order limitation in the conclusions, not only as a procedural choice in Section 3.1.
minor comments (5)
  1. [§5] The phrase 'The research failed to reject (both quantitatively and qualitatively) one out of the two hypotheses (RH1)' is ambiguous and seems to use 'failed to reject' in the opposite sense of the standard statistical meaning; the paper actually claims RH1 is confirmed, so the conclusion should state this directly.
  2. [§2.3.3] The citation 'Fang and Taylor (2021)' for the MGMT and PERF factors is missing from the bibliography.
  3. [Appendix A, Table 8] There are several spelling and notation errors in Table 8, including 'Amercian' (entries 6, 13, 24, 31, 36, 39), 'balance' for 'imbalance' (entry 9), and 'Sign' entries such as 'Impossible to detect' (entries 28, 43, 49, 56, 57, 62, 69) that are not signs and should be moved to a separate column or explained.
  4. [§2.1.1] The definitions in Eqs. (1)–(3) present sentiment as a price, return, or characteristic difference from a rational benchmark, but the text does not explain how this conceptual definition relates to the operational measures (e.g., BW index, survey measures, media-based indices) used in the reviewed studies.
  5. [§4.2.1] The text refers to 'medium-factor' models in the first sentence of Section 4.2.1, while elsewhere the paper uses 'medium complex' models; the terminology should be standardized.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this literature review transcribes external studies; its RH1 confirmation is questionable given the paper's own quantitative caveat, but no claim reduces to its own inputs by definition or by self-citation.

full rationale

This is a qualitative and quantitative literature review of 71 external papers. It fits no parameters, derives no predictive equations from its own hypotheses, and contains no self-citation chain that forces its conclusions. The central claim, RH1, is that augmenting models with sentiment proxies improves the coefficient of determination. The paper supports this in Section 4.1.7 by appeal to the frequency of significant sentiment coefficients (65 of 71 studies) and to reported coefficients and R-squared in the underlying papers, while Section 4.2.1 explicitly concedes: 'The first hypothesis (RH1) concerning the improvement of the coefficient of determination cannot be directly verified quantitatively due to the lack of appropriate data (e.g. incremental R-squared).' This is an internal evidentiary weakness, not circularity: the conclusion is not an input to the review, and it is not equivalent to any quantity the review itself constructed. Similarly, the abstract's claim that higher complexity improves R-squared is contradicted by Table 9's insignificant t-tests and by the paper's own caution, but that is a consistency and robustness problem rather than a self-referential derivation. Appendix A transcribes external studies' results, and the quantitative tables aggregate those reported values; no prediction is manufactured from a fit to the same data. Citations to Baker and Wurgler, De Long et al., and other authors are ordinary external literature and are not load-bearing self-citations, since the author does not rely on his own prior results. No equation in the paper (e.g., S_t = P_t - P*_t) is defined in terms of the hypotheses it is used to test. Therefore no specific circular step can be exhibited, and the appropriate honest finding is a score of 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No parameters are fitted and no entities are invented, but the review's conclusions depend on assumptions about sample representativeness, comparability of reported R-squared, and accuracy of the author's coding.

assumptions (3)
  • domain assumption The 60-citation threshold on Web of Science (as of January 2022) selects a representative sample of influential sentiment-asset pricing research.
    Section 3.1 states the minimum citation criterion without validating that it does not bias the sample toward older or more established measures.
  • domain assumption Reported adjusted R-squared values across studies are comparable despite differing models, data frequencies, and return horizons.
    Section 4.2.1 compares mean R-squared across model types using t-tests, implicitly assuming comparability.
  • domain assumption The author's subjective coding of each study's sentiment sign and significance is accurate.
    Section 3.2 describes qualitative analysis without coding validation or inter-rater reliability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Investor Sentiment in Asset Pricing Models: A Review of Empirical Evidence." pith.science (2026). https://pith.science/paper/YANBZEAR

@misc{pith2026241113180,
  author       = {Pith},
  title        = {Pith review of: Investor Sentiment in Asset Pricing Models: A Review of Empirical Evidence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YANBZEAR}},
  note         = {Machine review of arXiv:2411.13180}
}
read the original abstract

This study conducted a comprehensive review of 71 papers published between 2000 and 2021 that employed various measures of investor sentiment to model returns. The analysis indicates that higher complexity of sentiment measures and models improves the coefficient of determination. However, there was insufficient evidence to support that models incorporating more complex sentiment measures have better predictive power than those employing simpler proxies. Additionally, the significance of sentiment varies based on the asset and time period being analyzed, suggesting that the consensus relying on the BW index as a sentiment measure may be subject to change.

Figures

Figures reproduced from arXiv: 2411.13180 by the authors.

Figure 1
Figure 1. depicts the process of article selection for review [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 13 canonical work pages

  1. [1]

    the University of Michigan survey of consumer sentiment, 2. the Conference Board survey of consumer confidence Medium complex / Multifactor Negative Yes after 1977, and no before 42 # Author(s) Asset(s) Data period Frequency Investor sentiment measure(s) Model(s) Sign Significance 11 Cornelli, F., Goldreich, D., & Ljungqvist, A. (2006) Various countries -...

  2. [2]

    the ratio of odd-lot sales to purchases,

  3. [3]

    the net mutual fund redemption on stock returns Single-factor Positive 1, 3 – Yes 2 - No 4 Klibanoff, P., Lamont, O., & Wizman, T. A. (1998). Various countries January 1986 to March 1994 Weekly Based on media Medium complex - No 5 Simon, DP; Wiggins, RA (2001). American - S&P 500 January 1989 to June 1999 Daily 1. the VIX, 2. The put- call ratio, 3. the T...

  4. [4]

    S., Doran, J

    No 16 Banerjee, P. S., Doran, J. S., & Peterson, D. R. (2007) American - NYSE, AMEX and NASDAQ June 1986 to June 2005 Daily VIX Multifactor Positive Yes except for portfolios based on low beta, low book to market value and large size 17 Kurov, A. (2008). American - S&P 500 and Nasdaq-10 2002–2004 Daily

  5. [5]

    II sentiment index Medium complex Positive No for Bull market

    AAII sentiment index, 2. II sentiment index Medium complex Positive No for Bull market. Yes for bear market 18 Chang, S. C., Chen, S. S., Chou, R. K., & Lin, Y. H. (2008). American - NYSE 1994-2004 Intraday – hourly intervals Weather variables: wind speed, snowiness, raininess and temperature Multifactor - No 19 Fang, L., & Peress, J. (2009). American - N...

  6. [6]

    win dummy variable, 3

    goal difference, 2. win dummy variable, 3. loss dummy variable Single-factor positive for 1 and 2; negative for 3 Yes 22 Dorn, D. (2009). Germany - IPOs August 1999 to May 2000 Daily

  7. [7]

    gross day plus 1 purchases Medium complex Negative Yes 23 Kaplanski, G., & Levy, H

    gross when issued purchases, 2. gross day plus 1 purchases Medium complex Negative Yes 23 Kaplanski, G., & Levy, H. (2010). American - NYSE January 1950 to December 2007 Daily Aviation disasters Medium complex Negative Yes for the first day, no for the second day 24 Kurov, A. (2010). Amercian - S&P 500 January 1990 to November 2004 Daily Multiple Medium c...

  8. [8]

    Twitter GPOMS Machine learning Impossible to detect

    Twitter OpinionFinder, 2. Twitter GPOMS Machine learning Impossible to detect

Show all 13 references
  1. [9]

    F., Yu, J., & Yuan, Y

    Yes 29 Stambaugh, R. F., Yu, J., & Yuan, Y. (2012) American - NYSE, AMEX and NASDAQ from July 1965 to December 2007 Monthly BW Multifactor Positive Yes for 7 out of 11 anomalies 30 Moskowitz, T. J., Ooi, Y. H., & Pedersen, L. H. (2012) Various countries January 1965 to Decembe...

  2. [10]

    EU sentiment measure Multifactor Both depending on portfolio 1 – Yes, 2 - No 39 Xiong, G., & Bharadwaj, S

    BW, 2. EU sentiment measure Multifactor Both depending on portfolio 1 – Yes, 2 - No 39 Xiong, G., & Bharadwaj, S. (2013). American - all CRSP, Ken French’s website and Compustat November 2004 to February 2010 Monthly Based on media: positive and negative frequencies of news Mu...

  3. [11]

    of negative words in articles of SA, 2 freq

    freq. of negative words in articles of SA, 2 freq. of comments in SA, 3. the average fraction of negative words in DJNS Single-factor Negative 1,2 – Yes, 3 - No 42 Sprenger, T. O., Tumasjan, A., Sandner, P. G., & Welpe, I. M. (2014). Amercian - S&P 100 January 2010 to June 201...

  4. [12]

    The BW based on the PLS procedure Single-factor Negative 1 – No, 2 - Yes 52 Stambaugh, R

    the BW index, 2. The BW based on the PLS procedure Single-factor Negative 1 – No, 2 - Yes 52 Stambaugh, R. F., Yu, J., & Yuan, Y. (2015) American - NYSE, AMEX and NASDAQ August 1965 to January 2011 Monthly BW Multifactor Negative Yes 53 Goetzmann, W. N., Kim, D., Kumar, A., & ...

  5. [13]

    the Huang et al

    the BW index, 2. the Huang et al. (2015) investor sentiment index, 3. the CSI, 4. the Conference Board Consumer Confidence index, 5. FEARS Single-factor / Medium complex Negative Yes 71 Liu, H., Manzoor, A., Wang, C., Zhang, L., & Manzoor, Z. (2020). Various countries January ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.