REVIEW 3 major objections 4 minor 13 references
The Value of Information from Sell-side Analysts
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that the narrative text in sell-side analyst reports explains 10.19% of three-day abnormal stock returns out-of-sample, more than the 9.01% explained by quantitative forecast revisions, and that combining text with…
desk verdict The paper's headline text-vs-numbers comparison likely conflates samples, and the Shapley decomposition as written is infeasible; the robust out-of-sample text signal still deserves peer review after major fixes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is a three-stage pipeline: first, each report is mapped to a 5,120-dimensional embedding by averaging token and layer vectors from a large language model, with a sentence-segmented variant used to limit cross-topic contamination; second, out-of-sample ridge regressions of three-day cumulative abnormal returns on these embeddings produce $R^2_{\mathrm{oos}}$ values in an expanding-window design from 2000 through 2023; third, a Shapley value decomposition allocates the total out-of-sample $R^2$ across 17 predefined report topics by averaging each topic's marginal contribution over all subsets of topics, thereby attributing explanatory power while accounting for topic interactions.
What would settle it
Re-run the exact 131,072-subset Shapley computation on the published topic embeddings and check whether income-statement analysis contributes more than half of the total out-of-sample $R^2$; if it does not, the paper's central topic-importance claim fails.
Extended reading notes
Core claim
The central discovery is that the qualitative content of analyst reports, represented by LLM embeddings, has greater out-of-sample explanatory power for contemporaneous stock returns than the quantitative forecasts in the same reports. Text alone gives an out-of-sample $R^2$ of 10.19% for three-day cumulative abnormal returns, compared with 9.01% for forecast revisions and 9.08% for a broader set of 17 numerical measures; combining text and revisions raises the $R^2$ to 12.28%. This result survives removing all numbers from the text, which actually raises the text-only $R^2$ to 10.95%, and it is not explained by simple sentiment measures, which reach at most about 3.7%. A Shapley value decomposition attributes roughly two-thirds of the text's explanatory power to income-statement analysis, and within that topic, interpretation of realized income contributes about three times as much as reporting raw data. Translating the explained return variance into dollar terms through an imperfect-competition trading model, the paper estimates that early acquisition of analyst reports yields $0.38 million per event from text, $0.34 million from revisions, and $0.47 million combined.
Load-bearing premise
The topic-importance result rests on the unstated feasibility of the Shapley computation: the paper reports exact-looking contributions from all 131,072 subsets of 17 topics without describing an approximation, so the whole decomposition depends on an undisclosed computational step that must reproduce those numbers.
Editorial extensions
If this is right
- If analyst narratives carry more value-relevant information than forecast numbers, then consensus estimates and recommendation ratings are incomplete summaries of analyst output, and investors who read only the numbers miss a distinct, somewhat larger source of return predictability.
- The 12.28% combined $R^2$ being significantly higher than either input alone implies that quantitative forecasts and qualitative text are complements rather than substitutes.
- Because the text-only signal survives when all digits are stripped from the reports, the market is reacting to narrative framing and reasoning, not merely to the numbers embedded in the prose.
- The sharp peak in information content in the week after earnings announcements means the timing of report dissemination is economically important, not just its content.
- The dollar-value estimates imply that selective early distribution of analyst reports is worth millions of dollars per large-cap stock each year, providing a concrete measure of the incentive behind tipping practices.
Reading between the lines
- The paper does not build a tradeable strategy; a natural extension is to construct a long-short portfolio sorted on the text-only predicted abnormal return and measure net returns after transaction costs, testing whether the $0.47 million gross value survives trading frictions.
- Because the sample is restricted to S&P 100 stocks, the size-liquidity channel found in the paper may not generalize; for smaller, less liquid firms the information channel could dominate and make analyst text relatively more valuable per dollar traded.
- The topic hierarchy is defined relative to one set of 17 ChatGPT-generated categories; a different topic taxonomy might redistribute credit, so the 'income statement is half the value' conclusion should be tested with an independently constructed labeling scheme.
- A controlled experiment that presents investors with number-only summaries versus full interpretive narratives for the same earnings outcome would directly test the paper's claim that interpretation, not information acquisition, drives the market reaction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the information content of sell-side analyst reports by combining LLM-based text embeddings with out-of-sample Ridge regressions. Using 122,252 Investext reports for S&P 100 stocks from 2000 to 2023, it reports that text embeddings explain 10.19% of three-day cumulative abnormal returns out-of-sample, versus 9.01% for quantitative forecast revisions, and that a combined model reaches 12.28%. The paper also introduces a 17-topic taxonomy (generated with ChatGPT-4o and a fine-tuned BERT classifier) and uses a Shapley value decomposition (Eq. 3) to attribute the text-based R2_OOS across topics, finding that Income Statement Analysis dominates. Economic significance is assessed by translating the predictive R2 into dollar values through the Kadan and Manela (2025) framework, yielding an estimated annualized information value of about $6.89 million for an average S&P 100 stock. The main contribution is the claim that qualitative narrative in analyst reports carries information beyond the quantitative forecasts, with the decomposition and dollar-value estimates as secondary contributions.
Significance. If the central claim holds, the paper provides novel, large-sample evidence on the value of qualitative analyst text, complementing the established literature on quantitative revisions and sentiment. The empirical design has notable strengths: the out-of-sample evaluation uses an expanding-window scheme with a 2015-2023 test period; the main result is robust across several LLMs (BERT, OpenAI, LLaMA-3), multiple ML algorithms (PLS, XGBoost, neural networks), alternative CAR windows, and a 'numbers removed' analysis; and the 2023-only test addresses look-ahead leakage in a credible way. The economic translation through an external theoretical framework adds a practical dimension that is rare in this literature. The topic-decomposition and dollar-value results, however, rest on computational and sample-comparability assumptions that are not fully disclosed, as detailed in the major comments. With those points resolved, the paper would be a valuable contribution to the analyst-information and AI-in-finance literatures.
major comments (3)
- [Section 3.2.1, Table 3 and Table A3] The headline comparison between 'Text only' (R2_OOS = 10.19%) and 'Revision only' (R2_OOS = 9.01%) may not be evaluated on a common sample. Table A3 shows that EFREV has N=90,625 and TPREV has N=84,108, while the full report sample is N=122,252 and RECREV has N=120,673. The text-only model requires no forecast-revision fields, while the revision-only and 'Rev + text' models will drop any report missing RECREV, EFREV, or TPREV unless explicit imputation is performed. Table 3 does not report subsample sizes, and the text does not describe imputation for missing revisions. If the revision-based models are estimated on the roughly 84k-report subsample while text-only uses the full 122k sample, the 1.18 percentage-point gap and the 2.09pp incremental gain from adding text to revisions are not identified, because they conflate information content with sample composition. Please report all R2_OOS values and DM tests on a common set of observations, or clarify the imputation procedure.
- [Section 2.3, Eq. (3) and Section 3.2.3, Figure 2] The Shapley value decomposition as written in Eq. (3) requires computing out-of-sample R2 for every subset of the 17 topics, i.e., 2^17 = 131,072 Ridge regressions on 5,120-dimensional embeddings, each with cross-validated regularization over an expanding window. The paper reports exact-looking Shapley values in Figure 2 and claims that Income Statement Analysis accounts for 67% of total R2_OOS, yet it does not describe any approximation, such as Monte Carlo Shapley, truncation, or a linearity assumption. Without a feasible and disclosed approximation, the numerical topic contributions and the central decomposition result in Section 3.2.3 are not verifiable. Please describe the actual computational procedure, the approximation error bounds, or provide code/reproducibility details.
- [Abstract and Section 3.2.1] The abstract states that qualitative information is 'more economically significant than quantitative forecasts,' but the direct DM test comparing 'Text only' and 'Revision only' yields a t-statistic of 1.66 (Table 3, row 'Overall', column '(5)-(1)'), which is not statistically significant at conventional levels. While the body text appropriately says 'comparable or marginally larger,' the abstract overstates the evidence. Please temper the abstract claim or provide a same-sample test that would support a stronger statement.
minor comments (4)
- [Section 2.2] The word 'expriment' in the sentence 'I expriment a standard K-Means algorithm' should be 'experiment'.
- [Table A8 and Figure A8] The description of the 'CLNV' series in Figure A8 says it uses the trade-signing algorithm of Lee and Ready (1991), but the main analysis already uses Lee and Ready, and the table labels 'CLNV' as the algorithm of Chakrabarty et al. (2007). Please clarify which method corresponds to which series, and fix the spelling of 'volatility' in Table A8.
- [Section 3.2.3] The Shapley values are described as 'relative contribution' to total R2_OOS, but with 17 topics some Shapley values are negative (e.g., 'minimal or even negative'). Please state explicitly how negative Shapley values are normalized or whether the reported percentages are computed over the sum of positive values only.
- [Section 3.2.5, Table 7] The label 'ToneIncome,NB/BERT' in the table header is ambiguous about the two separate measures; consider writing 'ToneIncome,NB' and 'ToneIncome,BERT' explicitly to match the reporting in the text and panel notes.
Circularity Check
No significant circularity: the predictive and decomposition claims are genuine out-of-sample results, not reduced to their inputs by construction.
full rationale
The paper's main derivation chain is self-contained. Analyst report text is mapped to LLaMA-2 embeddings (Eq. 1-2), and three-day CARs are obtained from market data; the Ridge regressions are trained on expanding pre-2015 windows and evaluated out-of-sample on 2015-2023 data (Eq. 4-6). The R2_oos values in Table 3 are genuine forecasts against a zero benchmark: the text embeddings do not encode the target return, and no fitted parameter is renamed as a prediction. The Shapley decomposition (Eq. 3) is defined as the exact additive attribution of the fitted model's own out-of-sample R2; this is a mathematical property of the estimated model, not a circular reduction, because topic subsets are distinct input constructions evaluated by re-estimation. The dollar-value estimates in Section 3.3 are an external Kadan-Manela transformation of the same R2 into profit terms; this is an interpretive mapping, not an assumption that already contains the conclusion. Cited prior work (Gu et al. 2020; Huang et al. 2014; Kadan and Manela 2025; Li et al. 2024) is external to the author and is used for methodology rather than to import the central result. The reviewer concerns about unequal samples across revision and text models and the computational feasibility of 2^17 Shapley evaluations are substantive empirical and reproducibility risks, but they concern identification and implementation, not circularity. No load-bearing step reduces, by the paper's own equations or by self-citation, to its own inputs.
Assumptions & free parameters
free parameters (1)
- Ridge regularization parameter α =
Cross-validated over [10^-10, 10^10]
assumptions (4)
- domain assumption Contemporaneous abnormal returns around the report release date reflect the information content of the report.
- domain assumption The Kadan-Manela (2025) framework provides a valid mapping from explained return variance and price impact to the dollar value of information.
- ad hoc to paper The 17-topic taxonomy generated by ChatGPT-4o and the fine-tuned BERT classifier (89% accuracy) are valid partitions of analyst report content.
- domain assumption Sentence-segmented embeddings eliminate cross-topic contamination in the Shapley decomposition.
Cite this review
Pith. "Pith review of The Value of Information from Sell-side Analysts." pith.science (2026). https://pith.science/paper/OTP6GIZN
@misc{pith2026241113813,
author = {Pith},
title = {Pith review of: The Value of Information from Sell-side Analysts},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTP6GIZN}},
note = {Machine review of arXiv:2411.13813}
}
read the original abstract
I examine the value of information from sell-side analysts by analyzing a large corpus of their written reports. Using embeddings from state-of-the-art large language models, I show that qualitative information in analyst reports explains above 10% of contemporaneous stock returns out-of-sample, a value that is more economically significant than quantitative forecasts. I then perform a Shapley value decomposition to assess how much each topic within the reports contributes to explaining stock returns. The results show that analysts' income statement analyses account for more than half of the reports' explanatory power. Expressing these findings in economic terms, I estimate that early acquisition of analyst reports can yield significant profits. Analyst information value peaks in the first week following earnings announcements, highlighting their vital role in interpreting new financial data.
Reference graph
Works this paper leans on
-
[3]
Elective surgery trends exiting Q1, expectations for 2021 and an update on recent and upcoming new product approvals; and
work page 2021
-
[4]
Panel A provides annual statistics, including the number of reports, unique brokerage firms and analysts, and the average report length in pages and tokens. Panel B presents the same statistics aggregated by Fama-French 12 (FF12) industry. FF12 industry definitions are available on Kenneth French’s data library. Panel A: Sell-side analyst reports by year ...
work page 2020
-
[5]
The shaded area denotes the 2020 pandemic recession
Specifically, the red line isolates the contribution of the ‘Income Statement Analyses’ topic, as measured by its annual Shapley value. The shaded area denotes the 2020 pandemic recession. 10 Figure A7 Topic Importance via Shapley Value by Analyst Characteristics This figure illustrates how 17 different topics contribute to the analyst report information ...
work page 2000
-
[7]
Global trends and impact on EPD and Nutrition. Company Overview • Oracle Corporation, founded in 1977 and headquartered in Redwood Shores, California, is one of the largest and most prominent companies in the software space – and a technology bellwether. • As it has grown, Microsoft has expanded into enterprise software with Windows Server, SQL Server, Dy...
work page 1977
-
[10]
legislative action. • In addition to the expenses incurred by patent challenges, product liability and other legal suits could occur and lead to additional liabilities and revenue loss, which could substantially change our financial assumptions. Management and Governance • Top management changes can be unsettling, and the resulting uncertainty has caused ...
work page 1974
-
[12]
The Overall row reports theR2 OOS and t-statistics for 2015-2023
Panel A shows results for the full sample, while Panels B and C restrict the sample to reports released within one day (excluding same day) of an earnings announcement and those outside this window, respectively. The Overall row reports theR2 OOS and t-statistics for 2015-2023. The t-statistics for R2 OOS are calculated using zero benchmarks estimation fo...
work page 2020
-
[13]
The cross-validation process ensures that the chosen model generalizes well to unseen data, preventing overfitting while capturing the predictive power of the text embeddings. Partial Least Square Regression To mitigate the risk of overfitting inherent in high-dimensional text embeddings, I employ Partial Least Squares (PLS) for dimensionality reduction. ...
work page 2020
-
[1984]
Journal of Financial and Quantitative Analysis 19, 299–310
Firm size and the informational content of financial statements. Journal of Financial and Quantitative Analysis 19, 299–310. 35 Figure 1 Distribution of Topics over Time This figure shows the distribution of report sentences across 17 topics from 2000Q1 to 2023Q4. The stacked plot illustrates the proportional composition of sentences over time. The topic ...
Show all 13 references
-
[2000]
T-statistics for the R2 OOS are calculated against a zero benchmark following the procedure in Gu et al
The ‘Overall’ row summarizes the performance across the full 2015–2023 sample period. T-statistics for the R2 OOS are calculated against a zero benchmark following the procedure in Gu et al. (2020). PLS XGBoost NN1 NN2 NN3 NN4 NN5 year R2 OOS t-stat R2 OOS t-stat R2 OOS t-stat...
2020
-
[2002]
• Despite record prices, oil demand continues to grow, while supply growth lags and spare production and refining capacity is almost nonexistent
Industry Analysis • According to our global Immunology market model, US Psoriasis (PsO) represented a $7.7B market in 2016 and is expected to grow at a low-teens CAGR to $11.8B in 2019E and $13B by 2021E driven by more highly effective therapies. • Despite record prices, oil d...
2016
-
[2003]
The Journal of Finance 58, 1933–1967
An empirical analysis of analysts’ target prices: Short-term informativeness and long-term dynamics. The Journal of Finance 58, 1933–1967. Cao, S., Jiang, W., Wang, J., Yang, B.,
1933
-
[2010]
• Further, competition in the CDK-4/6 space is rising with Verzenio (abemaciclib) & Kisqali launches placing downward pressure on Ibrance trajectory
Competitive Landscape • According to Mercury Research, NVIDIA is now the 3rd largest chipset supplier (consisting of desktop and mobile chipsets, and integrated and non-integrated chipsets), shipping 5.4 million units in calendar Q3 for an 8.2% market share, versus Intel’s shi...
2010
-
[2023]
arXiv preprint arXiv:2304.07619
Can chatgpt forecast stock price movements? return predictability and large language models. arXiv preprint arXiv:2304.07619 . Lundberg, S. M., Lee, S.-I.,
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.