REVIEW 3 major objections 3 minor 36 references
Nowcasting the euro area with social media data
T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Reddit comment-voting signals, scored by a large language model, improve out-of-sample nowcasts of euro-area inflation and unemployment.
desk verdict A well-built Reddit-based nowcasting pipeline whose headline gains are inflated by selecting the best of 120 specifications on the evaluation sample; the economic signal looks real, but the magnitudes need a holdout or multiple-testing correction before they can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The device that carries the argument is the comment-weighted vote score $L_i = (S_i + \sum_{j=1}^{J} C_{i,j})/(J+1)$, where $S_i$ is the LLM's UP, DOWN, or NEUTRAL label for submission $i$ and $C_{i,j}$ are the LLM labels of its comments, optionally weighted by each comment's upvote-minus-downvote net score. A threshold $\tau$ maps $L_i$ back to a revised label $\bar{S}_i$: UP if $L_i>\tau$, DOWN if $L_i<-\tau$, NEUTRAL otherwise, and daily signals $\bar{X}_t = \sum_i \bar{S}_i$ are smoothed with backward-looking moving averages. The paper varies the comment set, the upvote weighting, the threshold, and the smoothing window to produce 120 candidate series, and evaluates each in a mixed-frequency MIDAS-AR regression with Almon-polynomial weights against a monthly AR(1) benchmark.
What would settle it
Run the specification search only on data through 2020 and then evaluate the chosen Reddit indicator on 2021-2023; if it no longer beats the AR(1) benchmark and the newspaper or swap indicators, the claimed out-of-sample gains are an artifact of in-sample selection.
Extended reading notes
Core claim
The central claim is that a comment-voting scheme applied to LLM-classified Reddit posts yields daily indicators that beat daily newspaper sentiment, inflation swaps, and oil prices in nowcasting eight euro-area series: overall HICP and its core, energy, food, and services components, plus total, over-25, and under-25 unemployment. The paper reports relative RMSFE improvements over the AR(1) benchmark ranging up to 13 percentage points for food price inflation and at least 5 percentage points for youth unemployment. The authors argue that the LLM's forward-looking, context-sensitive classification of informal Reddit text is what makes the raw signal work, and that letting comment votes revise a submission's original UP, DOWN, or NEUTRAL label regularizes the signal, reclassifying about 11 percent of inflation posts and 15 percent of unemployment posts.
Load-bearing premise
The reported gains assume that picking the best of 120 Reddit indicator designs on the same 2018-2023 period used to score them did not overfit that period.
Editorial extensions
If this is right
- Reddit-derived daily signals can be added to the nowcaster's toolkit for euro-area prices and labor markets, with the largest reported gains for food price inflation and meaningful gains for youth unemployment.
- Including comments rather than submissions alone improves nowcasting accuracy: the winning specification for every target variable uses the comment-voting scheme.
- The gains are concentrated in unusual periods, namely the COVID-19 recession and the high-inflation episode of 2021-2023, consistent with high-frequency information mattering most during turmoil.
- The LLM's classification of informal Reddit text reaches F1 scores around 0.71 to 0.75, well above dictionary-based baselines, so the signal extraction is reproducible with a general-purpose large language model.
- The information value persists when the daily information cutoff is moved earlier in the month, with gains clear up to about 14 days before the end of the nowcast period.
Reading between the lines
- (Editorial inference) If the gains survive a genuine holdout, the same pipeline should transfer to other European subreddits and to targets such as GDP or housing, giving country-level daily indicators; the paper itself only tests the euro area as a whole through r/europe.
- (Editorial inference) The comment-voting scheme acts like a cheap ensemble regularizer, but comment votes are not independent observations because later commenters have already read earlier comments; weighting comments by depth, time lag, or author diversity could either strengthen the signal or reveal that a few active users drive it.
- (Editorial inference) A direct testable extension would compare the Reddit signal against demographic-specific survey expectations, since the paper's larger gains for food inflation and youth unemployment fit the idea that personally experienced prices shape expectations more than abstract index numbers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper constructs daily Reddit-based indicators for euro-area inflation and unemployment by using LLaMa-3-70B to classify submissions and comments into UP/DOWN/NEUTRAL signals, then applies combinations of moving-average smoothing, thresholded comment voting, optional upvote/downvote weighting, and two comment sets, yielding 120 indicator specifications per target. These are evaluated one at a time as daily predictors in MIDAS-AR nowcasting regressions against a monthly AR(1) benchmark, newspaper sentiment indices, inflation/output swaps, and oil prices over January 2018 to December 2023. The authors report consistent RMSFE, MAFE, and CRPS gains, with improvements of up to 13 percentage points for food price inflation and at least 5 percentage points for youth unemployment, and they attribute the gains primarily to the social-interaction layer that reclassifies submissions using comment signals.
Significance. If the reported gains are genuine, the paper makes a useful contribution: it introduces a novel high-frequency, LLM-processed social media information set for euro-area nowcasting, and the social-interaction layer is an interesting and transferable idea. The authors are transparent about the LLM prompt, report F1 accuracy against human labels, and use standard Diebold-Mariano/Harvey and Giacomini-Rossi tests. However, the headline empirical claim is not yet identified because the best Reddit specification is selected on the same out-of-sample period used for evaluation, so the reported RMSFE improvements are minima over 120 candidates rather than the performance of a pre-specified model. This selection-overfitting concern is load-bearing for the paper's central claim and must be addressed before the empirical conclusions can be accepted.
major comments (3)
- [Section 2.4 and Table 3.2] The 'winning' Reddit specification is chosen by RMSFE computed on the January 2018 to December 2023 out-of-sample period, which is the same period on which the horse-race performance is reported. This makes the headline ratios (e.g., 0.679 for food, 0.723 for HICP) the minimum over 120 highly correlated statistics rather than the performance of a pre-specified rule; under the null of no predictive content, the minimum of 120 exchangeable statistics lies far below the median, so the reported gains are inflated. The heterogeneity of the winning configurations across targets (com 60/0.3 for HICP, com 365/0.1 for core, com 365/0.7 for food, com 90/0.3 for all unemployment rows) is more consistent with selection noise than with a stable economic signal. The F1 validation in Section 2.3 supports classification quality but does not establish that the selected indicator's time-series gains are free of selection overfitting. Please address this by reporting the full distribution of RMSFE over all 120 specifications, applying a multiple-testing correction such as White's reality check or a model confidence set, and/or selecting specifications on a pre-2018 subsample or a pre-registered rule and evaluating only those on 2018-2023.
- [Table 3.2, unemployment rows] The RMSFE gains for unemployment are not statistically significant against the AR(1) benchmark. For example, for the unemployment rate under 25, the RMSFE ratio is 0.906 with no asterisk and the CRPS ratio is 0.909 with no asterisk; only the MAFE ratio (0.872) is significant at the 5% level. The abstract and conclusion nevertheless claim 'consistent gains' and a minimum gain of 5 percentage points for youth unemployment, where the 5 percentage points is an RMSFE comparison against the best sentiment indicator, not a significant improvement over the benchmark. The wording should be qualified, or the evaluation should be extended (for example by pooling across targets or using a test with higher power), before claiming consistent unemployment gains.
- [Section 3.1 and Table 3.2] The comparison is asymmetric in candidate set size. Best Reddit is selected from 120 specifications, whereas 'Best Sentiment' is the best of 3 newspaper indices and 'Best Swap' is the best of 5 swap maturities. Even if the selection were performed honestly, comparing the best of 120 Reddit series against the best of a much smaller set tilts the comparison in Reddit's favor. The paper should either compare a fixed Reddit specification (or the distribution of all Reddit specifications) against the full set of competitors, or at least report how many of the 120 Reddit specifications beat the best sentiment and best swap series.
minor comments (3)
- [Section 2.1] The text contains a typo: 'Redddit' should be 'Reddit'.
- [Section 3.1] The sentence listing the availability of newspaper sentiment indices says 'Germany, France, Italy, and Germany'; the final country should presumably be Spain.
- [Table 3.2 and Figures 3.6, B.11] The shorthand in Table 3.2 (e.g., 'com 60 0.3 0 1' interpreted as 'firstlevel') is not clearly mapped to the figure legends, which use 'filter' and 'nofilter' labels such as 'comments_60_0.3_noscore_filter' and 'comments_365_0.1_noscore_nofilter'; please state explicitly which comment set corresponds to each label.
Circularity Check
Best-of-120 Reddit specification is selected on the same 2018-2023 out-of-sample window used to report nowcasting gains, so the headline improvements are partly a selection artifact.
-
fitted input called prediction
[Section 2.4 (indicator construction); Table 3.2 note; Section 3.1 (evaluation sample)]
"All of these choices are evaluated out-of-sample and results are reported in Appendix B.6. At the end we have a total of 120 series for each of the two target variables, inflation and unemployment. [...] “Best” indicators in each category are chosen on the basis of RMSFE."
The 120 Reddit series are generated as 2 comment sets x 2 scoring rules x 5 thresholds x 6 MA windows. Table 3.2 then picks the single best Reddit specification per target by RMSFE computed over January 2018 to December 2023, which is exactly the out-of-sample period on which the horse-race and the headline gains (13 percentage points for food, 5 for youth unemployment) are evaluated. Reporting the minimum of 120 exchangeable RMSFE ratios as the Reddit result means that even a useless indicator family would show apparent improvement over the AR(1); no holdout, nested validation, or multiple-testing correction is applied. The 'out-of-sample prediction' is therefore the fitted maximum of a search over the evaluation sample, not the performance of a pre-specified indicator.
full rationale
The central data-processing steps are not circular: the LLM classification is validated on human-labeled submissions (F1 around 0.71-0.75, Section 2.3), which is an external benchmark, and the MIDAS-AR nowcasting equations are standard. The main circularity is the selection protocol: Section 2.4 builds 120 indicator variants and says all choices are 'evaluated out-of-sample'; Table 3.2 chooses the best variant per target by RMSFE on the same 2018-2023 window that defines the out-of-sample evaluation. The reported improvements are the selected best, not the expected performance of a fixed Reddit signal. The heterogeneity of winning configurations across targets (com 60/0.3 for HICP, com 365/0.7 for food, com 90/0.3 for all unemployment targets) is consistent with this selection noise. The cited Barbaglia et al. (2023) self-citation is contextual and not load-bearing. Overall the paper contains real independent content, but the headline nowcasting gains are partially an artifact of fitting the indicator specification to the evaluation sample, warranting a score of 6.
Assumptions & free parameters
free parameters (5)
- MA smoothing window length =
60 or 365 days for inflation targets, 90 days for unemployment targets
- Threshold tau for comment-vote reclassification =
0.3 for most targets, 0.7 for food, 0.1 for core and services
- Upvote/downvote weighting flag =
0 (no scoring) for all winners in Table 3.2
- Comment set (first-level vs keyword-filtered) =
First level for inflation, keyword-filtered for unemployment
- LLM temperature =
0.5
assumptions (5)
- domain assumption Reddit users' forward-looking statements in r/europe contain information about future euro area inflation and unemployment.
- domain assumption The LLM's three-way classification of submissions and comments is accurate enough, and the labels on comments (never human-validated) are as reliable as labels on submissions.
- domain assumption The English-language subreddit r/europe is a valid proxy for euro area sentiment.
- domain assumption The MIDAS-AR(1) with unrevised data and a fixed lag structure is a fair environment for comparing information sets.
- domain assumption No material look-ahead bias from comment timing.
Cite this review
Pith. "Pith review of Nowcasting the euro area with social media data." pith.science (2026). https://pith.science/paper/QYONZCL2
@misc{pith2026250610546,
author = {Pith},
title = {Pith review of: Nowcasting the euro area with social media data},
year = {2026},
howpublished = {\url{https://pith.science/paper/QYONZCL2}},
note = {Machine review of arXiv:2506.10546}
}
read the original abstract
Using a state-of-the-art large language model, we extract forward-looking and context-sensitive signals related to inflation and unemployment in the euro area from millions of Reddit submissions and comments. We develop daily indicators that incorporate, in addition to posts, the social interaction among users. Our empirical results show consistent gains in out-of-sample nowcasting accuracy relative to daily newspaper sentiment and financial variables, especially in unusual times such as the (post-)COVID-19 period. We conclude that the application of AI tools to the analysis of social media, specifically Reddit, provides useful signals about inflation and unemployment in Europe at daily frequency and constitutes a useful addition to the toolkit available to economic forecasters and nowcasters.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Aliaj, T., Ciganovic, M., and Tancioni, M. (2023). Nowcasting inflation with lasso-regularized vector autoregressions and mixed frequency data. Journal of Forecasting , 42(3):464--480
work page 2023
-
[2]
Angelico, C., Marcucci, J., Miccoli, M., and Quarta, F. (2022). Can we measure inflation expectations using twitter? Journal of Econometrics , 228(2):259--277
work page 2022
-
[3]
Aprigliano, V., Emiliozzi, S., Guaitoli, G., Luciani, A., Marcucci, J., and Monteforte, L. (2023). The power of text-based indicators in forecasting italian economic activity. International Journal of Forecasting , 39(2):791--808
work page 2023
-
[4]
Ashwin, J., Kalamara, E., and Saiz, L. (2024). Nowcasting euro area gdp with news sentiment: a tale of two crises. Journal of Applied Econometrics , 39(5):887--905
work page 2024
-
[5]
Babii, A., Ghysels, E., and Striaukas, J. (2022). Machine learning time series regressions with an application to nowcasting. Journal of Business & Economic Statistics , 40(3):1094--1106
work page 2022
-
[6]
Baker, S. R., Bloom, N., and Davis, S. J. (2016). Measuring economic policy uncertainty. The quarterly journal of economics , 131(4):1593--1636
work page 2016
-
[7]
Ba \'n bura, M., Belousova, I., Bodn \'a r, K., and T \'o th, M. B. (2023). Nowcasting employment in the euro area . Number 2815. ECB Working Paper
work page 2023
-
[8]
Ba \'n bura, M., Giannone, D., Modugno, M., and Reichlin, L. (2013). Now-casting and the real-time data flow. In Handbook of economic forecasting , volume 2, pages 195--237. Elsevier
work page 2013
Show all 36 references
-
[9]
Barbaglia, L., Consoli, S., and Manzan, S. (2024). Forecasting gdp in europe with textual data. Journal of Applied Econometrics , 39(2):338--355
2024
-
[10]
M., Ratto, M., and Pezzoli, L
Barbaglia, L., Frattarolo, L., Onorante, L., Pericoli, F. M., Ratto, M., and Pezzoli, L. T. (2023). Testing big data in a big crisis: Nowcasting under covid-19. International Journal of Forecasting , 39(4):1548--1563
2023
-
[11]
and Lissona, C
Barigozzi, M. and Lissona, C. (2024). Ea-md-qd: Large euro area and euro member countries datasets for macroeconomic research
2024
-
[12]
Barsky, R. B. and Sims, E. R. (2012). Information, animal spirits, and the meaning of innovations in consumer confidence. American Economic Review , 102(4):1343--1377
2012
-
[13]
W., Carstensen, K., Menz, J.-O., Schnorrenberger, R., and Wieland, E
Beck, G. W., Carstensen, K., Menz, J.-O., Schnorrenberger, R., and Wieland, E. (2024). Nowcasting consumer price inflation using high-frequency scanner data: Evidence from germany. ECB Working Paper
2024
-
[14]
D., and Osiewicz, M
Benatti, N., Botelho, V., Consolo, A., Da Silva, A. D., and Osiewicz, M. (2020). High-frequency data developments in the euro area labour market. Economic Bulletin Boxes , 5
2020
-
[15]
and Roling, C
Breitung, J. and Roling, C. (2015). Forecasting inflation rates using daily data: A nonparametric midas approach. Journal of Forecasting , 34(7):588--603
2015
-
[16]
Bybee, L. (2023). Surveying generative ai's economic expectations. arXiv preprint arXiv:2305.02823
2023 arXiv
-
[17]
Carriero, A., Pettenuzzo, D., and Shekhar, S. (2024). Macroeconomic forecasting with large language models. arXiv preprint arXiv:2407.00890
2024
-
[18]
R., Giannone, D., and Modugno, M
Cascaldi-Garcia, D., Ferreira, T. R., Giannone, D., and Modugno, M. (2024). Back to the present: Learning about the euro area through a now-casting model. International Journal of Forecasting , 40(2):661--686
2024
-
[19]
Clark, T. E. and Ravazzolo, F. (2015). Macroeconomic forecasting performance under alternative specifications of time-varying volatility. Journal of Applied Econometrics , 30(4):551--575
2015
-
[20]
Consoli, S., Barbaglia, L., and Manzan, S. (2022). Fine-grained, aspect-based sentiment analysis on economic and financial lexicon. Knowledge-Based Systems , 247:108781
2022
-
[21]
D’Acunto, F., Malmendier, U., Ospina, J., and Weber, M. (2021). Exposure to grocery prices and inflation expectations. Journal of Political Economy , 129(5):1615--1639
2021
-
[22]
and Leibovici, F
Faria-e Castro, M. and Leibovici, F. (2024). Artificial intelligence and inflation forecasts. Technical report
2024
-
[23]
and Simoni, A
Ferrara, L. and Simoni, A. (2023). When are google data useful to nowcast gdp? an approach via preselection and shrinkage. Journal of Business & Economic Statistics , 41(4):1188--1202
2023
-
[24]
Ghysels, E., Kvedaras, V., and Zemlys, V. (2016). Mixed frequency data sampling regression models: The r package midasr. Journal of statistical software , 72:1--35
2016
-
[25]
and Rossi, B
Giacomini, R. and Rossi, B. (2010). Forecast comparisons in unstable environments. Journal of Applied Econometrics , 25(4):595--620
2010
-
[26]
Giannone, D., Reichlin, L., and Simonelli, S. (2009). Nowcasting euro area economic activity in real time: the role of confidence indicators. National Institute Economic Review , 210:90--97
2009
-
[27]
Goyal, S., Rosenkranz, S., Weitzel, U., and Buskens, V. (2017). Information acquisition and exchange in social networks. The economic journal , 127(606):2302--2331
2017
-
[28]
H., Meggiorini, G., and Melosi, L
Granziera, E., Larsen, V. H., Meggiorini, G., and Melosi, L. (2025). Speaking of inflation: the influence of fed speeches on expectations
2025
-
[29]
and Yang, L
Han, B. and Yang, L. (2013). Social networks, information acquisition, and asset prices. Management Science , 59(6):1444--1457
2013
-
[30]
L., Horton, J
Hansen, A. L., Horton, J. J., Kazinnik, S., Puzzello, D., and Zarifhonarvar, A. (2024). Simulating the survey of professional forecasters. Available at SSRN
2024
-
[31]
Harvey, D., Leybourne, S., and Newbold, P. (1997). Testing the equality of prediction mean squared errors. International Journal of forecasting , 13(2):281--291
1997
-
[32]
Horton, J. J. (2023). Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research
2023
-
[33]
i just like the stock
Long, S., Lucey, B., Xie, Y., and Yarovaya, L. (2023). “i just like the stock”: The role of reddit sentiment in the gamestop share rally. Financial Review , 58(1):19--37
2023
-
[34]
and McDonald, B
Loughran, T. and McDonald, B. (2011). When is a liability not a liability? textual analysis, dictionaries, and 10-ks. The Journal of finance , 66(1):35--65
2011
-
[35]
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023). Llama: Open and efficient foundation language models. CoRR , abs/2302.13971
2023 arXiv
-
[36]
N., Kaiser, ., and Polosukhin, I
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems , 30
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.