REVIEW 4 major objections 5 minor 126 references
Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Daily interviews of AI-built digital twins of finfluencers recover stock-level beliefs even when no post is made, and those beliefs predict future returns for large-cap stocks.
desk verdict Real-time digital-twin interviews are a genuinely new measurement instrument with a strong validation suite; the missing generic-LLM baseline keeps the persona-attribution claim short of established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The digital twin—an LLM interviewee built from an account's public profile, recent monitored posts, and finance-relevant context—queried daily under a fixed protocol, with a 3:30pm ET cutoff and real-time archival. The workhorse signal is Net Buy Share: the difference between the share of meaningful buy and sell recommendations across the 81 twins for a stock-date. The timed archival is what converts LLM output into a forward-looking instrument.
What would settle it
Run the same daily protocol on the same model with no account conditioning (or with the recent-posts field withheld) and compare silent-region Net Buy Share against the same return windows; if the unconditional LLM signal predicts returns as well as the conditioned one, the persona is not the source of the signal. A second check: re-prompt the model retrospectively on archived pre-return prompts after outcomes are known and see whether the ex-post responses fit returns better than the archived ones, which would indicate residual look-ahead in the protocol.
Extended reading notes
Core claim
The paper claims that standardized, real-time interviews of LLM-based digital twins of finfluencers make the silent region of selective disclosure measurable. Each twin is conditioned on an account's profile and recent monitored posts and is asked a fixed battery of questions about 429 large-cap stocks and seven macro conditions every trading day from December 2025 to March 2026, with responses timestamped under a conservative 3:30pm ET cutoff and archived before any return window begins. Validation against observable public posts shows the interviews align with human recommendations 91.5% of the time (AUC 0.776 vs 0.522 placebo), lean toward recommendations that will only later be posted, a
Load-bearing premise
The results stand on the assumption that the digital-twin responses are not contaminated by the very public posts they are conditioned on, and that the persona conditioning—not the LLM's own pre-trained market knowledge—is what produces the silent-region signal.
Editorial extensions
If this is right
- If the measurement claim holds, the silent region of any selectively disclosing agent can be converted into a comparable, time-stamped belief panel; the paper explicitly extends the logic to sell-side analysts, executives, fund managers, journalists, and advisers.
- The same panel separates direction, disagreement, and uncertainty: direction predicts returns, informative disagreement predicts higher future volatility and lower returns, and a high uncertainty score predicts lower volatility.
- At the market level, average sentiment from the panel does not forecast the S&P 500, but disagreement among twins' macro views does: a 10-point IQR increase predicts 86.4 basis points lower ten-day market returns.
- The real-time archival protocol offers a template for avoiding LLM look-ahead bias in any predictive setting: timestamped, fixed-protocol interviews archive beliefs before outcomes.
Reading between the lines
- A natural control group is absent from the paper: the same interviews run on the same LLM without any persona conditioning. If that control also predicts returns, part of the silent-region signal could come from the model's out-of-the-box market knowledge rather than from the finfluencer persona. I would expect the authors' next test to run exactly this control.
- The contamination risk is concrete and testable: since each twin is conditioned on 'recent monitored posts,' a same-day public recommendation can enter the context and mechanically inflate the 91.5% alignment. A holdout design that conditions twins only on posts older than one trading day would settle this.
- The same instrument could be turned on other classes of selective communicators—central banks, executives in quiet periods, political campaigns, or firms' investor-relations personas—wherever publicly disclosed text is strategic and silence is informative. The paper suggests this but does not test it.
- If the effect is real, its 10-day horizon and absence at 1 day suggest gradual diffusion of retail-facing beliefs, and the natural next step is to test whether the signal survives when conditioned on voluminous LLM-generated content from thousands of accounts, i.e., whether it is an attention phenomenon.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a measurement instrument for selective disclosure in financial social media. For each of 81 monitored X finfluencer accounts, the authors construct an LLM 'digital twin' conditioned on account metadata and recent posts, and repeatedly interview it in real time under a fixed protocol. The stock-pick branch elicits buy/hold/sell scores for 429 large-cap stocks; the macro branch elicits seven broad-market beliefs. The paper validates the twins against same-day public recommendations, pre-disclosure lean, and account-specific structure, then shows that the interview-based directional signal (notably Net Buy Share) predicts cross-sectional future excess returns at 5- and 10-day horizons, with the strongest coefficients in the silent region (84.8% of stock-days, where no same-date public post exists). It also reports that disagreement predicts lower future returns and higher future volatility, and that macro disagreement predicts weaker market returns. The contribution is framed as a measurement method rather than a trading strategy.
Significance. If the measurement claim holds, this is a novel and potentially useful instrument. The real-time archival design is a genuine strength: interviews are timestamped before return windows, avoiding the ex-post look-ahead bias that plagues retrospective LLM queries. The paper is also transparent about its protocol, variable definitions, and limitations, and it includes a thoughtful battery of placebo tests for the return predictability (unrestricted, AR simulation, size-quartile, and FF12-industry reassignments). The separation of direction, disagreement, and uncertainty is conceptually appealing and yields distinct empirical patterns. However, the central attribution of the silent-region signal to the finfluencer personas is not yet established. The absence of a generic-LLM (no-persona) control means the return predictability could reflect the LLM's out-of-the-box market knowledge rather than account-conditioned beliefs, and the same-day validation alignment in Table 3 may be inflated by conditioning on the very posts used as benchmarks. These are load-bearing gaps, not presentation issues.
major comments (4)
- [Internet Appendix A.2 / Table 3 Panel A] The validation alignment of 91.5% and AUC of 0.776 may be mechanically inflated. Each twin is conditioned on the account's 'recent monitored posts' at interview time, but the paper does not state whether same-day public posts are excluded from that conditioning material. The exact-overlap sample is defined by same account, same ticker, same date; if the public post preceded the interview, the twin has already seen it. Please report timestamp ordering between posts and interviews, and either exclude same-day posts from conditioning or rerun the alignment on overlaps where the interview strictly precedes the public post.
- [Section 5.B / Table 8 / Internet Appendix A.2] The key silent-region result (Net Buy Share coefficient 0.503, se 0.218 at the 10-day horizon, Table 8 Panel B) cannot distinguish persona-conditioned beliefs from generic LLM market knowledge. The protocol explicitly permits web search for 'relevant, up-to-date information about the specific tickers,' and the same LLM, given ticker lists and market context, could generate buy/sell scores with similar return predictability. Add a no-persona control (identical protocol without account material) and, ideally, an ablation that withholds recent posts, to attribute the signal to the finfluencer persona.
- [Section 4.A / Table 3 Panel B] The pre-disclosure lean test compares future-public tickers with same-account, same-date placebo or matched control tickers, but these controls are not matched on ticker visibility, news intensity, or LLM familiarity. The positive signed tilt (5.32 vs 2.60, and 4.99 vs 2.28 in the matched comparison) could reflect the LLM's generic knowledge about companies that later receive public recommendations, rather than a persona-specific lean. Report the lean test under a no-persona baseline, or add explicit controls for ticker attention/news activity.
- [Section 4.A / Table 3 Panel C] The account re-identification and adjacent-day similarity tests demonstrate that macro responses retain account-specific structure after removing date-question means. However, these tests do not establish that the silent-region stock-pick return signal is driven by that persona structure. The return predictability in Table 8 could still be generic LLM knowledge. The paper should address this directly, for example by showing that a no-persona version of the stock-pick interviews does not reproduce the silent-region coefficients.
minor comments (5)
- [Section 7.B] The 50-trading-day stock-pick sample and 53-day macro sample are acknowledged as a resource constraint, but the macro regressions in Tables 10–11 have N=53 and several coefficients are significant only at the 10% level. Please state explicitly that these market-level results are exploratory and should be interpreted with caution.
- [Table 8] The public-signal rows in the overlap sample are based on few posts (median one per stock-date), so the public Net Buy Share often takes discrete values -1, 0, or 1. Reporting coefficients per +10 percentage points is awkward for such a variable; consider a standardized or per-unit scaling, or a supplementary table with robust standardized coefficients.
- [Section 3 / Internet Appendix A.5] The 'meaningful' cutoff of 15 points from neutral and the speculation cutoff of 40 are central to several variables. A short sensitivity analysis (e.g., 10/20 and 30/50 cutoffs) would help establish that the main results are not artifacts of these thresholds.
- [Figure 2] The caption mentions a 'solid red line' marking the actual-data coefficient, but the figure appears in grayscale; please adjust the caption or use a distinct line style.
- [Section 5.A.2] The staggered-vintage design is a clear way to address overlapping return windows, but the 'positive and significant' counts across vintages are described as descriptive because vintages share calendar time. Please make this caveat more prominent in the table notes.
Circularity Check
Twin-validation step is partly circular (conditioning input vs. same-day benchmark), while the forward-looking return tests themselves are self-contained.
-
self definitional
[Section 3 (digital-twin construction) and Section 4.A / Table 3 Panel A]
"For each finfluencer, the LLM is given information about the account's public persona: profile metadata, recent monitored posts, and finance-relevant context available at the time of the interview ... When a finfluencer account publicly recommends a stock on a given trading date, does the digital-twin interview about that stock point in the same buy or sell direction? ... the interview direction aligns with the public recommendation 91.5% of the time."
The twin's output is conditioned on the account's own recent monitored posts, and the headline validation (Table 3 Panel A) is the same-day match between twin and human recommendations. Stock-pick interviews run before the open or during trading hours (IA A.4) and the context is updated as new content is produced, so a recommendation posted before the interview is available to the model as conditioning input. The 91.5% alignment therefore restates the model's input rather than demonstrating independent recovery of the persona's view; the paper neither excludes same-day posts from the prompt nor reports alignment when the reference post is provably absent. The only contamination-free validation (Table 3 Panel B, pre-disclosure signed tilt 5.32 vs. 2.60 on a 0–100 scale) is much weaker and i
full rationale
The paper's central empirical contribution is the silent-region return predictability: digital-twin interview signals are archived before the return windows, and the Fama-MacBeth, calendar-time, and staggered-vintage designs with placebo benchmarks are genuinely forward-looking. Those return regressions do not reduce to the model's inputs, so the core derivation is self-contained. However, one load-bearing validation step is partially circular. The digital twin is defined by conditioning on 'recent monitored posts' (Section 3, IA A.2), and the headline validation (Table 3 Panel A) is the same-day, same-ticker alignment with the human's public recommendation. Because stock-pick interviews can occur after a recommendation has been posted, the reference post can be part of the conditioning context, inflating the 91.5% alignment rate. The paper never excludes same-day posts from the prompt, and the pre-disclosure lean test (Panel B), which is immune to this particular contamination, is far weaker. Separately, the attribution of the silent-region signal to the finfluencer personas is unverified: the protocol explicitly licenses web search and 'general knowledge of market conditions,' and no no-persona LLM baseline is run, so the silent-region predictability could originate from generic LLM market knowledge rather than persona beliefs. This is a missing control rather than a by-construction reduction; it does not make the return tests circular, but it weakens the measurement claim. Self-citations (Cerina and Duch 2025; De Bartolomeis et al. 2026; Sorescu and Subrahmanyam 2006) are minor and not load-bearing. Overall: one genuine partial circularity in the validation pillar, with an independent forward-looking return analysis, so a moderate score of 4 is appropriate.
Assumptions & free parameters
free parameters (5)
- Meaningful recommendation cutoff =
15 points from neutral (scores >=65 buy, <=35 sell)
- Speculation score cutoff =
40
- Buy/sell classification thresholds =
65/35
- Net Sentiment classification rule =
At least 3 of 7 questions strong and low-speculation
- Informative Polarization minimum =
3 informative recommendations required
assumptions (5)
- domain assumption LLM responses conditioned on account profiles are valid representations of the finfluencer public persona
- domain assumption Real-time web search during interviews provides only information a market participant could access at the same time
- domain assumption The 3:30 p.m. ET cutoff correctly assigns interviews to event dates
- domain assumption The public-post comparison panel accurately classifies public recommendations
- standard math The placebo distributions correctly capture the null under short-sample time-series dependence
invented entities (1)
-
Digital twin (LLM interviewee conditioned on finfluencer account)
Cite this review
Pith. "Pith review of Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media." pith.science (2026). https://pith.science/paper/F44U7G7E
@misc{pith2026260801181,
author = {Pith},
title = {Pith review of: Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media},
year = {2026},
howpublished = {\url{https://pith.science/paper/F44U7G7E}},
note = {Machine review of arXiv:2608.01181}
}
read the original abstract
Social media affect financial markets, but public posts by financial media personas are voluntary disclosures. What is not disclosed is therefore usually unobserved. We address this measurement problem by conducting repeated, real-time interviews of "digital twins" built from monitored finfluencers' X accounts under a fixed protocol. The interviews recover stock-level public-persona belief proxies even when no public recommendation is made. Because the interviews are generated and archived before the relevant return windows, the design avoids the look-ahead bias that arises when LLMs are queried ex post. The evidence shows that information obtained from these digital-twin interviews predicts the cross section of large-cap stock returns in the expected direction. Repeated real-time interviews therefore show how selective disclosure can be turned into measurable panels of market views.
Figures
Reference graph
Works this paper leans on
-
[1]
Abdelrahman, Mahmoud, Edgardo Macatulad, Binyu Lei, Matias Quintana, Clayton Miller, and Filip Biljecki, 2025, What is a digital twin anyway? deriving the definition for the built environment from over 15,000 scientific publications, Building and Environment\/ 274, 112748
2025
-
[2]
Adams, Travis, Andrea Ajello, Diego Silva, and Francisco V \'a zquez-Grande, 2023, More than words: Twitter chatter and financial market sentiment, Technical Report FEDS 2023-034, Board of Governors of the Federal Reserve System, Accessed 2026-03-01
2023
-
[3]
Andersson, Charlotte, Andrew D Johnson, Emelia Benjamin, Daniel Levy, and Ramachandran Vasan, 2019, 70-year legacy of the Framingham Heart Study , Nature Reviews Cardiology\/ 16, 687--698
2019
-
[4]
Andrei, Daniel, and Michael Hasler, 2015, Investor attention and stock market volatility, Review of Financial Studies\/ 28, 33--72
2015
-
[5]
Angelopoulos, Anastasios N, Stephen Bates, Clara Fannjiang, Michael I Jordan, and Tijana Zrnic, 2023, Prediction-powered inference, Science\/ 382, 669--674
2023
-
[6]
Frank, 2004, Is all that talk just noise? T he information content of internet stock message boards, Journal of Finance\/ 59, 1259--1294
Antweiler, Werner, and Murray Z. Frank, 2004, Is all that talk just noise? T he information content of internet stock message boards, Journal of Finance\/ 59, 1259--1294
2004
-
[7]
Baker, Malcolm, and Jeffrey Wurgler, 2006, Investor sentiment and the cross-section of stock returns, Journal of Finance\/ 61, 1645--1680
2006
-
[8]
Barber, Brad, Reuven Lehavy, Maureen McNichols, and Brett Trueman, 2001, Can investors profit from the prophets? S ecurity analyst recommendations and stock returns, Journal of finance\/ 56, 531--563
2001
Show all 126 references
-
[9]
Barber, Brad M., Xing Huang, Terrance Odean, and Christopher Schwarz, 2022, Attention-induced trading and returns: Evidence from R obinhood users, Journal of Finance\/ 77, 3141--3190
2022
-
[10]
Barber, Brad M., and Terrance Odean, 2008, All that glitters: The effect of attention and news on the buying behavior of individual and institutional investors, Review of Financial Studies\/ 21, 785--818
2008
-
[11]
Barberis, Nicholas, Andrei Shleifer, and Robert Vishny, 1998, A model of investor sentiment, Journal of Financial Economics\/ 49, 307--343
1998
-
[12]
Barrot, Jean-Noël, Ron Kaniel, and David Sraer, 2016, Information and the trading behavior of retail investors, Review of Financial Studies\/ 29, 2667--2701
2016
-
[13]
Berger, Jonah, and Katherine L Milkman, 2012, What makes online content viral?, Journal of marketing research\/ 49, 192--205
2012
-
[14]
Bertrand, Marianne, and Sendhil Mullainathan, 2004, Are E mily and G reg more employable than L akisha and J amal? A field experiment on labor market discrimination, American Economic Review\/ 94, 991--1013
2004
-
[15]
Jones, and Xiaoyan Zhang, 2021, Retail trading and market quality, Journal of Finance\/ 76, 1029--1064
Boehmer, Ekkehart, Charles M. Jones, and Xiaoyan Zhang, 2021, Retail trading and market quality, Journal of Finance\/ 76, 1029--1064
2021
-
[16]
Bollen, Johan, Huina Mao, and Xiao-Jun Zeng, 2011, Twitter mood predicts the stock market, Journal of Computational Science\/ 2, 1--8
2011
-
[17]
Bordalo, Pedro, Nicola Gennaioli, Rafael La Porta, and Andrei Shleifer, 2019, Diagnostic expectations and stock returns, Journal of Finance\/ 74, 2839--2874
2019
-
[18]
Brick, J Michael, and Douglas Williams, 2013, Explaining rising nonresponse rates in cross-sectional surveys, The ANNALS of the American Academy of Political and Social Science\/ 645, 36--59
2013
-
[19]
Broska, David, Michael Howes, and Austin van Loon, 2025, The mixed subjects design: Treating large language models as potentially informative observations, Sociological Methods & Research\/ 54, 1074--1109
2025
-
[20]
Brown, Tom B., Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff...
2020
-
[21]
Cerina, Roberto, and Raymond Duch, 2025, The 2024 US presidential election PoSSUM poll, PS: Political Science & Politics\/ 58, 286--297
2025
-
[22]
Chen, Shuaiyu, T. Clifton Green, Huseyin Gulen, and Dexin Zhou, 2025, What does chatgpt make of historical stock returns? extrapolation and miscalibration in llm stock return forecasts, Purdue University Working Paper\/
2025
-
[23]
Kelly, and Dacheng Xiu, 2022, Expected returns and large language models, Journal of Finance\/ forthcoming
Chen, Yifei, Bryan T. Kelly, and Dacheng Xiu, 2022, Expected returns and large language models, Journal of Finance\/ forthcoming
2022
-
[24]
Claassen, Ryan L, and John Barry Ryan, 2025, Biased polls: I nvestigating the pressures survey respondents feel: Rl claassen, jb ryan, Acta Politica\/ 60, 821--843
2025
-
[25]
Anthony, Zhaoran Lu, and Marina Niessner, 2022, The gamestop short squeeze, Review of Asset Pricing Studies\/ 12, 1--40
Cookson, J. Anthony, Zhaoran Lu, and Marina Niessner, 2022, The gamestop short squeeze, Review of Asset Pricing Studies\/ 12, 1--40
2022
-
[26]
Anthony, and Marina Niessner, 2019, When saving is gambling, Journal of Finance\/ 74, 1887--1939
Cookson, J. Anthony, and Marina Niessner, 2019, When saving is gambling, Journal of Finance\/ 74, 1887--1939
2019
-
[27]
Anthony, and Marina Niessner, 2020, Why don't we agree? E vidence from a social network of investors, Journal of Finance\/ 75, 173--228
Cookson, J. Anthony, and Marina Niessner, 2020, Why don't we agree? E vidence from a social network of investors, Journal of Finance\/ 75, 173--228
2020
-
[28]
Da, Zhi, Joseph Engelberg, and Pengjie Gao, 2011, In search of attention, Journal of Finance\/ 66, 1461--1499
2011
-
[29]
Daniel, Kent, Alexander Klos, and Simon Rottke, 2023, The dynamics of disagreement, Review of Financial Studies\/ 36, 2431--2467
2023
-
[30]
De Bartolomeis, Piersilvio, Javier Abad, Guanbo Wang, Konstantin Donhauser, Raymond Duch, Fanny Yang, and Issa Dahabreh, 2026, Efficient randomized experiments using foundation models, Advances in Neural Information Processing Systems\/ 38, 159470--159509
2026
-
[31]
Bradford, Andrei Shleifer, Lawrence H
De Long, J. Bradford, Andrei Shleifer, Lawrence H. Summers, and Robert J. Waldmann, 1990, Noise trader risk in financial markets, Journal of Political Economy\/ 98, 703--738
1990
-
[32]
De Veirman, Marijke, Veroline Cauberghe, and Liselot Hudders, 2017, Marketing through instagram influencers: the impact of number of followers and product divergence on brand attitude, International journal of advertising\/ 36, 798--828
2017
-
[33]
Kelly, Mohammad Pourmohammadi, and Hanqing Tian, 2026, The inefficient pricing of news,, National Bureau of Economic Research\/ working paper
Didisheim, Antoine, Bryan T. Kelly, Mohammad Pourmohammadi, and Hanqing Tian, 2026, The inefficient pricing of news,, National Bureau of Economic Research\/ working paper
2026
-
[34]
Malloy, and Anna Scherbina, 2002, Differences of opinion and the cross section of stock returns, Journal of Finance\/ 57, 2113--2141
Diether, Karl B., Christopher J. Malloy, and Anna Scherbina, 2002, Differences of opinion and the cross section of stock returns, Journal of Finance\/ 57, 2113--2141
2002
-
[35]
Dye, Ronald A, 1985, Disclosure of nonproprietary information, Journal of accounting research\/ 123--145
1985
-
[36]
Eadie, Ashley L, Holly Fernandez Lynch, Naomi Scheinerman, and Ravi B Parikh, 2026, The arrival of digital twins and in silico trials in drug development, Nature Medicine\/ 1--5
2026
-
[37]
Ellsberg, Daniel, 1961, Risk, ambiguity, and the savage axioms, Quarterly Journal of Economics\/ 75, 643--669
1961
-
[38]
Engelberg, Joseph, Arzu Ozoguz, and Shenghua Wang, 2018, The social transmission of information in financial markets, Review of Financial Studies\/ 31, 3849--3879
2018
-
[39]
Parsons, 2011, The causal impact of media in financial markets, Journal of Finance\/ 66, 67--97
Engelberg, Joseph E., and Christopher A. Parsons, 2011, The causal impact of media in financial markets, Journal of Finance\/ 66, 67--97
2011
-
[40]
Preece, 2024, The finfluencer appeal: Investing in the age of social media, Technical report, CFA Institute Research and Policy Center
Espeute, Serena, and Rhodri G. Preece, 2024, The finfluencer appeal: Investing in the age of social media, Technical report, CFA Institute Research and Policy Center
2024
-
[41]
Fainmesser, Itay P, and Andrea Galeotti, 2021, The market for online influence, American Economic Journal: Microeconomics\/ 13, 332--372
2021
-
[42]
French, 1992, The cross-section of expected stock returns, Journal of Finance\/ 47, 427--465
Fama, Eugene F., and Kenneth R. French, 1992, The cross-section of expected stock returns, Journal of Finance\/ 47, 427--465
1992
-
[43]
Fama, Eugene F, and Kenneth R French, 2008, Dissecting anomalies, Journal of Finance\/ 63, 1653--1678
2008
-
[44]
MacBeth, 1973, Risk, return, and equilibrium: Empirical tests, Journal of Political Economy\/ 81, 607--636
Fama, Eugene F., and James D. MacBeth, 1973, Risk, return, and equilibrium: Empirical tests, Journal of Political Economy\/ 81, 607--636
1973
-
[45]
Fang, Lily, and Joel Peress, 2009, Media coverage and the cross-section of stock returns, The journal of finance\/ 64, 2023--2052
2009
-
[46]
Financial Industry Regulatory Authority , 2025, Social media-influenced investing, FINRA report, Published December 11, 2025
2025
-
[47]
Gentzkow, Matthew, and Jesse M Shapiro, 2006, Media bias and reputation, Journal of political Economy\/ 114, 280--316
2006
-
[48]
Giglio, Stefano, Matteo Maggiori, Johannes Stroebel, and Stephen Utkus, 2021, Five facts about beliefs and portfolios, American Economic Review\/ 111, 1481--1522
2021
-
[49]
Glasserman, Paul, and Caden Lin, 2023, Assessing look-ahead bias in stock return predictions generated by gpt sentiment analysis, arXiv preprint arXiv:2309.17322\/
2023 arXiv
-
[50]
Gordillo Mart \' nez, Lizeth, 2024, The impact of the Social Media Sentiment Index on S&P 500 returns , The An \'a huac Journal\/ https://doi.org/10.36105/theanahuacjour.2024v24n1.08
2024 doi
-
[51]
Goyal, Amit, and Ivo Welch, 2008, A comprehensive look at the empirical performance of equity premium prediction, Review of Financial Studies\/ 21, 1455--1508
2008
-
[52]
Goyal, Amit, Ivo Welch, and Athanasse Zafirov, 2024, A comprehensive 2022 look at the empirical performance of equity premium prediction, Review of Financial Studies\/ 37, 3490--3557
2024
-
[53]
Graham, John R, 1999, Herding among investment newsletters: Theory and evidence, The Journal of Finance\/ 54, 237--268
1999
-
[54]
Greenwood, Robin, and Andrei Shleifer, 2014, Expectations of returns and expected returns, Review of Financial Studies\/ 27, 714--746
2014
-
[55]
Stiglitz, 1980, On the impossibility of informationally efficient markets, American Economic Review\/ 70, 393--408
Grossman, Sanford J., and Joseph E. Stiglitz, 1980, On the impossibility of informationally efficient markets, American Economic Review\/ 70, 393--408
1980
-
[56]
Groves, Robert M, and Emilia Peytcheva, 2008, The impact of nonresponse rates on nonresponse bias: a meta-analysis, Public opinion quarterly\/ 72, 167--189
2008
-
[57]
Hassan, Tarek A., Stephan Hollander, Laurence van Lent, and Ahmed Tahoun, 2019, Firm-level political risk: Measurement and effects, Quarterly Journal of Economics\/ 134, 2135--2202
2019
-
[58]
Hayek, Friedrich A., 1945, The use of knowledge in society, American Economic Review\/ 35, 519--530
1945
-
[59]
He, Songrun, Linying Lv, Asaf Manela, and Jimmy Wu, 2025, Chronologically consistent large language models, arXiv preprint\/ 2502.21206
2025 arXiv
-
[60]
Hoberg, Gerard, and Gordon Phillips, 2016, Text-based network industries and endogenous product differentiation, Journal of Political Economy\/ 124, 1423--1465
2016
-
[61]
Hong, Harrison, Jeffrey D Kubik, and Amit Solomon, 2000, Security analysts' career concerns and herding of earnings forecasts, The Rand journal of economics\/ 121--144
2000
-
[62]
Kubik, and Jeremy C
Hong, Harrison, Jeffrey D. Kubik, and Jeremy C. Stein, 2004, Social interaction and stock-market participation, Journal of Finance\/ 59, 137--163
2004
-
[63]
Hong, Harrison, and Jeremy Stein, 2007, Disagreement and the stock market, Journal of Economic Perspectives\/ 21, 109--128
2007
-
[64]
Stein, 1999, A unified theory of underreaction, momentum trading, and overreaction in asset markets, Journal of Finance\/ 54, 2143--2184
Hong, Harrison, and Jeremy C. Stein, 1999, A unified theory of underreaction, momentum trading, and overreaction in asset markets, Journal of Finance\/ 54, 2143--2184
1999
-
[65]
Manning, 2023, Large language models as simulated economic agents: What can we learn from homo silicus?, Technical report, National Bureau of Economic Research
Horton, John J, Apostolos Filippas, and Benjamin S. Manning, 2023, Large language models as simulated economic agents: What can we learn from homo silicus?, Technical report, National Bureau of Economic Research
2023
-
[66]
Iliut \,a , Miruna-Elena, Mihnea-Alexandru Moisescu, Eugen Pop, Anca-Daniela Ionita, Simona-Iuliana Caramihai, and Traian-Costin Mitulescu, 2024, Digital twin---a review of the evolution from concept to technology and its analytical perspectives on applications in various fiel...
2024
-
[67]
Jadhav, Aakanksha, and Vishal Mirza, 2025, Large language models in equity markets: Applications, techniques, and insights, Frontiers in Artificial Intelligence\/ 1608365
2025
-
[68]
Jegadeesh, Narasimhan, and Woojin Kim, 2006, Value of analyst recommendations: International evidence, Journal of Financial Markets\/ 9, 274--309
2006
-
[69]
Jegadeesh, Narasimhan, and Sheridan Titman, 1993, Returns to buying winners and selling losers: Implications for stock market efficiency, Journal of Finance\/ 48, 65--91
1993
-
[70]
Jones, David, Chris Snider, Aydin Nassehi, Jason Yon, and Ben Hicks, 2020, Characterising the digital twin: A systematic literature review, CIRP journal of manufacturing science and technology\/ 29, 36--52
2020
-
[71]
Kakhbod, Ali, Seyed Mohammad Kazempour, Dmitry Livdan, and Norman Schuerhoff, 2023, Finfluencers, Swiss Finance Institute\/ Research Paper
2023
-
[72]
Keeter, Scott, Nick Hatley, Courtney Kennedy, and Arnold Lau, 2017, What low response rates mean for telephone surveys, Pew Research Center\/ 15, 1--39
2017
-
[73]
Kennedy, Courtney, and Hannah Hartig, 2019, Response rates in telephone surveys have resumed their decline, Pew Research Center\/ 27
2019
-
[74]
Kessler, Judd B, Corinne Low, and Colin D Sullivan, 2019, Incentivized resume rating: Eliciting employer preferences without deception, American Economic Review\/ 109, 3713--3744
2019
-
[75]
Nikolaev, 2024, Financial statement analysis with large language models, Working paper
Kim, Alex G., Maximilian Muhn, and Valeri V. Nikolaev, 2024, Financial statement analysis with large language models, Working paper
2024
-
[76]
Knight, Frank H., 1921, Risk, Uncertainty and Profit\/ (Houghton Mifflin, Boston, MA)
1921
-
[77]
Koijen, Ralph S.J., and Bradford Lynch Levy, 2026, Assessing the benefits of optimized agentic AI systems for asset pricing, Available at SSRN\/ 6474601
2026
-
[78]
Kuhn, Peter, and Kailing Shen, 2013, Gender discrimination in job ads: Evidence from china, Quarterly Journal of Economics\/ 128, 287--336
2013
-
[79]
Lalwani, Vaibhav, 2025, Finfluencer recommendations, Economics Letters\/ 112511
2025
-
[80]
Lattanzi, Luca, Roberto Raffaeli, Margherita Peruzzini, and Marcello Pellicciari, 2021, Digital twin for smart manufacturing: A review of concepts towards a practical industrial implementation, International Journal of Computer Integrated Manufacturing\/ 34, 567--597
2021
-
[81]
Lehner, Sebastian, and Alejandro Lopez-Lira, 2026, Chatgpt as a time capsule: The limits of price discovery, arXiv preprint arXiv:2604.21433\/
2026 arXiv
-
[82]
u ttler, Mike Lewis, Wen tau Yih, Tim Rockt \
Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen tau Yih, Tim Rockt \"a schel, Sebastian Riedel, and Douwe Kiela, 2020, Retrieval-augmented generation for knowledge-intensive nlp tasks, Advanc...
2020
-
[83]
Li, Weixian Waylon, Hyeonjun Kim, Mihai Cucuringu, and Tiejun Ma, 2026, Can llm-based financial investing strategies outperform the market in long run?, in Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1\/ , KDD '26 (Association for Com...
2026
-
[84]
Lopez-Lira, Alejandro, and Yuehua Tang, 2024, Can ChatGPT forecast stock price movements? R eturn predictability and large language models, arXiv preprint 2304.07619
2024
-
[85]
Loughran, Tim, and Bill McDonald, 2011, When is a liability not a liability? T extual analysis, dictionaries, and 10- K s, Journal of Finance\/ 66, 35--65
2011
-
[86]
Ludwig, Jens, Sendhil Mullainathan, and Ashesh Rambachan, 2024, Large language models: An applied econometric framework, Annual Review of Economics\/ 18
2024
-
[87]
Mahdavi, Sedigheh, Pradeep Kumar Joshi, Lina Huertas Guativa, Upmanyu Singh, et al., 2025, Integrating large language models in financial investments and market analysis: A survey, arXiv preprint\/ 2507.01990
2025 arXiv
-
[88]
Mahmood, Syed S, Daniel Levy, Ramachandran S Vasan, and Thomas J Wang, 2014, The Framingham Heart Study and the epidemiology of cardiovascular disease: A historical perspective, Lancet\/ 383, 999--1008
2014
-
[89]
Mercer, Andrew, and Arnold Lau, 2023, Comparing two types of online survey samples, Pew Research Center\/ 40--50
2023
-
[90]
Miller, Edward M., 1977, Risk, uncertainty, and divergence of opinion, Journal of Finance\/ 32, 1151--1168
1977
-
[91]
Mitchell, Mark L., and Erik Stafford, 2000, Managerial decisions and long-term stock price performance, Journal of Business\/ 73, 287--329
2000
-
[92]
Muchnik, Lev, Sinan Aral, and Sean J Taylor, 2013, Social influence bias: A randomized experiment, Science\/ 341, 647--651
2013
-
[93]
Mullainathan, Sendhil, and Andrei Shleifer, 2005, The market for news, American economic review\/ 95, 1031--1053
2005
-
[94]
West, 1987, A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix, Econometrica\/ 55, 703--708
Newey, Whitney K., and Kenneth D. West, 1987, A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix, Econometrica\/ 55, 703--708
1987
-
[95]
Mendoza-Urdiales, 2023, Social sentiment and impact in US equity market: A n automated approach, Social Network Analysis and Mining\/ 13
N \'u \ n ez-Mora, Jos \'e Antonio, and Rom \'a n A. Mendoza-Urdiales, 2023, Social sentiment and impact in US equity market: A n automated approach, Social Network Analysis and Mining\/ 13
2023
-
[96]
Ontario Securities Commission , 2025, OSC research uncovers concerns about finfluencers' power of persuasion
2025
-
[97]
Ouyang, Long, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, an...
2022
-
[98]
Pallais, Amanda, 2014, Inefficient hiring in entry-level labor markets, American Economic Review\/ 104, 3565--3599
2014
-
[99]
Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S
Park, Joon Sung, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S. Bernstein, 2024, Generative agent simulations of 1,000 people, arXiv preprint arXiv:2411.10109\/ 52
2024 arXiv
-
[100]
Petrocik, John R, 1996, Issue ownership in presidential elections, with a 1980 case study, American Journal of Political Science\/ 825--850
1996
-
[101]
Ranco, Gabriele, Darko Aleksovski, Guido Caldarelli, Miha Gr c ar, and Igor Mozeti c , 2015, The effects of twitter sentiment on stock price returns, PLoS ONE\/ 10, e0138441
2015
-
[102]
Reis, Ricardo, 2006, Inattentive consumers, Journal of Monetary Economics\/ 53, 1761--1800
2006
-
[103]
Riker, William H., 1996, The Strategy of Rhetoric: Campaigning for the American Constitution\/ (Yale University Press)
1996
-
[104]
Sarkar, Suproteem K., and Keyon Vafa, 2024, Lookahead bias in pretrained language models, Working Paper, University of Chicago
2024
-
[105]
Scharfstein, David S, and Jeremy C Stein, 1990, Herd behavior and investment, The American economic review\/ 465--479
1990
-
[106]
Shleifer, Andrei, 2000, Inefficient Markets: An Introduction to Behavioral Finance\/ (Oxford University Press)
2000
-
[107]
Sims, Christopher A., 2003, Implications of rational inattention, Journal of Monetary Economics\/ 50, 665--690
2003
-
[108]
Theriault, 2004, The structure of political argument and the logic of issue framing, Studies in Public Opinion: Attitudes, Nonattitudes, Measurement Error, and Change\/ 3, 133--65
Sniderman, Paul M., and Sean M. Theriault, 2004, The structure of political argument and the logic of issue framing, Studies in Public Opinion: Attitudes, Nonattitudes, Measurement Error, and Change\/ 3, 133--65
2004
-
[109]
Sorescu, Sorin, and Avanidhar Subrahmanyam, 2006, The cross section of analyst recommendations, Journal of Financial and Quantitative Analysis\/ 41, 139--168
2006
-
[110]
Spence, Michael, 1973, Job market signaling, Quarterly Journal of Economics\/ 87, 355--374
1973
-
[111]
Sandner, Andranik Tumasjan, and Isabell M
Sprenger, Timm O., Philipp G. Sandner, Andranik Tumasjan, and Isabell M. Welpe, 2014, Tweets and trades: The information content of stock microblogs, European Financial Management\/ 20, 926--957
2014
-
[112]
Sul, Hong Kee, Alan R Dennis, and Lingyao Yuan, 2017, Trading on T witter: U sing social media sentiment to predict stock returns, Decision Sciences\/ 48, 454--488
2017
-
[113]
Tao, Fei, He Zhang, Ang Liu, and Andrew YC Nee, 2018, Digital twin in industry: State-of-the-art, IEEE Transactions on industrial informatics\/ 15, 2405--2415
2018
-
[114]
Tetlock, Paul C., 2007, Giving content to investor sentiment: The role of media in the stock market, Journal of Finance\/ 62, 1139--1168
2007
-
[115]
Tetlock, Paul C., Maytal Saar-Tsechansky, and Sofus Macskassy, 2008, More than words: Quantifying language to measure firms' fundamentals, Journal of Finance\/ 63, 1437--1467
2008
-
[116]
Trueman, Brett, 1994, Analyst forecasts and herding behavior, The review of financial studies\/ 7, 97--124
1994
-
[117]
Securities and Exchange Commission Investor Advisory Committee , 2024, Recommendations regarding the protection of investors in their interactions with finfluencers
U.S. Securities and Exchange Commission Investor Advisory Committee , 2024, Recommendations regarding the protection of investors in their interactions with finfluencers
2024
-
[118]
Gomez, Łukasz Kaiser, and Illia Polosukhin, 2017, Attention is all you need, Advances in neural information processing systems\/ 30, I
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin, 2017, Attention is all you need, Advances in neural information processing systems\/ 30, I
2017
-
[119]
Verrecchia, Robert E, 1983, Discretionary disclosure, Journal of accounting and economics\/ 5, 179--194
1983
-
[120]
Vosoughi, Soroush, Deb Roy, and Sinan Aral, 2018, The spread of true and false news online, science\/ 359, 1146--1151
2018
-
[121]
Wang, Yue, Tianfan Fu, Yinlong Xu, Zihan Ma, Hongxia Xu, Bang Du, Yingzhou Lu, Honghao Gao, Jian Wu, and Jintai Chen, 2024, Twin-gpt: digital twins for clinical trials via large language model, ACM Transactions on Multimedia Computing, Communications and Applications\/
2024
-
[122]
Womack, Kent L, 1996, Do brokerage analysts' recommendations have investment value?, Journal of Finance\/ 51, 137--167
1996
-
[123]
Xiao, Mengxi, Zihao Jiang, Lingfei Qian, Zhengyu Chen, Yueru He, Yijing Xu, Yuecheng Jiang, Dong Li, Ruey-Ling Weng, Min Peng, Jimin Huang, Sophia Ananiadou, and Qianqian Xie, 2025, Retrieval-augmented large language models for financial time series forecasting, arXiv preprint...
2025 arXiv
-
[124]
Yang, Yuzhe, Yifei Zhang, Minghao Wu, Kaidi Zhang, Yunmiao Zhang, Honghai Yu, Yan Hu, and Benyou Wang, 2026, Twinmarket: A scalable behavioral and social simulation for financial markets, in The Thirty-ninth Annual Conference on Neural Information Processing Systems\/
2026
-
[125]
Yao, Jun-Feng, Yong Yang, Xue-Cheng Wang, and Xiao-Peng Zhang, 2023, Systematic review of digital twin technology and applications, Visual Computing for Industry, Biomedicine, and Art\/ 6, 10
2023
-
[126]
Yuan, Hang, Saizhuo Wang, and Jian Guo, 2024, Alpha-gpt 2.0: Human-in-the-loop ai for quantitative investment, arXiv preprint arXiv:2402.09746\/
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.