{"id":"783c283c-843e-49ba-a8df-50e30e20cb99","arxiv_id":"2608.01181","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Daily real-time LLM digital-twin interviews of finfluencer accounts predict cross-sectional large-cap returns over the next ten trading days, mainly in the silent region with no concurrent public post.","lead":"Researchers interviewed AI digital twins of stock influencers on X every day, asking fixed questions about stocks and the market. The answers, archived before the return windows, predicted the direction of large-cap stock returns, especially for stocks the influencers never mentioned publicly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing no-persona LLM baseline leaves attribution of silent-region predictability to finfluencer personas unverified","rationale":"The reader's weakest assumption identifies two threats: same-day contamination of the validation benchmark and absence of a generic-LLM baseline. I focus on the generic-LLM baseline as the single most load-bearing concern because it directly threatens the central interpretive claim, not just the validation alignment. If a no-persona LLM produces the same silent-region predictability, then the digital-twin instrument adds nothing beyond the underlying model's market knowledge, and the paper's contribution as a measurement device for selective disclosure collapses. The current validation evidence is indirect: account-specific macro structure (Table 3 Panel C) demonstrates persona conditioning affects responses but does not establish that the return-predictive content is persona-specific. The proposed test is a direct, feasible check that would settle whether the silent-region signal is attributable to the finfluencer personas. Because the paper is otherwise carefully designed with real-time archiving, timing rules, and thoughtful placebos, the appropriate verdict remains conditional pending this test; hence I recommend UNCHANGED relative to the reader's CONDITIONAL.","tokens_in":49393,"tokens_out":5001,"duration_ms":56368,"concrete_test":"Rerun the stock-pick branch on the same 50 event dates and 429-ticker universe with the identical prompt and web-search allowance but no account conditioning (e.g., a single neutral 'equity analyst' persona). Construct Net Buy Share from these responses and re-estimate the Table 8 Panel B no-post regressions. If the generic-LLM coefficient is positive and its 95% confidence interval contains the persona-conditioned estimate of 0.503, the silent-region signal is attributable to generic market knowledge and the persona interpretation fails. If the generic coefficient is near zero and significantly different from 0.503, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The return-predictability result may be statistically sound, but the paper's central measurement claim depends on attribution: the silent-region signal must reflect the finfluencer personas, not generic LLM market knowledge. The stock-pick protocol (Internet Appendix A.2) conditions on account metadata and recent posts but also permits web search for 'relevant, up-to-date information about the specific tickers...'. A generic LLM given the same ticker lists and real-time market context could generate buy/sell scores that predict returns with no persona information at all. The key silent-region estimate (Table 8, Panel B, Net Buy Share: 0.503, se 0.218 at 10 days) cannot distinguish these hypotheses. The validation tests do not close this gap: Table 3 Panel A is potentially contaminated by same-day posts entering the conditioning context; Panel B's pre-disclosure lean could arise from generic knowledge about the eventually-mentioned ticker; Panel C shows only that macro responses have account-specific structure, not that this structure drives stock-return predictability. Without a generic-LLM control, the claim that the instrument recovers persona beliefs in the silent region is unverified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a measurement instrument for selective disclosure in financial social media. For each of 81 monitored X finfluencer accounts, the authors construct an LLM 'digital twin' conditioned on account metadata and recent posts, and repeatedly interview it in real time under a fixed protocol. The stock-pick branch elicits buy/hold/sell scores for 429 large-cap stocks; the macro branch elicits seven broad-market beliefs. The paper validates the twins against same-day public recommendations, pre-disclosure lean, and account-specific structure, then shows that the interview-based directional signal (notably Net Buy Share) predicts cross-sectional future excess returns at 5- and 10-day horizons, with the strongest coefficients in the silent region (84.8% of stock-days, where no same-date public post exists). It also reports that disagreement predicts lower future returns and higher future volatility, and that macro disagreement predicts weaker market returns. The contribution is framed as a measurement method rather than a trading strategy.","tokens_in":49678,"tokens_out":4491,"duration_ms":53233,"significance":"If the measurement claim holds, this is a novel and potentially useful instrument. The real-time archival design is a genuine strength: interviews are timestamped before return windows, avoiding the ex-post look-ahead bias that plagues retrospective LLM queries. The paper is also transparent about its protocol, variable definitions, and limitations, and it includes a thoughtful battery of placebo tests for the return predictability (unrestricted, AR simulation, size-quartile, and FF12-industry reassignments). The separation of direction, disagreement, and uncertainty is conceptually appealing and yields distinct empirical patterns. However, the central attribution of the silent-region signal to the finfluencer personas is not yet established. The absence of a generic-LLM (no-persona) control means the return predictability could reflect the LLM's out-of-the-box market knowledge rather than account-conditioned beliefs, and the same-day validation alignment in Table 3 may be inflated by conditioning on the very posts used as benchmarks. These are load-bearing gaps, not presentation issues.","major_comments":[{"comment":"The validation alignment of 91.5% and AUC of 0.776 may be mechanically inflated. Each twin is conditioned on the account's 'recent monitored posts' at interview time, but the paper does not state whether same-day public posts are excluded from that conditioning material. The exact-overlap sample is defined by same account, same ticker, same date; if the public post preceded the interview, the twin has already seen it. Please report timestamp ordering between posts and interviews, and either exclude same-day posts from conditioning or rerun the alignment on overlaps where the interview strictly precedes the public post.","section":"Internet Appendix A.2 / Table 3 Panel A"},{"comment":"The key silent-region result (Net Buy Share coefficient 0.503, se 0.218 at the 10-day horizon, Table 8 Panel B) cannot distinguish persona-conditioned beliefs from generic LLM market knowledge. The protocol explicitly permits web search for 'relevant, up-to-date information about the specific tickers,' and the same LLM, given ticker lists and market context, could generate buy/sell scores with similar return predictability. Add a no-persona control (identical protocol without account material) and, ideally, an ablation that withholds recent posts, to attribute the signal to the finfluencer persona.","section":"Section 5.B / Table 8 / Internet Appendix A.2"},{"comment":"The pre-disclosure lean test compares future-public tickers with same-account, same-date placebo or matched control tickers, but these controls are not matched on ticker visibility, news intensity, or LLM familiarity. The positive signed tilt (5.32 vs 2.60, and 4.99 vs 2.28 in the matched comparison) could reflect the LLM's generic knowledge about companies that later receive public recommendations, rather than a persona-specific lean. Report the lean test under a no-persona baseline, or add explicit controls for ticker attention/news activity.","section":"Section 4.A / Table 3 Panel B"},{"comment":"The account re-identification and adjacent-day similarity tests demonstrate that macro responses retain account-specific structure after removing date-question means. However, these tests do not establish that the silent-region stock-pick return signal is driven by that persona structure. The return predictability in Table 8 could still be generic LLM knowledge. The paper should address this directly, for example by showing that a no-persona version of the stock-pick interviews does not reproduce the silent-region coefficients.","section":"Section 4.A / Table 3 Panel C"}],"minor_comments":[{"comment":"The 50-trading-day stock-pick sample and 53-day macro sample are acknowledged as a resource constraint, but the macro regressions in Tables 10–11 have N=53 and several coefficients are significant only at the 10% level. Please state explicitly that these market-level results are exploratory and should be interpreted with caution.","section":"Section 7.B"},{"comment":"The public-signal rows in the overlap sample are based on few posts (median one per stock-date), so the public Net Buy Share often takes discrete values -1, 0, or 1. Reporting coefficients per +10 percentage points is awkward for such a variable; consider a standardized or per-unit scaling, or a supplementary table with robust standardized coefficients.","section":"Table 8"},{"comment":"The 'meaningful' cutoff of 15 points from neutral and the speculation cutoff of 40 are central to several variables. A short sensitivity analysis (e.g., 10/20 and 30/50 cutoffs) would help establish that the main results are not artifacts of these thresholds.","section":"Section 3 / Internet Appendix A.5"},{"comment":"The caption mentions a 'solid red line' marking the actual-data coefficient, but the figure appears in grayscale; please adjust the caption or use a distinct line style.","section":"Figure 2"},{"comment":"The staggered-vintage design is a clear way to address overlapping return windows, but the 'positive and significant' counts across vintages are described as descriptive because vintages share calendar time. Please make this caveat more prominent in the table notes.","section":"Section 5.A.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is original and the real-time archival design is commendable. The missing no-persona control is an omission that is feasible to address and is central to the paper's contribution as a belief-measurement instrument. I do not see grounds for rejection, but the current evidence does not yet establish that the silent-region signal reflects finfluencer personas rather than generic LLM knowledge."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this paper does something genuinely new—runs a real-time, pre-registered panel of digital-twin interviews, archived before return windows begin—and backs it with a serious validation suite. It deserves peer review even though the central attribution claim isn't fully pinned down.\n\nThe innovation is real. Instead of scraping posts or querying an LLM ex post, the authors interview persona-conditioned twins daily, on a fixed protocol, and only then look at returns. That kills the look-ahead bias that plagues ex post LLM studies. The silent region is a genuine measurement gap, and the coverage gain is large: 84.8% of stock-days have no public post. The validation work is careful: 91.5% alignment with public recommendations, a pre-disclosure lean test, and an account re-identification test that shows twin responses are not generic boilerplate. The return predictability survives several placebo distributions, including an industry-restricted one, which is more than most papers do.\n\nThe soft spot is the missing no-persona baseline. The protocol conditions on recent posts and permits web search for ticker-specific information, so a generic LLM given the same ticker lists and real-time market context could plausibly produce buy/sell scores that predict returns. The return regressions never test this. The validation tests soften the concern—the pre-disclosure lean and account-specific structure are inconsistent with pure generic noise—but they don't settle the attribution question. The Table 3 alignment could also be inflated if same-day public posts enter the conditioning context; the paper should state explicitly whether that ever happens. The 50-day sample is short, but the staggered-vintage design and placebos are a reasonable response, and the limitations section is honest.\n\nBottom line: this is a paper for anyone working on social media and belief measurement, and it deserves a serious referee. The missing generic-LLM control is the first thing I'd ask for in a revision, along with timing details and an archive of the interviews and code.","headline":"Real-time digital-twin interviews are a genuinely new measurement instrument with a strong validation suite; the missing generic-LLM baseline keeps the persona-attribution claim short of established.","tokens_in":50131,"tokens_out":2596,"would_cite":true,"duration_ms":29017,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Daily interviews of AI-built digital twins of finfluencers recover stock-level beliefs even when no post is made, and those beliefs predict future returns for large-cap stocks.","keywords":["digital twins","selective disclosure","finfluencers","social media","return predictability","large language models","silent region","belief measurement"],"falsifier":"Run the same daily protocol on the same model with no account conditioning (or with the recent-posts field withheld) and compare silent-region Net Buy Share against the same return windows; if the unconditional LLM signal predicts returns as well as the conditioned one, the persona is not the source of the signal. A second check: re-prompt the model retrospectively on archived pre-return prompts after outcomes are known and see whether the ex-post responses fit returns better than the archived ones, which would indicate residual look-ahead in the protocol.","tokens_in":49298,"feed_emoji":"🤖","tokens_out":6585,"duration_ms":63788,"temperature":0.7,"pith_summary":"Public posts by financial influencers are voluntary disclosures, so the absence of a post is ambiguous: it could mean no view, a weak view, or a strategically withheld view. The paper builds a measurement instrument for this 'silent region' by interviewing 'digital twins' of 81 X accounts—language models conditioned on each account's profile and recent posts—with a fixed daily protocol in real time, archiving every answer before the return windows it will be tested against. The central claim: interview responses contain economically meaningful belief proxies even when nothing is posted, and a simple buy-minus-sell aggregation of the twins' stock recommendations predicts the cross-section of large-cap returns—24 basis points of five-day and 50 basis points of ten-day excess return per ten-percentage-point change in Net Buy Share on the 84.8% of stock-days with no public post. If correct, the method turns a previously unobservable quantity—what a selective discloser would have said when asked—into a comparable, timestamped panel, with a built-in safeguard against the look-ahead bias of retrospective LLM queries.","feed_headline":"AI interviews of finfluencer twins predict stock returns","feed_subtitle":"Daily digital-twin interviews reveal silent-region stock views that forecast large-cap moves up to ten days out.","key_machinery":"The digital twin—an LLM interviewee built from an account's public profile, recent monitored posts, and finance-relevant context—queried daily under a fixed protocol, with a 3:30pm ET cutoff and real-time archival. The workhorse signal is Net Buy Share: the difference between the share of meaningful buy and sell recommendations across the 81 twins for a stock-date. The timed archival is what converts LLM output into a forward-looking instrument.","core_discovery":"The paper claims that standardized, real-time interviews of LLM-based digital twins of finfluencers make the silent region of selective disclosure measurable. Each twin is conditioned on an account's profile and recent monitored posts and is asked a fixed battery of questions about 429 large-cap stocks and seven macro conditions every trading day from December 2025 to March 2026, with responses timestamped under a conservative 3:30pm ET cutoff and archived before any return window begins. Validation against observable public posts shows the interviews align with human recommendations 91.5% of the time (AUC 0.776 vs 0.522 placebo), lean toward recommendations that will only later be posted, a","pith_inferences":["A natural control group is absent from the paper: the same interviews run on the same LLM without any persona conditioning. If that control also predicts returns, part of the silent-region signal could come from the model's out-of-the-box market knowledge rather than from the finfluencer persona. I would expect the authors' next test to run exactly this control.","The contamination risk is concrete and testable: since each twin is conditioned on 'recent monitored posts,' a same-day public recommendation can enter the context and mechanically inflate the 91.5% alignment. A holdout design that conditions twins only on posts older than one trading day would settle this.","The same instrument could be turned on other classes of selective communicators—central banks, executives in quiet periods, political campaigns, or firms' investor-relations personas—wherever publicly disclosed text is strategic and silence is informative. The paper suggests this but does not test it.","If the effect is real, its 10-day horizon and absence at 1 day suggest gradual diffusion of retail-facing beliefs, and the natural next step is to test whether the signal survives when conditioned on voluminous LLM-generated content from thousands of accounts, i.e., whether it is an attention phenomenon."],"forward_implications":["If the measurement claim holds, the silent region of any selectively disclosing agent can be converted into a comparable, time-stamped belief panel; the paper explicitly extends the logic to sell-side analysts, executives, fund managers, journalists, and advisers.","The same panel separates direction, disagreement, and uncertainty: direction predicts returns, informative disagreement predicts higher future volatility and lower returns, and a high uncertainty score predicts lower volatility.","At the market level, average sentiment from the panel does not forecast the S&P 500, but disagreement among twins' macro views does: a 10-point IQR increase predicts 86.4 basis points lower ten-day market returns.","The real-time archival protocol offers a template for avoiding LLM look-ahead bias in any predictive setting: timestamped, fixed-protocol interviews archive beliefs before outcomes."],"supporting_citations":[{"why":"Supplies the cross-sectional regression method used for all return-predictability tests.","marker":"(Fama and MacBeth, 1973)"},{"why":"Provides the autocorrelation-robust standard errors that make overlapping-window inference credible.","marker":"(Newey and West, 1987)"},{"why":"Provides the calendar-time portfolio method used as an alternative estimator for the long-short signal.","marker":"(Mitchell and Stafford, 2000)"},{"why":"Supplies the disagreement framework that predicts lower subsequent returns when views are split, anchoring the polarization result.","marker":"(Miller, 1977)"},{"why":"Documents lookahead bias in pretrained LLMs, motivating the real-time archival design.","marker":"(Sarkar and Vafa, 2024)"},{"why":"Shows LLM return predictability and look-ahead concerns, the benchmark the design avoids.","marker":"(Lopez-Lira and Tang, 2024)"},{"why":"Provides retrieval-augmented conditioning, the mechanism by which twins are built from account material at query time.","marker":"(Lewis et al., 2020)"},{"why":"Instruction tuning that lets the model follow the fixed interview protocol and return structured answers.","marker":"(Ouyang et al., 2022)"},{"why":"Documents finfluencer recommendation behavior, motivating the monitoring universe and the attention-driven posting pattern.","marker":"(Kakhbod et al., 2023)"}],"fun_headline_variants":["Silent finfluencer views surfaced by AI twin interviews","AI twin chats predict large-cap returns from silent posts","Daily AI interviews of finfluencer twins forecast large-cap stocks","Unspoken finfluencer market views captured by AI twins","AI queries finfluencer twins to predict large-cap moves"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The results stand on the assumption that the digital-twin responses are not contaminated by the very public posts they are conditioned on, and that the persona conditioning—not the LLM's own pre-trained market knowledge—is what produces the silent-region signal.","fun_headline_variants_meta":{"raw":{"variants":["Silent finfluencer views surfaced by AI twin interviews","AI twin chats predict large-cap returns from silent posts","Daily AI interviews of finfluencer twins forecast large-cap stocks","Unspoken finfluencer market views captured by AI twins","AI queries finfluencer twins to predict large-cap moves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000682,"raw_usage":{"total_tokens":2888,"prompt_tokens":651,"completion_tokens":2237,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":395,"completion_tokens_details":{"reasoning_tokens":2156}},"tokens_in":395,"tokens_out":2237,"duration_ms":17328,"temperature":1.0,"reasoning_tokens":2156,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:26:37.229522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same daily protocol on the same model with no account conditioning (or with the recent-posts field withheld) and compare silent-region Net Buy Share against the same return windows; if the unconditional LLM signal predicts returns as well as the conditioned one, the persona is not the source of the signal. A second check: re-prompt the model retrospectively on archived pre-return prompts after outcomes are known and see whether the ex-post responses fit returns better than the archived ones, which would indicate residual look-ahead in the protocol.","supporting_citations":[{"cited_title":"West, 1987, A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix, Econometrica\\/ 55, 703--708","cited_arxiv_id":null,"evidence_quote":"Provides the autocorrelation-robust standard errors that make overlapping-window inference credible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the calendar-time portfolio method used as an alternative estimator for the long-short signal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the disagreement framework that predicts lower subsequent returns when views are split, anchoring the polarization result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents lookahead bias in pretrained LLMs, motivating the real-time archival design."},{"cited_title":"u ttler, Mike Lewis, Wen tau Yih, Tim Rockt \\","cited_arxiv_id":null,"evidence_quote":"Provides retrieval-augmented conditioning, the mechanism by which twins are built from account material at query time."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Instruction tuning that lets the model follow the fixed interview protocol and return structured answers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents finfluencer recommendation behavior, motivating the monitoring universe and the attention-driven posting pattern."}],"review_version":1}