{"id":"7a374601-0ad1-4754-bc1e-216d7c068c97","arxiv_id":"2608.01322","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM-based similarity of 10-K filings does not predict peer-firm stock reactions on acquisition announcements: pooled rank correlation +0.07, p=0.37.","lead":"This paper tests whether a large language model can identify 'shadow trading' peer firms from public SEC filings, comparing 10-K text similarity to stock price reactions across 30 acquisitions. It finds no reliable signal: similar filings do not predict which peers' stocks move, despite the model matching the known target in one famous case.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated single-shot LLM similarity scores: measurement error alone could explain the null, so the claim that this pipeline finds no association is not yet interpretable.","rationale":"The paper's central claim is the null association between LLM-derived 10-K similarity and announcement-day abnormal returns, interpreted as bearing on the SEC's ex ante identifiability premise. The most load-bearing condition for that claim is that the similarity scores are meaningful, reliable measurements. The paper itself admits this condition is unmet: each event was scored once, no stability or agreement data are reported, and the model is closed-weight. Measurement error in the independent variable is not a minor caveat; it directly attenuates rank correlations, so the observed rho approx 0 could arise even if a stable, human-validated similarity signal would predict returns. This is distinct from, and more fundamental than, the outcome-proxy concern raised by the reader: even if announcement-day AR were a perfect measure of economic linkage, an unreliable similarity score would still produce a null. The paper's candid limitation statements strengthen our confidence in its narrow mechanical claims (the statistics are computed correctly), but they also underscore that the broader conclusion about shadow trading administrability depends on a score reliability assumption that is untested. The proposed repeated-runs experiment would settle whether the null is a measurement artifact; if scores are highly unstable, the paper's headline result cannot be interpreted, though the verdict remains conditional rather than reject because the authors have explicitly scoped their claim and identified the stability study as the needed next step. We therefore agree partially with the reader's assessment: the reader focused on the dependent-variable proxy, while we identify the independent-variable measurement validity as the more load-bearing soft spot, but both point to the same conclusion that the paper's broader claim is not yet fully supported.","tokens_in":1036,"tokens_out":2190,"duration_ms":52323,"concrete_test":"Rerun Stage 2 (document-grounded ranking) 10 times per event for all 30 events using the identical prompts, rubric, and configuration (Gemini 3.1 Pro, Thinking Level High, top-P 0.90) in isolated sessions. Compute the intraclass correlation ICC(2,1) of similarity scores across runs and the standard deviation of the per-event Spearman correlations across runs. If the mean per-event correlation varies by more than +/- 0.2 across runs or ICC(2,1) < 0.7, the observed rho_pooled = +0.066 is consistent with pure measurement noise, and the null cannot be interpreted as evidence against ex ante identifiability. If scores are stable (ICC >= 0.7 and per-event rho SD < 0.1), the measurement-error concern is resolved and the reader's condition on baselines remains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 and Section 7 disclose that each event's Stage 2 similarity scores were produced once, by a single closed-weight model (Gemini 3.1 Pro), with no run-to-run variance, no prompt/rubric sensitivity analysis, and no agreement check against human judgment. Section 4.5 nevertheless treats these scores as exact ranks when computing Spearman correlations and the pooled rho=+0.066. If the scores contain substantial random measurement error, the observed correlations are attenuated toward zero regardless of whether a true association exists. The authors themselves say the scores are 'structured qualitative judgments rather than deterministic measurements' (Section 7) and identify a stability study as the first addition they would make. Without such a study, the headline null (rho approx 0, 95% CI [-0.08,+0.18]) is equally consistent with (a) genuinely absent linkage, (b) an LLM whose similarity judgments are internally inconsistent, or (c) a rubric so prompt-sensitive that the pipeline is essentially a noisy measurement instrument. This concern precedes the outcome-proxy problem: it bears on the internal validity of the independent variable, not just on interpretation of the dependent variable. The reader's weakest assumption (announcement-day AR as proxy for economic linkage) is important, but a noisy regressor can manufacture the null even when the proxy is perfect. Thus the paper's central claim that 'this pipeline does not recover economic linkage' is not yet supported until the scores are shown to be reliable measures.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether semantic similarity of 10-K MD&A text, computed by a two-stage Gemini 3.1 Pro pipeline, can identify economically linked peer firms ex ante, as the SEC's shadow-trading theory in SEC v. Panuwat presumes. The authors construct 30 M&A events (217 peer observations) across five industries, score each target's peer candidates with an LLM rubric (40/30/30 weights), compute announcement-day abnormal returns against sector ETFs, and report a pooled within-event Spearman correlation of +0.066 (permutation p=0.37) and a mean per-event Spearman correlation of +0.046 with 95% CI [-0.084, +0.175]. They interpret the CI as excluding any moderate association, report a case-level reading of 14/30 supportive, and provide a correction showing Incyte fell outside standard mid-cap definitions in 2016. The paper is explicitly exploratory and repeatedly scopes the claim to this pipeline, corpus, and return measure, while drawing broader implications for fair notice and CAT surveillance.","tokens_in":38969,"tokens_out":5222,"duration_ms":56489,"significance":"If the headline null were supported by validated similarity measurements and appropriate baselines, this would be a novel and important empirical input to live legal debates about shadow trading, fair notice, and the constitutionality of mass market surveillance. The paper's strengths are real: full prompts and per-event outputs are reproduced, the permutation tests are exact for small events, the mid-cap market-capitalization correction is a concrete factual contribution, and the limitations section is unusually candid. However, the central inference is currently blocked by an unvalidated measurement instrument: single-shot, closed-weight LLM scores with no variance, no prompt sensitivity analysis, and no human agreement check. Absent a reliability study, the reported null cannot be distinguished from attenuation caused by measurement error. The paper's own Section 7 identifies this and the absence of baselines as severe gaps. The contribution is therefore better framed as a documented null result for this specific pipeline than as evidence against the SEC theory's empirical premise.","major_comments":[{"comment":"The headline correlation analysis treats one-shot Gemini 3.1 Pro similarity scores as exact values. Section 3.3 reports a single run per event, and Section 7 concedes there is no run-to-run variance, no prompt/rubric sensitivity analysis, no second model, and no agreement with human expert judgments, describing the scores as 'structured qualitative judgments rather than deterministic measurements.' Classical measurement error in the regressor attenuates Spearman correlations toward zero. Therefore the observed rho=+0.066 and the CI [-0.084, +0.175] bound the association between this particular set of LLM outputs and returns, not the association between latent filing similarity and returns. Attenuation alone can produce the reported null even if a true association exists. The Section 4.5 claim that the CI 'excludes any moderate relationship' is not yet supported. A stability study, repeat","section":"§3.3, §4.5, §7"},{"comment":"The paper's wider relevance to the shadow-trading debate depends on showing the null is not an artifact of this one pipeline. Section 7 itself calls the absence of baselines 'the most consequential gap in the study' and lists three alternative explanations: filings do not encode linkage, the LLM pipeline fails to extract it, or announcement-day abnormal returns are too noisy. Without TNIC-style bag-of-words, embedding, SIC/GICS, or random-peer comparisons on the same return measure, the paper cannot distinguish these. Because the LLM is the only text measure tested, the conclusion should be narrowed to 'this pipeline yields no signal' rather than implying broader pressure on the SEC theory. Adding even one simple baseline would materially change the interpretability of the null.","section":"§7 (No baselines)"},{"comment":"The event set was generated by prompting Gemini 3.1 Pro and was not pre-registered or drawn mechanically from a deal database, as Section 3.1 discloses. Stage 1 peer discovery also ran with Google Search grounding enabled and without logging retrieved sources, so the claimed filings-only ex ante condition is not cleanly tested. The authors acknowledge parametric recall of widely reported deals, which is especially problematic for the Panuwat sanity check (Section 4.1). These issues bear directly on the external-validity claim that 'an insider could not determine from public disclosures which securities are off-limits.' A mechanical sample of deals from a database and a filings-only peer-discovery run with logged sources would be needed to make that claim load-bearing.","section":"§3.1, §3.2, §7 (event selection and peer discovery)"},{"comment":"The paper equates 'economic linkage' with same-day abnormal return benchmarked to a sector ETF, and Section 7 concedes the expected sign is deal-dependent and that effects could appear over longer windows or through options markets. This is not a minor caveat: the Panuwat case itself involved call options, and the SEC's own expert relied on announcement-day price movement. If the true economic linkage manifests in options volumes or multi-day windows, the current null says nothing about identifiability. The abstract and conclusions partially acknowledge this, but the broader statement that the results 'put pressure on the empirical premise of shadow trading enforcement' remains too strong without either an options-market outcome test or an explicit multi-day-window robustness check.","section":"§3.4, §7 (outcome proxy)"}],"minor_comments":[{"comment":"The sector is labeled 'Automotives' in the tables but 'automotive' in Section 3.1; please standardize.","section":"Tables 4, 5, 6"},{"comment":"The footnote reports 318 peer-event rows, while Section 4.5 and Table 4 use 217 peer observations. The relationship between these numbers (e.g., initial candidates vs. final usable sample) should be stated explicitly to avoid an apparent inconsistency.","section":"§3.4, footnote 1"},{"comment":"The trendline arrows use the opposite sign convention from the Spearman rho because rank runs opposite to score. The explanation appears only in Section 4.5; a one-sentence footnote to Table 4 would prevent reader confusion.","section":"Table 4 / §4.2 / §4.5"},{"comment":"The supports/contradicts/mixed labels are assigned after inspecting returns with researcher degrees of freedom. The paper is transparent about this and does not rest the main claim on the labels, but the phrase 'we state plainly that step (iii) involves judgment' is buried; consider moving this caveat immediately before Table 4.","section":"§4.2"},{"comment":"The ticker 'FRBKQ' appears in the abnormal returns table for WSFS/Beneficial; verify that this is a deliberate ticker for a delisted entity and not a typo.","section":"Appendix D.5.5"}],"recommendation":"major_revision","confidential_remarks":"I agree with the reader's conditional assessment. The mechanical rank-correlation analysis is honest and the appendices are valuable, but the central claim is currently undercut by the unvalidated LLM scores and the absence of baselines. This is fixable within the manuscript's scope through additional experiments (stability runs, a baseline comparator, and ideally an options-window outcome), which is why I recommend major revision rather than rejection. The paper's legal-policy framing is appropriate for an interdisciplinary venue, though for a core CS audience the methodological validation would need to be substantially strengthened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time as a cautionary example, not as evidence. The paper is the first to operationalize the SEC's shadow-trading 'economic linkage' standard with long-context LLM similarity on 10-K MD&A text and test it against announcement-day returns. That's a genuinely new application, and the dataset of 30 M&A events with per-event scores and returns is a useful resource. The authors also do several things right: the rank-correlation analysis is mechanical and clean, with permutation tests and a confidence interval; the Panuwat sanity check is presented with a caveat about pretraining contamination; and the Limitations section is unusually candid.\n\nBut the central null cannot be interpreted as it stands. Each event's similarity scores were produced once by one closed-weight model, with no run-to-run variance, no prompt-sensitivity analysis, no second model, and no human agreement check. The authors explicitly call the scores 'structured qualitative judgments' and list a stability study as the first thing they would add. Until that exists, the observed rho of +0.05 (CI [-0.08, +0.18]) is as consistent with a noisy measurement instrument as with a genuinely absent linkage. Random measurement error in the regressor attenuates rank correlations toward zero; the narrow CI is on a possibly biased point estimate, so it does not 'exclude a moderate relationship.' The no-baselines gap is the other load-bearing weakness: without TNIC, SIC, embedding, or random-peer comparisons, you can't tell whether the disclosure text lacks the signal, this pipeline fails to extract it, or the return proxy is too noisy. The authors concede both points. The outcome-proxy problem is real but secondary; the measurement problem alone would be enough to invalidate the broader claim.\n\nIf I were editing, I would send it to review — the legal stakes and novelty justify referee time — but I would frame the revision around measuring the instrument and adding baselines. As submitted, it is a well-documented pilot with an uninterpretable headline. Take it to a reading group as a case study in why LLM-as-measure needs validation, not as a finding about shadow trading.","headline":"A model null is honestly reported but not yet interpretable: single-shot unvalidated LLM scores could explain the zero correlation by themselves.","tokens_in":39374,"tokens_out":2294,"would_cite":false,"duration_ms":25686,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 30 M&A deals and 217 peers, 10-K similarity shows no link to announcement-day returns, undercutting the SEC's shadow trading premise.","keywords":["shadow trading","insider trading","SEC v. Panuwat","economic linkage","10-K MD&A","LLM semantic similarity","announcement-day abnormal returns","Consolidated Audit Trail"],"falsifier":"Re-run the same 30 events with an open-weight long-context model and a TNIC-style bag-of-words baseline, using five-day cumulative abnormal returns benchmarked to firm-specific market models; if either baseline yields a mean per-event Spearman correlation whose confidence interval excludes zero and exceeds the paper's upper bound of $\\rho \\approx 0.18$, the claim that filing similarity cannot flag economically linked peers is overturned.","tokens_in":38527,"feed_emoji":"⚖️","tokens_out":25801,"duration_ms":194765,"temperature":0.7,"pith_summary":"Shadow trading is the SEC's theory that an insider holding confidential news about one company is barred from trading in any 'economically linked' peer's securities, first prosecuted in SEC v. Panuwat (2023). The paper asks whether natural-language analysis of the disclosures an insider could actually read — the Management's Discussion and Analysis (MD&A, Item 7) sections of 10-K filings — can identify those linked peers in advance. It operationalizes 'economic linkage' as semantic similarity scored by a two-stage long-context LLM pipeline and tests whether those scores predict announcement-day abnormal returns across 217 peer observations and 30 M&A events in five industries. The result is null: pooled within-event rank correlation $+0.07$ (permutation $p = 0.37$), mean per-event Spearman correlation $+0.05$ with a 95% confidence interval of $[-0.08, +0.18]$ — narrow enough to exclude any moderate association rather than merely failing to detect one — with a case-level reading of 14 supportive, 12 contradictory, and 4 mixed events. A sympathetic reader cares because the same ex ante identifiability premise is what justifies the SEC's market-wide Consolidated Audit Trail; if it fails, the fair-notice and constitutional challenges to that surveillance infrastructure gain empirical footing, and the authors are explicit that the null is bound to this pipeline, corpus, and return measure.","feed_headline":"217 peers show no 10-K signal for shadow trading targets","feed_subtitle":"Across 30 M&A deals, 10-K similarity and peer returns barely correlate, challenging the SEC's shadow-trading premise.","key_machinery":"The load-bearing machinery is a two-stage LLM pipeline that converts the SEC's 'economic linkage' doctrine into a measurable quantity. Stage 1 gives the long-context model (Gemini 3.1 Pro) only the target's Item 7 (MD&A) text from its latest pre-announcement 10-K and asks for ten comparable public peers. Stage 2 loads all peers' Item 7 sections into one prompt and scores each peer 0–1 on a fixed rubric (technical/product overlap 40%, financial stage 30%, risk/market exposure 30%), weighting technical overlap to mirror the Panuwat testimony. Scores are paired with announcement-day abnormal returns benchmarked to sector ETFs, and the verdict rests on the within-event Spearman rank correlation","core_discovery":"Central discovery: 'economic linkage,' operationalized as filing-text similarity, does not identify the firms that react to a deal announcement. On the Panuwat fact pattern the pipeline recovers Incyte as the third-closest peer (+5.04% return), a sanity check the authors qualify: the case is widely reported and that event's correlation is slightly negative ($\\rho = -0.10$). Across 30 events the association is absent: pooled within-event rank correlation $+0.07$ ($p = 0.37$), mean per-event Spearman $+0.05$ with 95% CI $[-0.08, +0.18]$, excluding a moderate association. It also corrects the record: Incyte's pre-announcement capitalization was 14.3B, outside the 2B–10B mid-cap band the 'mid-ca","pith_inferences":["The null may indict the outcome measure as much as the text: testing longer windows, options-implied signals, or firm-specific beta benchmarks could recover an association this design shades, and that is the most direct next experiment rather than a refutation.","The paper's suggested inversion — start from the largest abnormal movers around each announcement, then ask which were publicly knowable in advance — is the sharper test of the doctrine because it removes the pipeline's peer-discovery stage from the chain of inference.","If the missing baselines (TNIC bag-of-words, SIC/GICS peers, embedding similarity) all reproduce the null, the finding would generalize from 'this pipeline fails' to 'public filings do not encode economic linkage,' a much stronger statement against the doctrine that the authors deliberately did not run.","The sector asymmetry — zero supporting cases in Technology, zero contradictions in Automotive — is suggestive enough to motivate a targeted study of supplier-chain industries, where spillover co-movement may be text-detectable in a way platform industries are not; five events per sector is too few to conclude anything."],"forward_implications":["If the null holds, the fair-notice premise of shadow trading liability fails: an insider of ordinary intelligence could not have determined from public filings which peer securities were off-limits before trading.","The confidence interval bounds any real association below roughly 0.18, far weaker than an enforcement heuristic would need; a standard right about as often as it is wrong cannot justify targeted watchlists replacing market-wide surveillance.","Because a linkage argument is cheap to construct once the price move is known, the post hoc / ex ante asymmetry survives the SEC's January 2026 CAT amendments, which removed names and taxpayer identifiers but not the search-first, identify-later architecture.","The mid-cap correction is independent of the NLP result: under every mainstream mid-cap definition in force in 2016, Incyte at 14.3B was not mid-cap, so the category could not have put an insider on notice ex ante."],"supporting_citations":[{"why":"Supplies the text-based network industry classification framework the pipeline augments; its TNIC logic is the prior evidence that filing similarity captures competitive structure.","marker":"Hoberg & Phillips, 2016"},{"why":"The SEC v. Panuwat opinion that defines 'economic linkage' and 'mid-cap oncology,' the legal standard the paper operationalizes and tests.","marker":"United States District Court, N.D. Cal., 2023"},{"why":"Introduced 'shadow trading' in the economics literature and estimated per-event profits of $139,400–$678,000, setting the stakes the paper evaluates.","marker":"Mehta et al., 2021"},{"why":"Establishes document-scale comparison of 10-K pairs as an existing task, grounding the paper's claim that the architecture is not the novelty.","marker":"Koval et al., 2024"},{"why":"Supplies the fair-notice / due-process doctrine that makes ex ante identifiability the legal hinge of the paper's argument.","marker":"Supreme Court of the United States, 2018b"},{"why":"Documents abnormal options volume in roughly 25% of U.S. takeovers, quantifying the shadow-trading enforcement problem the theory addresses.","marker":"Augustin et al., 2019"},{"why":"The Davidson v. Atkins complaint challenging the Consolidated Audit Trail's constitutionality, the litigation the results bear on.","marker":"New Civil Liberties Alliance, 2024"},{"why":"Demonstrates that pseudonymized records are re-identifiable, supporting the paper's argument that CAT's anonymization amendments leave the surveillance defect intact.","marker":"de Montjoye et al., 2013"}],"fun_headline_variants":["10-K similarity shows zero edge for shadow trading targets","NLP pipeline finds no peer signal in 10-K text","Shadow-trading peer prediction fails with 10-K embeddings","No link between 10-K similarity and shadow-trading returns","217 peers, zero correlation: 10-K text can't spot shadow trades"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The result rests on treating a one-day abnormal return, measured against a sector ETF, as an adequate stand-in for the legal construct of 'economic linkage'; the paper concedes (Section 7) that if linked firms react on longer horizons, through options, or in ways a sector benchmark masks, the null indicts the proxy rather than identifiability.","fun_headline_variants_meta":{"raw":{"variants":["10-K similarity shows zero edge for shadow trading targets","NLP pipeline finds no peer signal in 10-K text","Shadow-trading peer prediction fails with 10-K embeddings","No link between 10-K similarity and shadow-trading returns","217 peers, zero correlation: 10-K text can't spot shadow trades"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2650,"prompt_tokens":939,"completion_tokens":1711,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":1625}},"tokens_in":683,"tokens_out":1711,"duration_ms":12941,"temperature":1.0,"reasoning_tokens":1625,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:18:18.963152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 30 events with an open-weight long-context model and a TNIC-style bag-of-words baseline, using five-day cumulative abnormal returns benchmarked to firm-specific market models; if either baseline yields a mean per-event Spearman correlation whose confidence interval excludes zero and exceeds the paper's upper bound of $\\rho \\approx 0.18$, the claim that filing similarity cannot flag economically linked peers is overturned.","supporting_citations":[{"cited_title":"Journal of Political Economy , volume =","cited_arxiv_id":null,"evidence_quote":"Supplies the text-based network industry classification framework the pipeline augments; its TNIC logic is the prior evidence that filing similarity captures competitive structure."},{"cited_title":"and Reeb, David M","cited_arxiv_id":null,"evidence_quote":"Introduced 'shadow trading' in the economics literature and estimated per-event profits of $139,400–$678,000, setting the stakes the paper evaluates."},{"cited_title":"Findings of the Association for Computational Linguistics: EACL 2024 , pages =","cited_arxiv_id":null,"evidence_quote":"Establishes document-scale comparison of 10-K pairs as an existing task, grounding the paper's claim that the architecture is not the novelty."},{"cited_title":"and Verleysen, Michel and Blondel, Vincent D","cited_arxiv_id":null,"evidence_quote":"Demonstrates that pseudonymized records are re-identifiable, supporting the paper's argument that CAT's anonymization amendments leave the surveillance defect intact."}],"review_version":1}