{"id":"04260e94-ba86-45b2-9c05-f023a860a03c","arxiv_id":"2506.06329","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper defines a news-count share index and a capitalization-adjusted version, then overstates their validated market-signaling power.","lead":"This paper introduces the Hype Index, a measure of how much financial news attention each S&P 100 company or sector receives relative to its market capitalization. The authors claim these indices can signal short-term market moves, but the paper does not actually run the predictive tests it promises.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract and conclusion promise lagged associations with returns, volatility, and VIX, but Section 5.2 states no direct relationship is computed, so the central claim is unsupported by the manuscript's own evidence.","rationale":"The reader's verdict of REJECT is appropriate, but my load-bearing concern differs from the reader's stated weakest_assumption. The reader focused on LSEG's proprietary entity recognition as the weakest point; I agree that is a real data-quality risk, but it is secondary. The more decisive problem is internal: the abstract and conclusion make specific empirical promises about associations with returns, volatility, and VIX, and about signaling power, while Section 5.2 explicitly disclaims the direct computation. This is not a matter of noisy tagging or an unmodeled confounder; it is a gap between the central claim and the evidence presented. Even a perfect index would not validate the paper's advertised contributions without the promised lagged analysis. The reader's rationale does mention Section 5.2's disclaimer, so there is partial agreement, but the weakest_assumption field understates the issue. A concrete regression test would settle whether the central claim can be rehabilitated; if the hype coefficients are robust, the paper would need substantial rewriting but could have a real result, whereas if they are not, the rejection stands. My recommendation is to keep the reader's REJECT verdict, hence UNCHANGED.","tokens_in":9700,"tokens_out":4326,"duration_ms":44656,"concrete_test":"Recompute the analysis promised in the abstract: build a panel of the 101 tickers over the 326 trading days and regress forward 5-day realized volatility (or absolute return) on lagged daily changes in HypeIndex and CapHypeIndex, including firm or sector fixed effects, lagged volatility, and market controls, with standard errors clustered by date. If the hype coefficients are not statistically significant after controls, or if significance does not survive a simple multiple-testing correction across the 11 sectors, the advertised signaling power is unsupported. As a secondary check, repeat for returns at 1- and 5-day horizons.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central advertised claim is that the Hype Index family provides 'signaling power for short-term market movements' and shows 'associations with returns, volatility, and VIX index at various lags' (Abstract). That claim is not tested anywhere in the manuscript. Section 5.2 states: 'We do not directly compute the relationship between changes in the Hype Index and subsequent market dynamics.' The empirical sections are descriptive: Section 4.2 annotates a few dates around August 2024; Section 4.4 reports that VIX and hype 'appear to move in tandem'; Section 5.1 defines Hype Momentum but never estimates it. The only quantitative links are a scatterplot of news weight vs market weight (Figures 10 and 11) and high correlations between the two indices (Table 4), neither of which establishes signaling power. Thus even if LSEG's ticker tagging and the article-count proxy were perfect, the manuscript would not support its central claim. The condition that must hold for the claim—an actual lagged association between hype changes and market outcomes—is absent from the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines a News Count-Based Hype Index (the share of daily news mentions for each S&P 100 stock or sector) and a Capitalization Adjusted Hype Index (that share divided by the stock's or sector's market-capitalization weight). Using LSEG/Refinitiv news headlines for roughly 101 S&P 100 constituents over 2023–2025, it reports sector-level time series, clusterings into hype groups, normality tests, correlations between the two index variants, and visual comparisons with VIX and GPR. The abstract and conclusion assert that the index family has associations with returns, volatility, and the VIX at various lags, and that it has signaling power for short-term market movements.","tokens_in":9903,"tokens_out":6120,"duration_ms":63395,"significance":"The construction is transparent and the definitions are explicit, with a clear link to a simple, interpretable attention measure. If the advertised lagged associations and signaling power were actually demonstrated, the contribution would be of practical interest for volatility analysis and market monitoring. However, the manuscript as submitted does not compute those associations: Section 5.2 explicitly disclaims the direct test, Section 4.4 is a visual comparison, and Section 4.3 delegates forecasting evidence to a separate paper. The significance of the current version therefore rests on the descriptive index construction rather than on the predictive claims that motivate the abstract and conclusion.","major_comments":[{"comment":"The abstract advertises 'associations with returns, volatility, and VIX index at various lags' and 'signaling power for short-term market movements,' and the Conclusion repeats the association claim. Section 5.2 states, 'We do not directly compute the relationship between changes in the Hype Index and subsequent market dynamics.' No lagged regression, VAR, panel model, or predictive accuracy statistic appears in the empirical sections. This is a load-bearing mismatch: the paper's central advertised result is absent from its own evidence. The authors should either add formal tests (for example, panel regressions of future returns or realized volatility on lagged hype changes with appropriate controls and clustered standard errors) or rewrite the abstract and conclusion to report only the descriptive findings actually presented.","section":"Section 5.2, Abstract, Conclusion"},{"comment":"The claimed association with the VIX is supported only by a visual inspection: the text says the indices 'appear to move in tandem' and the figure shows 7-day rolling means. No correlation coefficient, co-movement statistic, or test of statistical significance is reported for the VIX relationship. As written, the Conclusion's statement that the index 'exhibits meaningful associations with ... market sentiment indicators such as the VIX' is not supported by any quantitative result.","section":"Section 4.4"},{"comment":"Hype Momentum is defined in Definition 5.2 but never operationalized or estimated. The text says its empirical roles are 'further explored in Section 5,' yet Section 5.2 explicitly disclaims a direct computation of hype-to-market relationships. The Conclusion sentence that 'persistent deviations from neutrality often precede significant movements in price or volatility' is therefore unsupported. The authors need to provide an estimator for Hype Momentum and report its empirical performance, or remove the claim.","section":"Section 5.1"},{"comment":"The sentiment forecasting evidence is delegated to a separate paper by the same authors; the current manuscript states that 'the authors of this paper ... have investigated' the relationship in 2025. This means one of the four advertised evaluation lenses, namely signaling power, is not actually evaluated here. If this is intended as background literature, it should be presented as such and removed from the list of evaluations claimed for this paper.","section":"Section 4.3"},{"comment":"The high correlations in Table 4 are largely structural: by Definition 3.1, CapHypeIndex_{i,t} = HypeIndex_{i,t} / MarketCapWeight_{i,t}, and the paper itself notes in Section 4.1 that market-cap weights are relatively stable for large caps. The claim that one index 'can serve as an effective proxy for the other' therefore needs a stronger test of incremental information, for example whether the capitalization-adjusted index explains any variation in future volatility or returns beyond the raw index and a market-cap control.","section":"Section 4.1, Table 4"}],"minor_comments":[{"comment":"The text contains an unresolved editorial instruction: 'Mention what tickers are removed We remove X.TSLA...' The final sample is described as 101 companies, which does not match the stated S&P 100 universe; please clarify the exact list of included and removed tickers.","section":"Section 1.2"},{"comment":"The paper states that around August 5, 2024, the S&P 500 experienced a decline of 'over 10% over a single weekend.' The actual S&P 500 move on that Monday was roughly 3%; please correct the figure or provide a source for the 10% claim.","section":"Section 4.2"},{"comment":"The figure caption contains a typo: 'Hyp' should read 'Hype'.","section":"Figure 9"},{"comment":"The normality tests are reported without the sample size or a clear connection to any downstream modeling choice. If the normality assumption is not used later, the subsection should be shortened or explicitly framed as descriptive.","section":"Section 3.4"},{"comment":"The normalization description is inconsistent: the caption says 'Scaled by Overall Avg = 1,' while the text says 'scaling by daily average enforces...' Please clarify whether the scaling uses the overall sample average or a daily average.","section":"Section 2.2"},{"comment":"The sector Hype Index counts a multi-ticker news item once for each mentioned ticker. This is a legitimate modeling choice, but its effect on sector shares and cross-sector comparisons should be discussed explicitly.","section":"Section 2.1"},{"comment":"The regression in Figure 10 reports p-values of 0.0000 for firm-day observations, but the standard errors are not clustered by firm or date. Given the repeated-observation structure, those p-values are likely overstated.","section":"Section 5.2, Figure 10"}],"recommendation":"major_revision","confidential_remarks":"The paper needs substantial additional empirical work to support the abstract's claims. If the editor prefers an enforce-the-advertised-evaluation standard, a reject-and-resubmit may be more appropriate than a standard major revision. There is also a scope question: the contribution as currently demonstrated is a descriptive index construction, not a statistical study of market signaling; the authors should either reposition the paper accordingly or add the missing tests."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a descriptive attention index that reads like a validated forecasting tool, and that mismatch is the whole story. The raw Hype Index is a news-count share; the adjusted version divides it by market-cap weight. Both are clean definitions, and the data-collection section is honest about the LSEG pipeline and its entity-recognition limitations. But the abstract and conclusion promise associations with returns, volatility, and VIX, and Section 5.2 says outright: 'We do not directly compute the relationship between changes in the Hype Index and subsequent market dynamics.' So the central claim is not tested in the manuscript.\n\nWhat is actually new? Not much. Media-coverage share is a known attention measure—the paper doesn't cite Da, Engelberg, and Gao or the broader investor-attention literature—and the cap-adjusted variant is a mechanical ratio. The high correlations in Table 4 are mostly forced because the denominator (market-cap weight) is stable for mega-caps; the authors even acknowledge this structural relationship. The sector clustering and normality tests are competent descriptive work, but they don't establish signaling power.\n\nThe soft spots are proportionate to this: the conclusion contradicts the methods section, the forecasting claims lean on a self-citation to Cao and Geman (2025), the manuscript contains editorial artifacts ('Mention what tickers are removed,' 101 firms for the S&P 100), and there is no code or data. The LSEG ticker-mapping assumption is a real secondary concern, but even with perfect mapping, the paper still doesn't test its advertised claim.\n\nWho gets value from this? Someone who wants a transparent, count-based attention dashboard for descriptive purposes. As a scientific contribution, it's not ready. I would not bring it to a reading group, I would not cite it, and if I were the editor I would desk-reject or request major revision that actually estimates the lagged hype-market relationships advertised in the abstract.","headline":"The Hype Index is a clean descriptive metric, but the paper's advertised predictive claims are explicitly disclaimed in its own Section 5.2.","tokens_in":10489,"tokens_out":3477,"would_cite":false,"duration_ms":33065,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G80","62P05","91B84"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims a stock's share of financial headlines, divided by its share of market capitalization, reveals when media attention is out of proportion to economic size.","keywords":["Hype Index","media attention","natural language processing","market signaling","stock volatility","S&P 100","capitalization adjustment","investor sentiment"],"falsifier":"Take a random sample of roughly 500 headlines that the LSEG pipeline tagged to S&P 100 tickers and have independent annotators judge whether the tag matches the company actually discussed. If the false-tag rate is high or systematically concentrated in certain sectors, then correcting the tags would move the capitalization-adjusted cluster rankings and the index would fail as an unbiased attention measure.","tokens_in":9445,"feed_emoji":"📈","tokens_out":5833,"duration_ms":58407,"temperature":0.7,"pith_summary":"The paper sets out to show that media 'hype' can be measured directly as the share of financial headlines mentioning a stock or sector, without reading tone. Because large companies naturally appear in the news more often, it constructs a second measure: that headline share divided by the firm's share of S&P 100 market capitalization, so that a value above 1 flags attention out of proportion to economic size. Both indices are built for the S&P 100 over 326 trading days and examined through sector clusters, correlations with returns and volatility, and spikes around events such as the August 2024 market selloff. The payoffs would be a simple, interpretable attention metric for volatility analysis and market signaling.","feed_headline":"One ratio reveals when a stock is overhyped","feed_subtitle":"The Hype Index divides a stock's share of headlines by its market-cap weight to flag attention far from economic size.","key_machinery":"The load-bearing object is the ratio $CapHypeIndex_{i,t} = (N_{i,t}/\\sum_j N_{j,t}) / (MC_{i,t}/\\sum_j MC_{j,t})$. The numerator is a stock's share of daily news mentions; the denominator is its market-capitalization weight within the S&P 100 universe. The ratio converts raw attention into attention per unit of economic size, with 1 as the neutrality benchmark, and the paper reads sustained deviations from 1 as hype or neglect. Sector-level versions are computed by summing constituent ticker hype indices within each GICS sector, and cluster labels are assigned from the resulting trajectories.","core_discovery":"The central claim is that the Hype Index family quantifies attention distortions: $HypeIndex_{i,t} = N_{i,t}/\\sum_j N_{j,t}$ gives the fraction of all S&P 100 news mentions going to stock $i$, and $CapHypeIndex_{i,t}$ divides that fraction by the stock's market-capitalization weight. The paper argues that this ratio is a valid signal of over- or under-hyping and supports it with three empirical observations: the raw index puts Financials and Information Technology at three to four times the average coverage; the capitalization-adjusted version moves Information Technology into the 'less prominent' cluster and pushes Real Estate, Industrials, and Utilities into the 'relatively hyped' cluster; and the two indices are strongly correlated, with sector-level correlations from 0.82 to 0.98. The paper also reports that adjusted hype moves with VIX and GPR changes around stress events and that normality is rejected for the adjusted index and its percent changes.","pith_inferences":["The descriptive event analysis does not by itself prove forecasting power; a strict out-of-sample test asking whether a surprise jump in adjusted hype predicts next-day or next-week volatility would settle that question.","Because the adjusted index is a ratio of two weights, its daily variation is dominated by the news numerator when market caps are stable; the distinct information in the capitalization adjustment is likely concentrated in stress episodes rather than in the daily series.","Counting headlines treats every mention equally; weighting by source reach or combining with sentiment would separate 'hype' from 'information,' a natural extension of the same data pipeline.","The strong raw-adjusted correlation suggests the index would be most useful as an attention-risk factor for volatility modeling or event detection, not as a standalone directional return signal."],"forward_implications":["Financials and Information Technology receive three to four times the market-average news share over the sample, while Utilities, Real Estate, and Materials receive less than half the average.","Adjusting for market capitalization reverses the picture: Information Technology becomes less prominent relative to its size, while Real Estate, Industrials, and Utilities look relatively hyped.","Because sector-level correlations between the raw and capitalization-adjusted indices run from 0.82 to 0.98, raw news share can proxy for the adjusted measure whenever market-capitalization weights are slow-moving.","Spikes and troughs in capitalization-adjusted hype concentrate around identified market events, including the August 2024 selloff, the November 2024 election rally, and the April 2025 tariff shock.","The normality tests reject a normal model for the adjusted index and its percent changes, which matters for any later statistical use of the index."],"supporting_citations":[{"why":"Supplies the prior hype-adjusted probability measure and the sentiment-forecasting result (75% precision) that this paper's hype framework extends.","marker":"[2]"},{"why":"Provides the template of converting unstructured media content into a structured index relevant for financial decision-making.","marker":"[1]"},{"why":"Establishes the foundational empirical link between media tone and market behavior that this volume-based measure complements.","marker":"[8]"},{"why":"Shows that unusual news content forecasts market stress, motivating attention intensity rather than only sentiment as a signal.","marker":"[5]"},{"why":"Demonstrates that sentiment and topic features from headlines can predict next-day volatility, a baseline against which hype measures are positioned.","marker":"[4]"},{"why":"Defines the Cboe VIX methodology used as the fear-gauge benchmark in the comparison of hype indices with market uncertainty.","marker":"[9]"},{"why":"Supplies the behavioral overreaction premise that media-driven attention could precede price and volatility moves.","marker":"[3]"}],"fun_headline_variants":["Hype Index: News share vs. market cap flags overhyped stocks","NLP-driven Hype Index measures media attention vs. economic size","Ratio of news to market weight signals when a stock is overhyped","Attention distortion metric: Hype Index reveals overexposure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The index inherits the accuracy of LSEG/Refinitiv's proprietary entity-recognition tags: if a headline is mapped to the wrong ticker, or if a bare article count assigns equal weight to a one-word mention and a full analysis, every hype value, cluster, and correlation inherits that error.","fun_headline_variants_meta":{"raw":{"variants":["Hype Index: News share vs. market cap flags overhyped stocks","NLP-driven Hype Index measures media attention vs. economic size","Ratio of news to market weight signals when a stock is overhyped","Attention distortion metric: Hype Index reveals overexposure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1314,"prompt_tokens":948,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":290}},"tokens_in":564,"tokens_out":366,"duration_ms":4620,"temperature":1.0,"reasoning_tokens":290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:10:49.168868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly 500 headlines that the LSEG pipeline tagged to S&P 100 tickers and have independent annotators judge whether the tag matches the company actually discussed. If the false-tag rate is high or systematically concentrated in certain sectors, then correcting the tags would move the capitalization-adjusted cluster rankings and the index would fail as an unbiased attention measure.","supporting_citations":[{"cited_title":"A Hype-Adjusted Probability Measure for NLP Stock Return Forecasting","cited_arxiv_id":null,"evidence_quote":"Supplies the prior hype-adjusted probability measure and the sentiment-forecasting result (75% precision) that this paper's hype framework extends."},{"cited_title":"Measuring Geopolitical Risk","cited_arxiv_id":null,"evidence_quote":"Provides the template of converting unstructured media content into a structured index relevant for financial decision-making."},{"cited_title":"Does Unusual News Forecast Market Stress?","cited_arxiv_id":null,"evidence_quote":"Shows that unusual news content forecasts market stress, motivating attention intensity rather than only sentiment as a signal."},{"cited_title":"A sentiment analysis approach to the prediction of market volatil- ity","cited_arxiv_id":null,"evidence_quote":"Demonstrates that sentiment and topic features from headlines can predict next-day volatility, a baseline against which hype measures are positioned."},{"cited_title":"Does the Stock Market Overreact?","cited_arxiv_id":null,"evidence_quote":"Supplies the behavioral overreaction premise that media-driven attention could precede price and volatility moves."}],"review_version":1}