{"id":"f0b1c802-0155-46e7-ae48-67060ccc79fb","arxiv_id":"2608.04200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"QLoRA tuning improves financial sentiment classification, but no sentiment model shows robust next-session stock return predictability in the 2019 Benzinga sample.","lead":"Seven financial sentiment classifiers, from Naive Bayes to QLoRA-tuned large language models, are benchmarked on human labels and then tested on whether their scores predict stock returns. Fine-tuning improved accuracy on labels, but none of the classifiers produced statistically reliable next-session return forecasts after multiple-testing correction.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The economic-null half of the central claim depends on the next-session open alignment of Eq. (5); if Benzinga headlines move prices within the same session, the 'clear gap' is an artifact of the timing rule, not a documented property of the sentiment scores.","rationale":"The paper's strongest claim has two parts: QLoRA adaptation improves financial sentiment classification, and none of the sentiment models yields return rankings that survive proper inference, thereby documenting a gap. The first part is reasonably supported by a same-backbone zero-shot/QLoRA gain and consistent gains across three 7-8B backbones, although the comparison changes the inference head and there is no code or uncertainty quantification. The second part is where the central claim is least secure. Equation (5) deliberately discards same-session returns, and Section 6.1 concedes that liquid large-cap news may be priced within minutes or hours. If that is true, the one-day ICs measure only the residual after the main adjustment, so their FDR-cleared insignificance cannot establish a 'clear gap.' This is not an internal contradiction; the paper's limitations are unusually candid. But the abstract states the gap as a finding rather than as a design-dependent null. A timestamped intraday re-test is the decisive check: it separates 'stale-signal design' from 'no tradable signal.' The reader's weakest_assumption identifies the same alignment rule, so I agree. The verdict should remain CONDITIONAL because the checks that would resolve the concern—intraday timestamps, code or data release, and uncertainty quantification on classification metrics—are not in the manuscript.","tokens_in":22016,"tokens_out":4647,"duration_ms":45429,"concrete_test":"Re-run Experiment 2 on Benzinga headlines with reliable publication timestamps using event-time alignment: entry at the price immediately preceding publication, with returns over 5 minutes, 30 minutes, 1 hour, 4 hours, and same-session close; apply the same within-model ranking, Newey-West inference, and FDR correction. If any intraday horizon yields significant rank IC or portfolio return for the same seven models, the Section 5.2 null is an artifact of next-session alignment. If all intraday horizons remain insignificant, the gap claim is supported rather than merely assumed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.5.2's return alignment (Eq. 5) maps every calendar-date signal to the first trading session strictly after date d, with entry at that session's open. For a headline published during or before a trading session on date d, same-session price discovery is excluded entirely; even pre-market news waits until the next day's open. The abstract's second claim—'a clear gap between classification accuracy and tradable cross-sectional signals'—therefore rests on a zero IC being measured after the interval in which any signal would have been tradeable. Section 6.1 acknowledges this timing mismatch but does not test it, while the paper's own Future Work section identifies the relevant intraday alignment. With roughly 52 fresh-news stock-dates per day and a single year, the rank-IC tests also carry limited power, so the FDR-cleared null is consistent either with no signal or with a signal that expired before entry. The concern is not that the calculations are internally wrong; it is that the design cannot distinguish 'no tradable information' from 'information already traded before the measurement window,' and the central claim states the former as a documented gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports two experiments. Experiment 1 constructs a harmonized three-class benchmark from five financial text datasets and compares TF–IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B, using a fixed train–validation–test split. Experiment 2 applies the seven probability-producing classifiers to 13,115 headline–stock observations for a fixed S&P 100 universe in 2019, converting model probabilities to continuous sentiment scores, aggregating by stock–date, and aligning them with next-session open returns over horizons of 1, 2, 3, and 5 days. The classification results show large gains from QLoRA relative to zero-shot Qwen (macro-F1 from 0.7274 to 0.8615), with Mistral the best at 0.8771. The downstream results show positive but small one-day rank ICs (largest 0.0143) and no test significant after Newey–West and FDR correction; portfolio returns are mostly weak or attributable to market exposure. The paper concludes there is a gap between classification accuracy and tradable cross-sectional signals.","tokens_in":22285,"tokens_out":7844,"duration_ms":60827,"significance":"If the results hold, the paper makes a useful empirical contribution: it provides a unified multi-dataset benchmark for financial sentiment classification with a clean same-backbone QLoRA comparison, and it demonstrates a careful evaluation protocol for linking sentiment scores to returns, including horizon-aware alignment, Newey–West inference, and FDR correction. The paper's explicit discussion of limitations (timing mismatch, small sample, contamination risk) is a strength. However, the strength of the economic conclusion currently exceeds what the evidence supports, because the null result is measured after the interval in which the signal may have already been traded.","major_comments":[{"comment":"The alignment rule maps every calendar-date signal to the first trading session strictly after date d, with entry at that session's open. The paper acknowledges in Section 6.1 that public information may be incorporated into large-cap equity prices within minutes or hours; the chosen alignment therefore measures the residual signal after the primary price adjustment has occurred. Because the abstract's central claim asserts a 'clear gap between classification accuracy and tradable cross-sectional signals', and this null result is consistent with a signal that expired before the measurement window, the gap is not documented by the current design. Please provide a sensitivity analysis with same-session or intraday alignment (e.g., using timestamped headlines for a subset) or revise the abstract and conclusion to characterize the result as 'no robust evidence of next-session predictability under the present alignment'.","section":"Section 3.5.2, Eq. (5)"},{"comment":"The downstream sample consists of one calendar year, 72 S&P 100 constituents with usable headlines, and 13,115 headline–stock observations over 253 dates. The paper notes that some daily cross-sections contain only a small number of stocks with fresh news, reducing power. With this limited power, the FDR-cleared null (minimum adjusted q-value 0.9622) is consistent both with a genuine absence of signal and with a weak signal that the test cannot detect. The paper should report a power analysis or an upper bound on the effect sizes that could be detected, and it should temper the 'clear gap' language in the abstract accordingly.","section":"Section 6.1"},{"comment":"The paper acknowledges that the pretrained models may have seen the 2019 Benzinga headlines in their training corpora, so the downstream evaluation is not strictly out-of-sample with respect to model weights. The abstract's phrase 'temporally separate' is correct but does not rule out contamination. Please attempt an empirical check (e.g., comparing model predictions on held-out vs. memorized headlines, or probing for exact-match recall) or explicitly qualify the 'out-of-sample' claim throughout the manuscript.","section":"Section 6.1"}],"minor_comments":[{"comment":"The phrase 'clear gap between classification accuracy and tradable cross-sectional signals' is stronger than the paper's own conclusion in Section 6.1, which says the findings should be interpreted as evidence about the present experimental design rather than a general rejection of the economic value of sentiment. Please align the abstract with this qualification.","section":"Abstract"},{"comment":"In the sentence 'Words such asliability,capital, andtaxare frequently labeled as negative', there are missing spaces after 'as' and 'and'; please correct.","section":"Section 2.1"},{"comment":"The entry for Twitter Financial News Sentiment lists its year as '–' and gives no sample size; please complete the entry or note that the year is unknown.","section":"Table 1"},{"comment":"FinBERT and Financial-RoBERTa each appear in two rows ('Pretrained classification head' and 'Off-the-shelf sentiment checkpoint'), but the text says they are evaluated using the same publicly available checkpoints without fine-tuning. This duplication is confusing; please merge the rows or clarify the distinction.","section":"Table 3"},{"comment":"The classification comparisons are based on a single random seed (seed 42) without confidence intervals or repeated runs. The large zero-shot vs. QLoRA gap is clearly robust, but the small differences among Mistral, LLaMA3, and the two Qwen variants (macro-F1 0.8771 vs 0.8753 vs 0.8615) should not be over-interpreted. Reporting variance across several seeds would make the ranking claims more reliable.","section":"Section 3.4"},{"comment":"The manuscript does not include code or data. Given the paper's benchmarking contribution, a reproducibility statement or link to the harmonized benchmark and evaluation scripts would strengthen the paper.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical study with a clean classification benchmark and a carefully hedged economic evaluation. The main issue is that the abstract's 'clear gap' claim is not supported by the current alignment rule; the paper itself identifies the timing mismatch as a limitation. With a sensitivity analysis or a softened claim, this could be a good paper. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper in one line: a clean QLoRA sentiment benchmark followed by a null return-predictability test, both executed carefully and hedged honestly. The classification half is the stronger contribution. A unified five-dataset three-class protocol, fixed splits, same-backbone zero-shot versus QLoRA, three backbones, and a loss ablation is genuinely useful for anyone building financial sentiment classifiers. Mistral and LLaMA both land around 0.88 macro-F1; QLoRA lifts Qwen by 13 points. Those numbers will be cited.\n\nThe economic half is competent but weaker than the abstract implies. The design is disciplined: fresh-news-only cross-sections, rank IC, Newey–West, FDR, overlapping cohorts. The null—nothing survives correction—is believable under the stated rule that signals enter at the next session's open. But that rule is exactly the soft spot. If Benzinga headlines move prices intraday, a next-session open test can show zero even when the sentiment scores carry real information. The paper acknowledges this in Section 6.1 and even lists intraday horizons in Future Work, so it is not hidden. Still, calling the result a 'clear gap between classification accuracy and tradable cross-sectional signals' overstates what the design can establish. It is a gap at daily frequency, conditional on next-session entry, not a general property of the scores.\n\nOther soft spots are minor individually but add up: single-seed classification with no confidence intervals, no code or data release, a one-year downstream sample with only 72 names, and pretraining contamination that cannot be excluded. The portfolio results are gross and not risk-adjusted, which the paper says plainly. None of these is fatal; they are standard limitations for a first benchmark of this kind.\n\nWho gets value: people working on financial LLMs and news-driven strategies will want this as a reference point and a cautionary data point. It deserves a serious referee, not a desk reject. The main revision asks are reproducibility—ship the code and data or at least detailed training logs—and either multi-seed uncertainty or a softening of the abstract's claim about the gap. If those are addressed, this becomes a solid citable benchmark.","headline":"A careful, honestly hedged benchmark paper: the QLoRA classification results are solid and useful, while the economic null is credible under its next-session design but weaker than the abstract's 'clear gap' phrasing suggests.","tokens_in":22789,"tokens_out":1720,"would_cite":false,"duration_ms":17644,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QLoRA fine-tuning sharply improves financial sentiment classification, but the tuned models' daily sentiment scores do not reliably predict stock returns after statistical correction.","keywords":["financial sentiment analysis","QLoRA","parameter-efficient fine-tuning","large language models","return predictability","rank information coefficient","cross-sectional stock returns","news sentiment"],"falsifier":"Re-run the downstream evaluation with timestamped headlines aligned to prices immediately before publication and returns measured over 5-minute to 4-hour horizons, using market- or sector-adjusted abnormal returns; if any model's intraday rank ICs are positive and survive false-discovery-rate correction, the paper's next-session null would be overturned.","tokens_in":21820,"feed_emoji":"📊","tokens_out":11712,"duration_ms":92238,"temperature":0.7,"pith_summary":"This paper separates two questions that are often conflated: whether QLoRA fine-tuning makes 7-8B language models accurate financial sentiment classifiers, and whether that accuracy becomes cross-sectional return predictability. On a unified three-class benchmark combining five financial text sources, QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615, and the best tuned backbone (Mistral-7B) reaches 0.8840 accuracy and 0.8771 macro-F1. On a temporally separate 2019 sample of 10,637 headlines linked to S&P 100 stocks, all seven probability-producing classifiers show positive but small one-day rank information coefficients, the largest being FinBERT's 0.0143, and none of the 28 model-horizon tests survives autocorrelation-consistent inference with false-discovery-rate correction. The paper concludes that parameter-efficient adaptation is effective for label accuracy while documenting a clear gap between semantic accuracy and tradable market signals.","feed_headline":"QLoRA boosts sentiment accuracy but not return forecasts","feed_subtitle":"Best tuned model hits 88% label accuracy; none of 28 rank-IC tests survives correction in a 2019 sample.","key_machinery":"The carrying mechanism is QLoRA (a frozen 4-bit quantized backbone trained only through low-rank adapters), converted into a sequence-classification head from which each headline yields a continuous expected-polarity score $s = p_{\\mathrm{pos}} - p_{\\mathrm{neg}} \\in [-1,1]$. Scores are averaged by stock and calendar date for stocks with fresh news, aligned to the first trading session strictly after the signal date with return $R^{(h)}_{i,t} = P^{\\mathrm{close}}_{i,t+h-1}/P^{\\mathrm{open}}_{i,t} - 1$, and then ranked cross-sectionally each day. Rank IC means are tested with heteroskedasticity- and autocorrelation-consistent standard errors and false-discovery-rate correction, and overlapping cohorts convert multi-day holdings into a daily portfolio return series. This pipeline is what turns label accuracy into an economic test without assuming probability calibration across model families.","core_discovery":"The central claim is that QLoRA is effective for financial sentiment adaptation, while classification accuracy and economic predictability are distinct objectives: a model that reproduces human sentiment labels more accurately need not rank future returns better. Experiment 1 shows large supervised gains—Qwen2.5's macro-F1 rises from 0.7274 zero-shot to 0.8615 after QLoRA, while Mistral-7B and LLaMA3-8B reach 0.8771 and 0.8753—and that inverse-frequency class weighting does not help. Experiment 2 finds that all seven models produce positive but small one-day rank ICs between 0.0013 and 0.0143, that every model's IC turns negative at the two- or five-day horizon, and that no model-horizon test remains significant after multiple-testing correction. The paper interprets the positive multi-day long-only returns as market exposure rather than sentiment alpha, and treats the results as evidence of a gap between linguistic performance and economic usefulness.","pith_inferences":["Beyond the paper: the timing mismatch it identifies is directly testable—align each timestamped headline to returns starting minutes after publication over 5-minute to 4-hour horizons; if intraday rank ICs are positive and survive false-discovery-rate correction, the next-session null is an artifact of the alignment rule.","Beyond the paper: because the two Qwen variants produce nearly identical signals (score correlation 0.998), downstream behavior is driven by the QLoRA-adapted backbone rather than loss weighting; a return-aware training objective that maximizes rank IC would test whether the classification-to-predictability gap can be closed by design.","Beyond the paper: the sample's uneven coverage—usable headlines for only 72 constituents and small daily cross-sections—limits statistical power, so a multi-year, broader-universe replication could overturn the null even if the 2019 result stands.","Beyond the paper: the paper's own framing implies that a model trained to predict abnormal returns conditioned on expectations, rather than human sentiment labels, might show predictability that these label-trained classifiers miss."],"forward_implications":["QLoRA can take a general 7-8B instruction-tuned LLM from roughly 0.73 to 0.86 macro-F1 on financial sentiment, so parameter-efficient adaptation is a practical route for specialized classifiers on limited hardware.","The best label-accuracy model is not the best economic-signal model: FinBERT, with the largest one-day rank IC, underperforms the QLoRA models on labels, so classifier leaderboards cannot substitute for return-based evaluation.","Any predictable component in these sentiment scores is concentrated in the first trading session and decays or reverses by days 2-5, consistent with fast price discovery.","The near-zero adjusted q-values (minimum 0.9622) imply that claims of financial sentiment alpha need multiple-testing-aware inference and cannot rest on one favorable horizon or one model.","Because long-only multi-day returns were broadly positive while long-short returns were negative, naive backtests that ignore market exposure can mistake beta for sentiment alpha."],"supporting_citations":[{"why":"Supplies the QLoRA method that carries the parameter-efficient adaptation result.","marker":"[2]"},{"why":"Supplies the LoRA low-rank adapter technique underlying all QLoRA fine-tunes.","marker":"[4]"},{"why":"Financial PhraseBank is one of the five benchmark datasets and the de facto sentence-level financial sentiment reference.","marker":"[11]"},{"why":"The FOMC corpus supplies the monetary-policy text type in the unified benchmark.","marker":"[13]"},{"why":"SEntFiN 1.0 supplies entity-aware financial news annotations in the benchmark.","marker":"[15]"},{"why":"FinBERT is the off-the-shelf domain encoder baseline that also produces the largest one-day rank IC in Experiment 2.","marker":"[1]"},{"why":"Financial-RoBERTa is the second off-the-shelf encoder baseline used in both experiments.","marker":"[14]"}],"fun_headline_variants":["QLoRA lifts sentiment accuracy, not return signals","High accuracy, zero tradable edge: QLoRA study","Accuracy soars, alpha falls flat in sentiment benchmark","88% F1 doesn't forecast returns: QLoRA reality check","Sentiment fine-tuning: better labels, no better trades"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The economic null rests on the rule that every calendar-date headline signal is tested against returns from the first trading session strictly after the signal date, so if price discovery happens within the same session, the rank-IC and portfolio tests can be near zero even when sentiment carries real information.","fun_headline_variants_meta":{"raw":{"variants":["QLoRA lifts sentiment accuracy, not return signals","High accuracy, zero tradable edge: QLoRA study","Accuracy soars, alpha falls flat in sentiment benchmark","88% F1 doesn't forecast returns: QLoRA reality check","Sentiment fine-tuning: better labels, no better trades"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1634,"prompt_tokens":1092,"completion_tokens":542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":459}},"tokens_in":708,"tokens_out":542,"duration_ms":5026,"temperature":1.0,"reasoning_tokens":459,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:41:45.044295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the downstream evaluation with timestamped headlines aligned to prices immediately before publication and returns measured over 5-minute to 4-hour horizons, using market- or sector-adjusted abnormal returns; if any model's intraday rank ICs are positive and survive false-discovery-rate correction, the paper's next-session null would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Financial PhraseBank is one of the five benchmark datasets and the de facto sentence-level financial sentiment reference."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The FOMC corpus supplies the monetary-policy text type in the unified benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SEntFiN 1.0 supplies entity-aware financial news annotations in the benchmark."}],"review_version":2}