{"id":"7c5fcf41-dbda-45b5-a5bf-50283000773e","arxiv_id":"2507.01923","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A decision-oriented evaluation framework for generated text is tested on market digests, but the reported evidence lacks a random baseline and the abstract claims team results that do not appear in the experiments.","lead":"This paper proposes judging generated text by how well it helps people and AI agents make financial trading decisions, rather than by word overlap with reference texts. In a small study of market digests, the evidence is mixed and the abstract claims stronger findings than the experiment tables support.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central claim about collaborative human-LLM teams outperforming baselines is not implemented anywhere in the experimental protocol or results, and the 'random performance' companion claim has no defined baseline.","rationale":"The framework is plausible and the paper deserves credit for a decision-oriented protocol, transparent annotator details in Appendices C and D, and comparisons across three LLM investors. However, the central empirical assertion in the abstract is not internally supported: no human-LLM team condition exists in the experiment, and the random-performance comparison is undefined. These are internal-evidence problems, not disagreements with prevailing consensus. The reader's REJECT is appropriate because the strongest claims as stated are disconnected from the reported results. The paper could be revised by actually running a human-LLM joint condition, adding a well-defined random baseline, and reporting significance tests; without those, the current version does not support its headline conclusion.","tokens_in":7247,"tokens_out":4787,"duration_ms":58206,"concrete_test":"Run a claim-to-evidence audit: enumerate every claim in the abstract and Section 4.1, then map each claim to a specific experimental condition in Section 3.2 and a row in Tables 1–3. If no condition pairs a human with an LLM in a single joint decision (for example, a human making the final choice after reading an LLM recommendation, or an LLM and human vote being combined), the collaborative-team claim is unsupported. In the same audit, compute a per-day random baseline by sampling the same number of stocks with the same buy/sell split from the full Taiwan-listed universe and scoring with the 0.55%/0.50% thresholds; a paired or binomial test comparing each Table 1 cell against this baseline would determine whether the random-performance claim actually holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states that richer analytical commentaries enable collaborative human-LLM teams to outperform individual human or agent baselines significantly, but the body never operationalizes a human-LLM team. Section 3.2 defines only individual investors: three human annotators and three LLM agents each select stocks independently from a full universe. Tables 1 and 2 report individual accuracies and transaction counts; Appendix A (Table 3) varies only the digest generator. No condition combines a human decision with an LLM decision, so the collaborative-team result cannot be derived from the reported data. The companion claim that summaries do not beat random performance is also untestable as written because no random baseline is specified; the +0.55%/−0.50% thresholds in Section 3.2 measure directional accuracy, not chance level, unless the base rates of qualifying up and down moves in the selected universe are given. Without these base rates and per-decision significance tests, neither the negative summary claim nor the positive commentary claim is supported. The load-bearing empirical assertion of the paper, as stated, is therefore not evidenced by the presented results.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a decision-oriented framework for evaluating generated text, in which market digest texts (morning briefs and closing-bell reports) are scored by the accuracy of buy/sell decisions made by human investors and LLM agents who read only the text. The authors build a 30-day corpus from professional financial transcripts, generate alternative digests with GPT-4o under two asset-selection pipelines, and report thresholded prediction accuracy (+0.55% / -0.50%) and average transaction counts for three human and three LLM investors. The abstract's headline claims are that neither humans nor LLM agents consistently surpass random performance when relying solely on summaries, and that richer analytical commentaries enable collaborative human-LLM teams to outperform individual baselines significantly.","tokens_in":7437,"tokens_out":3019,"duration_ms":38058,"significance":"If the central claims were supported, decision-oriented evaluation would be a valuable complement to intrinsic metrics, and the human-LLM complementarity result would be of broad interest to the NLG and human-AI collaboration communities. The paper has useful ingredients: paid human annotators, explicit annotation guidelines, a thresholded accuracy metric, and a distinction between performance-based and professional-insight selection. However, the two load-bearing claims in the abstract are not supported by the experimental content: no human-LLM team condition appears anywhere in the protocol or results, and no random baseline is defined or measured. The reported differences also lack confidence intervals and significance tests. As presented, the paper does not establish its main conclusions.","major_comments":[{"comment":"The abstract and §1 state that richer analytical commentaries enable collaborative human-LLM teams to outperform individual human or agent baselines significantly, but no such team condition exists in the experiments. Section 3.2 defines only individual investors (three human annotators and three LLM agents) who independently select stocks, and Tables 1 and 2 report only individual accuracies and transaction counts. Appendix A (Table 3) varies the digest generator, not the decision-making agent. The manuscript therefore contains no data from which a human-LLM collaborative-team result can be derived.","section":"Abstract and §3.2"},{"comment":"The claim that 'neither humans nor LLM agents consistently surpass random performance' is untestable as written because no random baseline is defined. The thresholded accuracy labels each selected stock as correct or incorrect based on realized returns above +0.55% or below -0.50%; the expected accuracy of random selection depends on the base rates of qualifying upward and downward moves in the candidate universe and on the number and composition of stocks selected. Without these base rates or a random-selection control condition, the reported accuracy values cannot be compared with chance, and the negative result in the abstract is not evidenced.","section":"§3.2 and §4.1"},{"comment":"All accuracy and transaction results are reported as raw percentages or averages without confidence intervals, statistical tests, or per-cell sample sizes. With only three human investors and three LLM agents per condition, differences such as Human C's closing-bell accuracy of 42.24% on journalist text versus 75.00% on performance-based text (Table 1) may reflect small-sample variability. The words 'consistently' and 'significantly' in the abstract and §4.1 are therefore not justified by the presented statistics.","section":"Tables 1, 2, 3 and §4.1"}],"minor_comments":[{"comment":"The term 'collaborative human-LLM teams' is never defined or operationalized in the body; either add an explicit team protocol or remove the claim.","section":"Abstract"},{"comment":"The main text says the dataset covers a 30-day window, while Appendix A reports an extended 89-day dataset; clarify which corpus underlies Table 1 versus Table 3.","section":"§3.1 and Appendix A"},{"comment":"The annotation guidelines reference a 'provided companies.csv file' that is not included or linked; please provide the file or a full description of the stock universe.","section":"Appendix D"},{"comment":"The phrase 'frontier-scale LLM agents' is imprecise; 'frontier LLMs' or 'commercial LLMs' would be clearer.","section":"Abstract and §5"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract's collaborative-team and random-baseline claims and the experimental design is substantial: the paper would require new experiments and new statistical analysis to support its headline conclusions. The self-citation to Takayanagi et al. (2025) is relevant and not problematic, but the lack of any human-LLM team condition is a fundamental gap rather than a presentation issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract promises a collaborative human–LLM team result, but the body never runs that condition. Section 3.2 defines only individual investors — three humans and three LLMs acting alone. No table combines a human decision with an LLM decision, so the abstract's headline claim is unsupported by the presented protocol. The companion claim that summaries are no better than random is also untestable: the +0.55%/−0.50% thresholds measure directional accuracy, not chance, and the base rates of qualifying up/down moves in the selected universe are never given. No random baseline is defined anywhere.\n\nThat is the big soft spot, and it is load-bearing. But the paper does useful things too. The idea of scoring generated text by downstream decision outcomes is timely, and extending it to market digests with both human and LLM investors is a legitimate proof-of-concept. The dataset construction — two selection pipelines, professional transcripts aligned to market events, detailed annotation guidelines — is careful. The raw accuracy numbers in Table 1 are new, and the distinction between morning briefs and closing-bell reports is a sensible axis of variation.\n\nBeyond the missing team condition and random baseline, the statistics are thin. With three annotators and a 30-day window, the raw percentages need at least confidence intervals or simple tests. No data or code are released, which weakens a contribution that is partly empirical. The self-citation to Takayanagi et al. is fine; the exposure to Hsu and Tan and Pu et al. is honest.\n\nThis paper deserves a serious referee, but not because the current draft is sound. The question it asks is important, and the experimental machinery is mostly in place — the inference just overreaches. A revision that defines a proper random baseline, adds a real human–LLM team condition, reports significance tests, and releases the data could turn this into a solid short paper. My verdict on this version is reject-and-resubmit, not desk-reject.","headline":"The abstract sells a collaborative-team result the experiments never ran, and the random baseline is undefined; the underlying idea is worth attention but this draft's central claims don't hold.","tokens_in":7971,"tokens_out":2968,"would_cite":false,"duration_ms":38415,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The value of generated text should be measured by the decisions it enables, not by word overlap or fluency—and market digests show how.","keywords":["decision-oriented text evaluation","market digest evaluation","thresholded prediction accuracy","LLM investors","human-LLM collaboration","summarization evaluation","financial decision-making","NLG evaluation"],"falsifier":"Run the same decision protocol with a random baseline: for each market digest, have each LLM and human place the same number of buy/sell trades uniformly at random, then compare thresholded accuracy. If random trades match or beat text-informed accuracy on LLM-generated morning briefs, the claim that those briefs add decision value would be falsified; if text-informed decisions beat random, the framework's core measure is validated. A second check would pair humans with LLMs in a joint trading condition to test the collaborative-outperformance claim directly.","tokens_in":7039,"feed_emoji":"📈","tokens_out":6953,"duration_ms":78063,"temperature":0.7,"pith_summary":"The paper argues that generated text in high-stakes settings should be evaluated by the decisions it enables, not by surface similarity to reference text. To make this concrete, it treats daily market digests—objective morning briefs and analytical closing-bell reports—as a test bed, and judges each text by the accuracy of buy/sell decisions made by three human investors and three LLM agents who read only that text. On this measure, the paper reports that journalistic summaries are not the most decision-useful input: LLM-generated morning briefs improve decision accuracy, professional-insight asset selection helps both humans and agents, and richer analytical commentaries support the best overall decisions, including the claimed benefit of human-LLM collaboration. The point matters because if decision-oriented evaluation is right, the field's standard intrinsic metrics are measuring the wrong thing.","feed_headline":"Score generated text by the trades it inspires, not its prose","feed_subtitle":"A decision-oriented framework rates market digests on whether human and AI investors act on them profitably.","key_machinery":"The machinery is a decision-oriented evaluation protocol built around thresholded prediction accuracy. Each market digest is given to a participant—human annotator or LLM agent—who selects any subset of Taiwan-listed stocks to buy or sell without external references; a buy is scored correct if the stock closes above $+0.55\\%$ and a sell if it falls below $-0.50\\%$, and the participant's accuracy is the fraction of correct trades. Around this core, the paper builds two text-generation pipelines—one selecting assets by prior-day volatility, volume, and institutional flow ('performance-based') and one selecting assets named by professional journalists ('professional-insight')—so that the protocol can separate the effect of asset curation from the effect of text wording. Three large language models (GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet) act as both content generators and as autonomous investor-evaluators.","core_discovery":"The central discovery is that the decision-making utility of a text is separable from its lexical or factual quality: under a thresholded-accuracy protocol in which a buy counts as correct only if the stock rises above $+0.55\\%$ and a sell only if it falls below $-0.50\\%$, the paper finds that LLM-generated morning briefs consistently raise decision accuracy relative to verbatim professional transcripts, while closing-bell reports show a division of labour—professional texts help LLM investors most, LLM-generated texts help human investors most. The paper also finds that human curation of the asset list ('professional-insight' selection) is a stabilising factor that improves outcomes for both kinds of investor. Taken together, the paper claims these results demonstrate that traditional intrinsic evaluation misses what makes financial text valuable: its capacity to produce profitable decisions, either alone or in human-LLM teams.","pith_inferences":["A natural extension would be to add a random-selection baseline to the protocol, since the paper's abstract states that humans and LLMs do not consistently beat random performance but no such baseline appears in the reported tables.","The claimed collaborative advantage could be tested directly by pairing one human and one LLM in a joint trading session under the same texts and comparing their combined accuracy with the solo baselines reported in Table 1.","The framework transfers to other high-stakes domains—medical summaries, legal memos—by replacing the financial return thresholds with a domain-specific outcome measure such as correct triage or correct legal action.","If adopted as a standard, decision-oriented evaluation would push text generators to optimize for decision utility, possibly at the expense of surface faithfulness—an interesting shift for the field."],"forward_implications":["If decision-oriented evaluation is sound, fluency and lexical overlap cannot certify a generation system fit for high-stakes use; downstream decision accuracy must be part of the acceptance test.","LLM-generated objective summaries can outperform professional journalistic summaries for immediate trading decisions, so summary generation should be optimized for decision utility rather than reference fidelity.","Human expertise remains load-bearing at the asset-selection stage, suggesting human-in-the-loop pipelines will keep an edge even where text generation is automated.","Standardized LLM investors, run deterministically, offer a reproducible evaluation instrument that could be reused across future text-generation studies.","Because the same text can help one audience and hurt another, decision-oriented evaluation should be audience-aware, distinguishing human readers from autonomous LLM readers."],"supporting_citations":[{"why":"Supplies the evidence that BLEU correlates poorly with human judgment, motivating the need for alternative evaluation.","marker":"Callison-Burch et al., 2006"},{"why":"Documents weaknesses in intrinsic NLG metrics across domains, the baseline this paper argues against.","marker":"Novikova et al., 2017"},{"why":"Shows factuality errors in summarization that surface metrics miss, supporting the gap between intrinsic quality and decision value.","marker":"Pagnoni et al., 2021"},{"why":"Introduces decision-focused evaluation in healthcare, the direct predecessor this paper extends to financial commentary.","marker":"Hsu and Tan, 2021"},{"why":"Supplies the thresholded rise/fall labeling ($+0.55\\%$/$-0.50\\%$) used to score each trade.","marker":"Xu and Cohen, 2018"},{"why":"Shows that GPT-4-generated financial reports can sway investors, the behavioral effect this framework formalizes.","marker":"Takayanagi et al., 2025"},{"why":"Provides extrinsic human evaluation of summaries on downstream tasks, which this work extends to LLM agents and real trades.","marker":"Pu et al., 2024"}],"fun_headline_variants":["Judge text by the trades it inspires, not its prose","Score text by its impact on decisions, not its style","Forget n-gram overlap: measure text by decisions it drives","Text value lies in the trades it enables, not its grammar","Evaluate generated text by the decisions it improves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that short-horizon thresholded trade accuracy—buying stocks that rise above $+0.55\\%$ and selling ones that fall below $-0.50\\%$—is a valid and sufficient measure of a text's decision-making value, even though no chance-level baseline is measured.","fun_headline_variants_meta":{"raw":{"variants":["Judge text by the trades it inspires, not its prose","Score text by its impact on decisions, not its style","Forget n-gram overlap: measure text by decisions it drives","Text value lies in the trades it enables, not its grammar","Evaluate generated text by the decisions it improves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1338,"prompt_tokens":870,"completion_tokens":468,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":486,"tokens_out":468,"duration_ms":5312,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:39:34.772677+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same decision protocol with a random baseline: for each market digest, have each LLM and human place the same number of buy/sell trades uniformly at random, then compare thresholded accuracy. If random trades match or beat text-informed accuracy on LLM-generated morning briefs, the claim that those briefs add decision value would be falsified; if text-informed decisions beat random, the framework's core measure is validated. A second check would pair humans with LLMs in a joint trading condition to test the collaborative-outperformance claim directly.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the evidence that BLEU correlates poorly with human judgment, motivating the need for alternative evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces decision-focused evaluation in healthcare, the direct predecessor this paper extends to financial commentary."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that GPT-4-generated financial reports can sway investors, the behavioral effect this framework formalizes."}],"review_version":1}