REVIEW 3 major objections 4 minor 11 references
Decision-Oriented Text Evaluation
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The value of generated text should be measured by the decisions it enables, not by word overlap or fluency—and market digests show how.
desk verdict The abstract sells a collaborative-team result the experiments never ran, and the random baseline is undefined; the underlying idea is worth attention but this draft's central claims don't hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a decision-oriented evaluation protocol built around thresholded prediction accuracy. Each market digest is given to a participant—human annotator or LLM agent—who selects any subset of Taiwan-listed stocks to buy or sell without external references; a buy is scored correct if the stock closes above $+0.55\%$ and a sell if it falls below $-0.50\%$, and the participant's accuracy is the fraction of correct trades. Around this core, the paper builds two text-generation pipelines—one selecting assets by prior-day volatility, volume, and institutional flow ('performance-based') and one selecting assets named by professional journalists ('professional-insight')—so that the protocol can separate the effect of asset curation from the effect of text wording. Three large language models (GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet) act as both content generators and as autonomous investor-evaluators.
What would settle it
Run the same decision protocol with a random baseline: for each market digest, have each LLM and human place the same number of buy/sell trades uniformly at random, then compare thresholded accuracy. If random trades match or beat text-informed accuracy on LLM-generated morning briefs, the claim that those briefs add decision value would be falsified; if text-informed decisions beat random, the framework's core measure is validated. A second check would pair humans with LLMs in a joint trading condition to test the collaborative-outperformance claim directly.
Extended reading notes
Core claim
The central discovery is that the decision-making utility of a text is separable from its lexical or factual quality: under a thresholded-accuracy protocol in which a buy counts as correct only if the stock rises above $+0.55\%$ and a sell only if it falls below $-0.50\%$, the paper finds that LLM-generated morning briefs consistently raise decision accuracy relative to verbatim professional transcripts, while closing-bell reports show a division of labour—professional texts help LLM investors most, LLM-generated texts help human investors most. The paper also finds that human curation of the asset list ('professional-insight' selection) is a stabilising factor that improves outcomes for both kinds of investor. Taken together, the paper claims these results demonstrate that traditional intrinsic evaluation misses what makes financial text valuable: its capacity to produce profitable decisions, either alone or in human-LLM teams.
Load-bearing premise
The load-bearing premise is that short-horizon thresholded trade accuracy—buying stocks that rise above $+0.55\%$ and selling ones that fall below $-0.50\%$—is a valid and sufficient measure of a text's decision-making value, even though no chance-level baseline is measured.
Editorial extensions
If this is right
- If decision-oriented evaluation is sound, fluency and lexical overlap cannot certify a generation system fit for high-stakes use; downstream decision accuracy must be part of the acceptance test.
- LLM-generated objective summaries can outperform professional journalistic summaries for immediate trading decisions, so summary generation should be optimized for decision utility rather than reference fidelity.
- Human expertise remains load-bearing at the asset-selection stage, suggesting human-in-the-loop pipelines will keep an edge even where text generation is automated.
- Standardized LLM investors, run deterministically, offer a reproducible evaluation instrument that could be reused across future text-generation studies.
- Because the same text can help one audience and hurt another, decision-oriented evaluation should be audience-aware, distinguishing human readers from autonomous LLM readers.
Reading between the lines
- A natural extension would be to add a random-selection baseline to the protocol, since the paper's abstract states that humans and LLMs do not consistently beat random performance but no such baseline appears in the reported tables.
- The claimed collaborative advantage could be tested directly by pairing one human and one LLM in a joint trading session under the same texts and comparing their combined accuracy with the solo baselines reported in Table 1.
- The framework transfers to other high-stakes domains—medical summaries, legal memos—by replacing the financial return thresholds with a domain-specific outcome measure such as correct triage or correct legal action.
- If adopted as a standard, decision-oriented evaluation would push text generators to optimize for decision utility, possibly at the expense of surface faithfulness—an interesting shift for the field.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a decision-oriented framework for evaluating generated text, in which market digest texts (morning briefs and closing-bell reports) are scored by the accuracy of buy/sell decisions made by human investors and LLM agents who read only the text. The authors build a 30-day corpus from professional financial transcripts, generate alternative digests with GPT-4o under two asset-selection pipelines, and report thresholded prediction accuracy (+0.55% / -0.50%) and average transaction counts for three human and three LLM investors. The abstract's headline claims are that neither humans nor LLM agents consistently surpass random performance when relying solely on summaries, and that richer analytical commentaries enable collaborative human-LLM teams to outperform individual baselines significantly.
Significance. If the central claims were supported, decision-oriented evaluation would be a valuable complement to intrinsic metrics, and the human-LLM complementarity result would be of broad interest to the NLG and human-AI collaboration communities. The paper has useful ingredients: paid human annotators, explicit annotation guidelines, a thresholded accuracy metric, and a distinction between performance-based and professional-insight selection. However, the two load-bearing claims in the abstract are not supported by the experimental content: no human-LLM team condition appears anywhere in the protocol or results, and no random baseline is defined or measured. The reported differences also lack confidence intervals and significance tests. As presented, the paper does not establish its main conclusions.
major comments (3)
- [Abstract and §3.2] The abstract and §1 state that richer analytical commentaries enable collaborative human-LLM teams to outperform individual human or agent baselines significantly, but no such team condition exists in the experiments. Section 3.2 defines only individual investors (three human annotators and three LLM agents) who independently select stocks, and Tables 1 and 2 report only individual accuracies and transaction counts. Appendix A (Table 3) varies the digest generator, not the decision-making agent. The manuscript therefore contains no data from which a human-LLM collaborative-team result can be derived.
- [§3.2 and §4.1] The claim that 'neither humans nor LLM agents consistently surpass random performance' is untestable as written because no random baseline is defined. The thresholded accuracy labels each selected stock as correct or incorrect based on realized returns above +0.55% or below -0.50%; the expected accuracy of random selection depends on the base rates of qualifying upward and downward moves in the candidate universe and on the number and composition of stocks selected. Without these base rates or a random-selection control condition, the reported accuracy values cannot be compared with chance, and the negative result in the abstract is not evidenced.
- [Tables 1, 2, 3 and §4.1] All accuracy and transaction results are reported as raw percentages or averages without confidence intervals, statistical tests, or per-cell sample sizes. With only three human investors and three LLM agents per condition, differences such as Human C's closing-bell accuracy of 42.24% on journalist text versus 75.00% on performance-based text (Table 1) may reflect small-sample variability. The words 'consistently' and 'significantly' in the abstract and §4.1 are therefore not justified by the presented statistics.
minor comments (4)
- [Abstract] The term 'collaborative human-LLM teams' is never defined or operationalized in the body; either add an explicit team protocol or remove the claim.
- [§3.1 and Appendix A] The main text says the dataset covers a 30-day window, while Appendix A reports an extended 89-day dataset; clarify which corpus underlies Table 1 versus Table 3.
- [Appendix D] The annotation guidelines reference a 'provided companies.csv file' that is not included or linked; please provide the file or a full description of the stock universe.
- [Abstract and §5] The phrase 'frontier-scale LLM agents' is imprecise; 'frontier LLMs' or 'commercial LLMs' would be clearer.
Circularity Check
No significant circularity; the shared-author citation is motivational and the evaluation is externally benchmarked.
full rationale
This is an empirical evaluation paper, not a mathematical derivation, and no step reduces a claimed prediction to its own input by construction. The thresholded accuracy metric in Section 3.2 is externally sourced from Xu and Cohen (2018) and is applied to actual stock returns, so the evaluation has an external benchmark rather than being defined by the paper's own outputs. The one overlapping-author citation, Takayanagi et al. (2025) in Section 2, is used only to motivate outcome-based evaluation, noting that GPT-4-generated reports can sway investor choices; it does not justify a uniqueness theorem or supply any fitted parameter, so it is not load-bearing. The abstract's collaborative human-LLM team claim and the 'random performance' comparison do not appear in the experimental protocol; that is a gap between claims and evidence, not a circularity. No equation is reduced to itself, and no parameter is fitted and then renamed a prediction. Score 1 reflects the minor self-citation; the derivation chain itself is self-contained.
Assumptions & free parameters
free parameters (3)
- Rise/fall thresholds =
+0.55% rise, -0.50% fall
- Top K stocks =
not specified
- Prompt configurations =
not specified
assumptions (4)
- domain assumption Short-horizon prediction accuracy on self-selected stocks is a valid measure of text quality.
- domain assumption Three human annotators are representative of human investors.
- domain assumption The short time horizon isolates the effect of the text from other market factors.
- domain assumption The rise/fall thresholds from Xu and Cohen (2018) are appropriate for this market.
Cite this review
Pith. "Pith review of Decision-Oriented Text Evaluation." pith.science (2026). https://pith.science/paper/PCU674EP
@misc{pith2026250701923,
author = {Pith},
title = {Pith review of: Decision-Oriented Text Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/PCU674EP}},
note = {Machine review of arXiv:2507.01923}
}
read the original abstract
Natural language generation (NLG) is increasingly deployed in high-stakes domains, yet common intrinsic evaluation methods, such as n-gram overlap or sentence plausibility, weakly correlate with actual decision-making efficacy. We propose a decision-oriented framework for evaluating generated text by directly measuring its influence on human and large language model (LLM) decision outcomes. Using market digest texts--including objective morning summaries and subjective closing-bell analyses--as test cases, we assess decision quality based on the financial performance of trades executed by human investors and autonomous LLM agents informed exclusively by these texts. Our findings reveal that neither humans nor LLM agents consistently surpass random performance when relying solely on summaries. However, richer analytical commentaries enable collaborative human-LLM teams to outperform individual human or agent baselines significantly. Our approach underscores the importance of evaluating generated text by its ability to facilitate synergistic decision-making between humans and LLMs, highlighting critical limitations of traditional intrinsic metrics.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006. https://doi.org/10.18653/v1/E06-1032 Re-evaluating the role of bleu in machine translation research . In Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics, pages 249--256, Trento, Italy. Association for Computational Linguistics
-
[4]
Yen-Chun Hsu and Chenhao Tan. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.47 Decision-focused summarization . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 611--626, Online. Association for Computational Linguistics
-
[5]
Shrey Joshi, Jyoti Singh, Saket Singh, and Sushant Tyagi. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.349 Medicalsum: A guided clinical abstractive summarization model for generating medical reports from patient-doctor conversations . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4567--4578. Association for Comp...
-
[6]
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017. https://doi.org/10.18653/v1/D17-1238 Why we need new evaluation metrics for nlg . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2241--2252, Copenhagen, Denmark. Association for Computational Linguistics
-
[7]
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. https://doi.org/10.18653/v1/2021.naacl-main.383 Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, pages 4812--4829, Onl...
-
[8]
Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2024. https://aclanthology.org/2024.lrec-main.821 Is summary useful or not? an extrinsic human evaluation of text summaries on downstream tasks . In Proceedings of the 2024 Joint International Conference on Computational Linguistics and Language Resources and Evaluation (LREC-COLING 2024), pages 9389--9404, Torino, It...
work page 2024
Show all 11 references
-
[9]
Takehiro Takayanagi, Hiroya Takamura, Kiyoshi Izumi, and Chung-Chi Chen. 2025. https://doi.org/10.18653/v1/2025.findings-naacl.22 Can gpt-4 sway experts' investment decisions? In Findings of the Association for Computational Linguistics: NAACL 2025, pages 374--383, Albuquerque...
2025 doi
-
[10]
Yumo Xu and Shay B. Cohen. 2018. https://doi.org/10.18653/v1/P18-1183 Stock movement prediction from tweets and historical prices . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1970--1979, Melbourne, ...
2018 doi
-
[11]
Ningyu Zhang, Shumin Deng, Juan Li, Xi Chen, Wei Zhang, and Huajun Chen. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.2 Summarizing chinese medical answer with graph convolution networks and question-focused dual attention . In Findings of the Association for Computat...
2020 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.