Pith. sign in

REVIEW 4 major objections 4 minor 38 references

Frontier Financial Judgement shows no evaluated AI agent matches expert analyst news labels on more than 52.4% of cases.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 09:43 UTC pith:OT75OBSO

load-bearing objection A well-built, honest benchmark for agent news-flow filtering, but the 52.4% headline floats on unmeasured expert-label reliability. the 4 major comments →

arxiv 2607.20645 v1 pith:OT75OBSO submitted 2026-07-22 cs.CL cs.AIcs.LG

Frontier Financial Judgement: Can agents tell what might move a stock?

classification cs.CL cs.AIcs.LG
keywords financial news evaluationequity analyst judgementAI agentsnew information detectionmaterialityfalse positivesbenchmarkvaluation impact
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces Frontier Financial Judgement, a benchmark for testing whether AI research agents can reproduce professional equity analysts' core judgement about company news: whether an item is genuinely new, how much it should matter for valuation, and in which direction. The benchmark uses 82 realistic cases built from expert-designed synthetic events blended with real live articles and historical documents, with a fixed information cutoff to prevent leakage. Evaluating 14 agents, the paper finds that the strongest agent matches all three expert labels in only 52.4% of cases, and that estimated false-positive escalation on surrounding noise ranges from roughly 1% to 32% across agents. The authors argue the failures are not simply retrieval or sentiment errors but reflect missing contextual financial judgement, such as recognizing subtle changes inside repeated disclosures or choosing the financially relevant comparison. They conclude that practical deployment of news-flow filtering requires joint evaluation of accuracy, restraint, output reliability, and cost, not accuracy alone.

Core claim

The central claim is that current frontier agents cannot yet reliably replicate expert equity-analyst judgement on financial news flow. The strongest agent, GPT-5.5, achieves 52.4% all-label accuracy—matching the expert's newness, expected importance, and direction labels for the same event—while its atomic accuracy is 71.1%. On genuinely new events, all-label accuracy drops to 50.0% for the best agent and much lower for others. The paper also demonstrates that target-event accuracy does not determine false-positive restraint: agents with nearly identical target performance differ sharply in how often they escalate irrelevant live articles, from GPT-5.6 Sol's 1.0% to Claude Opus 4.8's 24.9%.

What carries the argument

The load-bearing device is a three-label judgement schema—information_new, expected importance, and direction—with accepted secondary labels for genuine boundary cases, and an all-label accuracy measure that requires all three classifications to match the expert labels. Synthetic events are designed by professional analysts to be realistic but nonexistent, rendered as web-like articles, and placed in bundles containing five recently collected real articles and two historical documents, all assessed at a fixed evidence cutoff with web-search access. This construction makes novelty detection, materiality calibration, and directional reasoning jointly measurable while reducing the risk that ans

Load-bearing premise

Expert-assigned labels are treated as ground truth, but the paper reports no inter-rater reliability or adjudication process, so if independent analysts would disagree on a substantial share of these 82 cases, the headline accuracy ceiling loses its anchor.

What would settle it

Take a random sample of the 82 cases and have several independent professional analysts label each one with the same newness-importance-direction schema, then measure pairwise agreement. If all-three-label agreement among experts is near 52% rather than near 100%, the benchmark ceiling would reflect noisy ground truth instead of agent failure; if expert agreement is high, the 52.4% figure would be confirmed as a genuine capability gap.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim holds, automated equity-news filtering cannot yet replace professional analyst judgement in consequential financial workflows, because even the best agent misses the complete expert assessment on almost half of events.
  • Because false-positive rates vary from about 1% to 32% among agents with similar target accuracy, practical deployment must evaluate restraint separately from accuracy, especially in low-base-rate news environments.
  • The systematic weakness in expected-importance labelling—every agent scores lower on importance than on newness—identifies materiality calibration as a key bottleneck for improving financial judgement agents.
  • The benchmark design of mixing synthetic expert-designed events with live articles and historical documents offers a contamination-resistant template for repeated evaluation of point-in-time financial judgement.
  • Cost, latency, and output reliability vary independently of accuracy, so performance comparisons that ignore these operational dimensions are insufficient for real-world deployment decisions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If expert labels on these 82 cases are measurably noisy—something the paper does not test—the 52.4% ceiling could understate agent capability rather than define it; an inter-rater reliability study would resolve this.
  • The shared failure on the sequential-revision case suggests a testable intervention: prompting agents to compute period-over-period deltas before assigning direction may improve accuracy on framing-masked reversals.
  • The benchmark's current concentration on semiconductor supply-chain companies means sector-specific vocabulary and analyst conventions could inflate or deflate measured difficulty; extending to other sectors would clarify how general the capability gap is.
  • The observed cost-accuracy frontier implies that deployment choices may be driven less by model intelligence than by engineering around latency and escalation budgets, a direction the paper leaves implicit.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Frontier Financial Judgement, a benchmark of 82 expert-designed synthetic news events embedded in realistic article/distractor bundles, and uses it to evaluate 14 LLM-based agents. Each case requires three labels—newness, expected importance, and direction—and the headline result is that the best agent, GPT-5.5, matches the full expert label set in only 52.4% of cases. The paper also reports an approximate false-positive proxy, along with cost, token, and search-volume metrics, and argues that accuracy, cost, and restraint form a multi-dimensional trade-off for real-world deployment.

Significance. If the expert labels are shown to be reliable, this is a valuable contribution to financial-agent evaluation. The benchmark design has genuine strengths: the synthetic-event approach controls contamination and realism; the evidence cutoff is frozen; the answer schema is validated; and the joint measurement of accuracy, cost, output reliability, and false-positive behaviour addresses aspects that many NLP benchmarks ignore. The paper also transparently acknowledges that its false-positive rates are an approximate proxy and that the sector scope is limited. However, the central claim—that no agent can reliably reproduce expert judgement—rests entirely on the stability and reproducibility of the expert-assigned labels, and the manuscript currently provides no evidence on that point. The absence of a human-baseline measurement and of a validation that the rendered articles unambiguously convey the intended event makes the headline accuracy difficult to interpret. These are fixable with additional experiments and reporting.

major comments (4)
  1. [§3.1, Table 4, §6] The headline 52.4% all-label accuracy is measured against expert labels, but no inter-rater reliability or adjudication is reported. Section 6 concedes that 'some expert judgements are inherently subjective.' If independent experts agree only imperfectly on importance or direction, the agent accuracy numbers may be partially an artifact of label noise rather than a measure of agent capability. Please add a second-expert annotation study on a random subset (with per-label agreement such as Cohen's kappa or percent agreement), including adjudication, and report a human-expert ceiling measured on the final rendered articles. Without this, the central quantitative claim lacks a critical anchor.
  2. [§3.1, §3.2] Labels are assigned to the intended event description before the LLM renders the article and before web-page chrome is applied. There is no validation step showing that the final rendered article, as presented to agents, unambiguously conveys the event and supports the same labels. If rendering introduces ambiguity or accidentally obscures a key fact, agents are being scored against an unobservable target. Please have experts annotate the rendered articles (or a sample) to confirm that the gold labels remain recoverable from the presented text, and report the agreement between labels assigned to the event specification and labels assigned to the rendered article.
  3. [§3.1, Table 1 note] The scoring allows multiple accepted importance or direction labels, but the frequency and distribution of such secondary labels are never reported. For example, Table 1 notes that a high importance is also accepted for the 'Vera folded into Rubin opportunity' event, but the reader cannot tell how many of the 82 cases have multiple accepted labels or what fraction of the 'correct' predictions use a secondary label. Without this information, the effective difficulty of the target set is unquantified and the reported accuracies may be inflated. Please report, per label and overall, the number of cases with multiple accepted labels and the proportion of agent-correct responses that rely on a secondary label.
  4. [§3.4, Table 5, Abstract] The abstract and discussion describe a spread in 'false-positive rates' from ~1% to ~32%, but these are approximate proxy rates computed on unlabeled distractors: an item is counted as a false positive when a model labels it both new and more important than none. This proxy is not calibrated against expert judgments of whether those distractor articles are actually new and important, so the absolute numbers and the operational trade-off claim rest on an unvalidated assumption. The authors acknowledge the proxy in §3.4 and the Table 5 notes, but the abstract presents the numbers without that context in the first bullet. Please either (a) validate the proxy on a random sample of distractors via expert annotation, or (b) explicitly reframe these as 'escalation rates' throughout the paper and de-emphasise the raw comparison.
minor comments (4)
  1. [§4.1, Table 3] Nemotron 3 Super produced only 28 valid answers out of 82 cases, so its accuracy scores are not comparable with other agents. The paper acknowledges this, but consider reporting Nemotron's results separately or excluding it from the headline ranking to avoid misleading visual comparisons.
  2. [Abstract and §3.3] The abstract says '656 items for assessment,' which is 82 cases × 8 items; making this explicit would help the reader. Also, '82 synthetic events' vs. '656 items' is slightly confusing because each case contains eight items, not one.
  3. [General] There is no data availability statement. The conclusion calls the benchmark 'a reproducible foundation,' but the paper does not state whether the benchmark items, expert labels, or agent harness configurations will be released. Please add an explicit availability statement and, if applicable, a link to a public repository.
  4. [Figure 1] Figure 1 is informative but a label or legend identifying a few named points would help readability; the log-scale cost axis makes it difficult to distinguish points in the crowded lower-left region.

Circularity Check

0 steps flagged

No circularity: headline accuracy and false-positive figures are direct measurements against externally assigned expert labels, with no fitted parameters, no definitional reduction, and no load-bearing self-citation.

full rationale

This paper is a benchmark evaluation, not a derivation. The central claim — that the strongest agent matches all expert labels in only 52.4% of cases (Section 4, Table 3) — is a direct measurement of agent outputs against ground-truth labels assigned by professional analysts in Section 3.1 ('Experts assign three labels... and a label rationale'), i.e., labels defined independently of and prior to the evaluated agents' outputs. No parameter is fitted to any subset of data and then 'predicted' on a closely related quantity; no equation reduces to its own input; and no uniqueness theorem or ansatz is imported from the authors' prior work. The non-trivial dependency is the ground-truth construct itself: Section 3.1 allows multiple accepted labels for boundary cases ('experts can specify more than one accepted importance or direction label') and Section 6 concedes 'some expert judgements are inherently subjective, and other financial professionals may reasonably have assigned different labels.' That is a validity/reliability risk — the target may be noisy or imperfectly anchored to the rendered articles — but it is not circularity, because the target remains external to the systems being scored. The false-positive metric is explicitly presented as an 'approximate escalation proxy' (Section 3.4) computed on unlabelled distractors, i.e., a descriptive statistic of agent behavior rather than a fitted prediction. The Harbor harness citation [13] is an open-source, containerized framework, and the accuracy comparison against expert labels does not depend on it for the headline numbers. Evaluation via the paper's own limitations statement shows no step in which the result is equivalent to its inputs by construction, so the honest finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The benchmark's load-bearing inputs are expert judgement, the curation of distractors, and the realism of LLM-generated articles. No free parameters are fitted, but these assumptions are unmeasured and directly affect the validity of both the accuracy and false-positive numbers.

axioms (4)
  • domain assumption Expert labels are treated as ground truth.
    Section 3.1: 'experts assign three labels... gold labels'; no inter-rater agreement or adjudication is reported, so the accuracy numbers inherit whatever noise is in the expert judgements.
  • domain assumption LLM-generated synthetic articles faithfully represent realistic news flow and do not leak the answer.
    Section 3.1 combines expert event descriptions with primer paragraphs and LLM rendering; the benchmark's validity depends on the result being realistic and not artificially easy or hard, but no human realism check is reported.
  • domain assumption All live articles and historical documents in each case are distractors (not new and not important) for the false-positive proxy.
    Section 3.4 explicitly states 'live articles and historical documents are curated distractors rather than individually expert-labelled negative examples,' so the reported false-positive rates are controlled estimates, not gold-labelled rates.
  • domain assumption A primary or accepted secondary expert label is the correct answer.
    Scoring (Section 3.4) counts a prediction correct when it matches the primary or an accepted secondary label; if other analysts would choose different labels, reported accuracies would change.

pith-pipeline@v1.3.0-alltime-deepseek · 15371 in / 8185 out tokens · 71050 ms · 2026-08-01T09:43:52.339279+00:00 · methodology

0 comments
read the original abstract

We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements. Rapidly identifying new information, evaluating its implications and determining its valuation impact is one of the most time-consuming and challenging aspects of real-world equity coverage. This is becoming ever more difficult and important as AI rapidly increases the quantity of new information to process. The strongest agent we evaluate on Frontier Financial Judgement matches all expert labels in only 52.4% of cases. We also find significant divergence in estimated false-positive rates among frontier agents, ranging from ~1% for GPT-5.6 Sol to ~32% for Claude Sonnet 4.6. To construct the benchmark and make it representative of real-world settings, we combine human-designed and labelled synthetic articles with live news articles and historical documents, creating 656 items for assessment. The resulting task requires agents to distinguish genuinely new, valuation-relevant financial information from stale, immaterial or misleading news under realistic conditions. We find substantial trade-offs among agent accuracy, cost, false positives and reliability that continue to hinder the reliable deployment of news-flow filtering in practice.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 19 canonical work pages · 4 internal anchors

  1. [1]

    Mikhail, and Andrea S

    Paul Asquith, Michael B. Mikhail, and Andrea S. Au. Information content of equity analyst reports.Journal of Financial Economics, 75(2):245–282, February 2005. doi: 10.1016/j.jfineco.2004.01.002. URL https://doi.org/10.1016/j.jfineco.2004.01.002

  2. [2]

    Finance agent benchmark: Benchmarking LLMs on real-world financial research tasks.arXiv preprint arXiv:2508.00828, 2025

    Antoine Bigeard, Langston Nashold, Rayan Krishnan, and Shirley Wu. Finance agent benchmark: Benchmarking LLMs on real-world financial research tasks.arXiv preprint arXiv:2508.00828, 2025. doi: 10.48550/arXiv.2508.00828. URL https://arxiv.org/abs/2508.00828

  3. [3]

    Disclosure processing costs, investors’ information choice, and equity market outcomes: A review.Journal of Accounting and Economics, 70(2–3):101344, November

    Elizabeth Blankespoor, Ed deHaan, and Iván Marinovic. Disclosure processing costs, investors’ information choice, and equity market outcomes: A review.Journal of Accounting and Economics, 70(2–3):101344, November

  4. [4]

    Enhancing ESG news annotation: Leveraging GPT for the analysis of ESG news and events.SSRN Electronic Journal, February 2025

    Keven Bluteau, Frank Coggins, and Gilles Boevi Koumou. Enhancing ESG news annotation: Leveraging GPT for the analysis of ESG news and events.SSRN Electronic Journal, February 2025. doi: 10.2139/ssrn.5128896. URL https://ssrn.com/abstract=5128896

  5. [5]

    Information, trading, and volatility: Evidence from firm-specific news.The Review of Financial Studies, 32(3):992–1033, March 2019

    Jacob Boudoukh, Ronen Feldman, Shimon Kogan, and Matthew Richardson. Information, trading, and volatility: Evidence from firm-specific news.The Review of Financial Studies, 32(3):992–1033, March 2019. doi: 10.1093/rfs/hhy083. URL https://doi.org/10.1093/rfs/hhy083

  6. [6]

    Brown, Andrew C

    Lawrence D. Brown, Andrew C. Call, Michael B. Clement, and Nathan Y. Sharp. The activities of buy-side analysts and the determinants of their stock recommendations.Journal of Accounting and Economics, 62(1): 139–156, August 2016. doi: 10.1016/j.jacceco.2016.06.002. URL https://doi.org/10.1016/j.jacceco.2016.06.002

  7. [7]

    EFSA: Towards event-level financial sentiment analysis

    Tianyu Chen, Yiming Zhang, Guoxin Yu, Dapeng Zhang, Li Zeng, Qing He, and Xiang Ao. EFSA: Towards event-level financial sentiment analysis. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7455–7467, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18...

  8. [8]

    FinAgentBench: A benchmark dataset for agentic retrieval in financial question answering.arXiv preprint arXiv:2508.14052, 2025

    Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim, Minjae Kim, Juneha Hwang, Jaeseon Ha, Hojun Choi, Suyeol Yun, Yongjin Kim, and Yongjae Lee. FinAgentBench: A benchmark dataset for agentic retrieval in financial question answering.arXiv preprint arXiv:2508.14052, 2025. doi: 10.48550/arXiv.2508.14052. URL https://arxiv.org/abs/2508.14052

  9. [9]

    Demirakos, Norman C

    Efthimios G. Demirakos, Norman C. Strong, and Martin Walker. What valuation models do analysts use? Accounting Horizons, 18(4):221–240, December 2004. doi: 10.2308/acch.2004.18.4.221. URL https://doi.org/10.2308/acch.2004.18.4.221

  10. [10]

    Exa Search API, 2026

    Exa. Exa Search API, 2026. URL https://exa.ai/docs/reference/search. API documentation, accessed 19 July 2026

  11. [11]

    When can the market identify old news?Journal of Financial Economics, 149(1):92–113, July 2023

    Anastassia Fedyk and James Hodson. When can the market identify old news?Journal of Financial Economics, 149(1):92–113, July 2023. doi: 10.1016/j.jfineco.2023.04.008. URL https://doi.org/10.1016/j.jfineco.2023.04.008

  12. [12]

    Assessing look-ahead bias in stock return predictions generated by GPT sentiment analysis.The Journal of Financial Data Science, 6(1):25–42, 2024

    Paul Glasserman and Caden Lin. Assessing look-ahead bias in stock return predictions generated by GPT sentiment analysis.The Journal of Financial Data Science, 6(1):25–42, 2024. doi: 10.3905/jfds.2023.1.143. URL https://doi.org/10.3905/jfds.2023.1.143

  13. [13]

    Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026

    Harbor Framework Team. Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026. URL https://github.com/harbor-framework/harbor. Software, version 0.15.0

  14. [14]

    Driven to distraction: Extraneous events and underreaction to earnings news.The Journal of Finance, 64(5):2289–2325, October 2009

    David Hirshleifer, Sonya Seongyeon Lim, and Siew Hong Teoh. Driven to distraction: Extraneous events and underreaction to earnings news.The Journal of Finance, 64(5):2289–2325, October 2009. doi: 10.1111/j.1540-6261.2009.01501.x. URL https://doi.org/10.1111/j.1540-6261.2009.01501.x

  15. [15]

    14 FinSearchComp: Towards a realistic, expert-level evaluation of financial search and reasoning.arXiv preprint arXiv:2509.13160, 2025

    Liang Hu, Jianpeng Jiao, Jiashuo Liu, Yanle Ren, Zhoufutu Wen, Kaiyuan Zhang, Xuanliang Zhang, Xiang Gao, Tianci He, Fei Hu, Yali Liao, Zaiyuan Wang, Chenghao Yang, Qianyu Yang, Mingren Yin, Zhiyuan Zeng, Ge Zhang, Xinyi Zhang, Xiying Zhao, Zhenwei Zhu, Hongseok Namkoong, Wenhao Huang, and Yuwen Tang. 14 FinSearchComp: Towards a realistic, expert-level ev...

  16. [16]

    FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023

    Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023. doi: 10.48550/arXiv.2311.11944. URL https://arxiv.org/abs/2311.11944

  17. [17]

    Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

    Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5075–5084, Singapore, Decem...

  18. [18]

    McCurdy, and Xiaofei Zhao

    Yoontae Jeon, Thomas H. McCurdy, and Xiaofei Zhao. News as sources of jumps in stock returns: Evidence from 21 million news articles for 9000 companies.Journal of Financial Economics, 145(2):1–17, August 2022. doi: 10.1016/j.jfineco.2021.08.002. URL https://doi.org/10.1016/j.jfineco.2021.08.002

  19. [19]

    All that glisters is not gold: A benchmark for reference-free counterfactual financial misinformation detection

    Yuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He, Ziyang Xu, Chen Xu, Zhiyang Deng, Prayag Tiwari, Xi Chen, Alejandro Lopez-Lira, Jimin Huang, Junichi Tsujii, and Sophia Ananiadou. All that glisters is not gold: A benchmark for reference-free counterfactual financial misinformation detection. InProceedings of the 64th Annual Meeting of the Association for ...

  20. [20]

    S. P. Kothari, Eric So, and Rodrigo Verdi. Analysts’ forecasts and asset pricing: A survey.Annual Review of Financial Economics, 8:197–219, October 2016. doi: 10.1146/annurev-financial-121415-032930. URL https://doi.org/10.1146/annurev-financial-121415-032930

  21. [21]

    FinMarBa: A Market-Informed Dataset for Financial Sentiment Classification

    Baptiste Lefort, Eric Benhamou, Beatrice Guez, Jean-Jacques Ohana, Ethan Setrouk, and Alban Etienne. FinMarBa: A market-informed dataset for financial sentiment classification.arXiv preprint arXiv:2507.22932, 2025. doi: 10.48550/arXiv.2507.22932. URL https://arxiv.org/abs/2507.22932

  22. [22]

    ExAnte: A benchmark for ex-ante inference in large language models

    Yachuan Liu, Xiaochun Wei, Lin Shi, Xinnuo Li, Bohan Zhang, Paramveer Dhillon, and Qiaozhu Mei. ExAnte: A benchmark for ex-ante inference in large language models. InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1551–1571, Rabat, Morocco, March 2026. Association f...

  23. [23]

    FMDLlama: Financial misinformation detection based on large language models

    Zhiwei Liu, Xin Zhang, Kailai Yang, Qianqian Xie, Jimin Huang, and Sophia Ananiadou. FMDLlama: Financial misinformation detection based on large language models. InCompanion Proceedings of the ACM on Web Conference 2025, pages 1153–1157. Association for Computing Machinery, 2025. doi: 10.1145/3701716.3715599. URL https://doi.org/10.1145/3701716.3715599

  24. [24]

    AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements

    Zhiwei Liu, Yueru He, Qing Ou, Tianlei Zhu, Xiaorui Guo, Xueqing Peng, and Sophia Ananiadou. AuditFraudBench: Benchmarking audit judgment in detecting fraudulent misstatements.arXiv preprint arXiv:2606.08345, 2026. doi: 10.48550/arXiv.2606.08345. URL https://arxiv.org/abs/2606.08345

  25. [25]

    Mikhail, Beverly R

    Michael B. Mikhail, Beverly R. Walther, and Richard H. Willis. The effect of experience on security analyst underreaction.Journal of Accounting and Economics, 35(1):101–116, April 2003. doi: 10.1016/S0165-4101(02)00099-X. URL https://doi.org/10.1016/S0165-4101(02)00099-X

  26. [26]

    DiligenceBench: An equity-research agent evaluation, July 2026

    Malthe Have Musaeus, Faisal Sayed, Mersad Abbasi, Daanish Khazi, and Karina Nguyen. DiligenceBench: An equity-research agent evaluation, July 2026. URL https://www.paperinstruments.com/blog/diligence-bench. Paper Instruments and Thoughtful Lab

  27. [27]

    Patell and Mark A

    James M. Patell and Mark A. Wolfson. The intraday speed of adjustment of stock prices to earnings and dividend announcements.Journal of Financial Economics, 13(2):223–252, June 1984. doi: 10.1016/0304-405X(84)90024-2. URL https://doi.org/10.1016/0304-405X(84)90024-2. 15

  28. [28]

    FinResearchBench: A logic tree based agent-as-a-judge evaluation framework for financial research agents.arXiv preprint arXiv:2507.16248, 2025

    Rui Sun, Zuo Bai, Wentao Zhang, Yuxiang Zhang, Li Zhao, Shan Sun, and Zhengwen Qiu. FinResearchBench: A logic tree based agent-as-a-judge evaluation framework for financial research agents.arXiv preprint arXiv:2507.16248, 2025. doi: 10.48550/arXiv.2507.16248. URL https://arxiv.org/abs/2507.16248

  29. [29]

    Paul C. Tetlock. All the news that’s fit to reprint: Do investors react to stale information?The Review of Financial Studies, 24(5):1481–1512, May 2011. doi: 10.1093/rfs/hhq141. URL https://doi.org/10.1093/rfs/hhq141

  30. [30]

    Tetlock, Maytal Saar-Tsechansky, and Sofus Macskassy

    Paul C. Tetlock, Maytal Saar-Tsechansky, and Sofus Macskassy. More than words: Quantifying language to measure firms’ fundamentals.The Journal of Finance, 63(3):1437–1467, June 2008. doi: 10.1111/j.1540-6261.2008.01362.x. URL https://doi.org/10.1111/j.1540-6261.2008.01362.x

  31. [31]

    BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

    Alex Wang, Georg Meinhardt, Jacob Katz, Joseph H. Kim, Pratyush K. Chaudhary, Chase Blagden, and Eric Xu. BigFinanceBench: A workflow-grounded benchmark for financial-research agents.arXiv preprint arXiv:2606.03829, 2026. doi: 10.48550/arXiv.2606.03829. URL https://arxiv.org/abs/2606.03829

  32. [32]

    Livebench: A challenging, contamination-limited LLM benchmark

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum. Livebench: A challenging, contamination-limited LLM benchmark. In...

  33. [33]

    FinDVer: Explainable claim verification over long and hybrid-content financial documents

    Yilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, and Arman Cohan. FinDVer: Explainable claim verification over long and hybrid-content financial documents. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 14739–14752, Miami, Florida, USA, No...

  34. [34]

    Zijie Zhao and Roy E. Welsch. Point-in-time financial RAG with frozen LLMs and market-feedback adaptive retrieval.arXiv preprint arXiv:2605.31201, 2026. doi: 10.48550/arXiv.2605.31201. URL https://arxiv.org/abs/2605.31201

  35. [35]

    Trade the event: Corporate events detection for news-based event-driven trading

    Zhihan Zhou, Liqian Ma, and Han Liu. Trade the event: Corporate events detection for news-based event-driven trading. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2114–2124, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.186. URL https://aclanthology.org/2021.findin...

  36. [36]

    Towards temporal-aware multi-modal retrieval augmented generation in finance

    Fengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, and Tat-Seng Chua. Towards temporal-aware multi-modal retrieval augmented generation in finance. InProceedings of the 33rd ACM International Conference on Multimedia, pages 6289–6297. Association for Computing Machinery, 2025. doi: 10.1145/3746027.3755723. URL https://...

  37. [2020]

    URL https://doi.org/10.1016/j.jacceco.2020.101344

    doi: 10.1016/j.jacceco.2020.101344. URL https://doi.org/10.1016/j.jacceco.2020.101344

  38. [2025]

    ICLR 2025 Spotlight

    URL https://proceedings.iclr.cc/paper_files/paper/2025/hash/ e4a46394ba5378b3f9a186a5b4c650d1-Abstract-Conference.html. ICLR 2025 Spotlight