{"id":"37ac1b4c-214a-4eaf-a129-632bc50fa636","arxiv_id":"2607.20645","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"No evaluated AI agent can fully match professional analysts' newness/importance/direction labels on the new 82-case Frontier Financial Judgement benchmark; GPT-5.5 tops out at 52.4%.","lead":"This paper introduces a benchmark where AI agents read news and documents about semiconductor companies and must flag what is genuinely new and could move a stock; even the best tested agent gives the full correct answer only about half the time. A generalist should care because it measures whether AI can do a real analyst's filtering job, not just answer quiz questions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Expert-label reliability is the load-bearing assumption: the 52.4% headline compares agents to a ground truth whose inter-rater agreement and anchoring to the final rendered articles are unmeasured.","rationale":"The reader identified the absence of inter-rater reliability as the weakest assumption. This is indeed the most load-bearing concern for the central claim, because the headline accuracy is a comparison against expert labels. Without a measure of how reproducible those labels are across experts, and without evidence that the labels describe the article actually shown to agents, the 52.4% number cannot be separated from label noise. The paper's own Section 6 admits subjectivity, but does not quantify it. The multiple-accepted-label mechanism is an attempt to handle ambiguity, yet its usage is unreported, so the difficulty of the benchmark is not calibrated. A dual-annotation study on the final rendered articles would directly test the stability of the target and also catch any event-article mismatch. This concern does not change the reader's verdict: the paper still presents a useful benchmark, but should be CONDITIONAL pending the reliability evidence and release. No other concern (e.g., false-positive proxy) is as fundamental to the headline claim. Agreement with the reader: agree.","tokens_in":15622,"tokens_out":6265,"duration_ms":59399,"concrete_test":"Select a random subset of 30 of the 82 final articles (with chrome and navigation stripped), and have 2–3 independent professional equity analysts annotate each with the same three-label scheme, without access to the original event descriptions or gold labels. Compute per-label and all-label agreement (e.g., Fleiss' kappa) between the independent analysts and the original gold labels. Also compute the maximum achievable all-label accuracy if an agent always chose the modal expert label (or any label within the set accepted by a majority of experts). If agreement is low (e.g., kappa < 0.6) or if the gold labels are not consistently recoverable from the rendered articles, the 52.4% figure should be reinterpreted as close to the label-noise ceiling and the headline claim substantially weakened. Report the secondary-label usage rate as well.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that no agent exceeds 52.4% all-label accuracy is only interpretable if the expert-assigned labels are a stable, reproducible target. Section 3.1 states experts 'assign three labels... and a label rationale,' but no inter-rater reliability or adjudication is reported. Section 6 concedes 'some expert judgements are inherently subjective.' Moreover, labels appear to be attached to the intended event description, not necessarily to the final LLM-rendered article; the paper does not report a validation step where experts read the rendered article and confirm that it unambiguously conveys the event. If another expert would assign different labels, or if the rendered article contains ambiguities not in the event spec, the true 'expert ceiling' may be close to or below 52.4%, making the headline an artifact of label noise rather than agent failure. The multiple-accepted-labels mechanism (Section 3.1) further loosens the target set, but the frequency and distribution of secondary labels are never reported, so the effective difficulty is unquantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Frontier Financial Judgement, a benchmark of 82 expert-designed synthetic news events embedded in realistic article/distractor bundles, and uses it to evaluate 14 LLM-based agents. Each case requires three labels—newness, expected importance, and direction—and the headline result is that the best agent, GPT-5.5, matches the full expert label set in only 52.4% of cases. The paper also reports an approximate false-positive proxy, along with cost, token, and search-volume metrics, and argues that accuracy, cost, and restraint form a multi-dimensional trade-off for real-world deployment.","tokens_in":15846,"tokens_out":4031,"duration_ms":41904,"significance":"If the expert labels are shown to be reliable, this is a valuable contribution to financial-agent evaluation. The benchmark design has genuine strengths: the synthetic-event approach controls contamination and realism; the evidence cutoff is frozen; the answer schema is validated; and the joint measurement of accuracy, cost, output reliability, and false-positive behaviour addresses aspects that many NLP benchmarks ignore. The paper also transparently acknowledges that its false-positive rates are an approximate proxy and that the sector scope is limited. However, the central claim—that no agent can reliably reproduce expert judgement—rests entirely on the stability and reproducibility of the expert-assigned labels, and the manuscript currently provides no evidence on that point. The absence of a human-baseline measurement and of a validation that the rendered articles unambiguously convey the intended event makes the headline accuracy difficult to interpret. These are fixable with additional experiments and reporting.","major_comments":[{"comment":"The headline 52.4% all-label accuracy is measured against expert labels, but no inter-rater reliability or adjudication is reported. Section 6 concedes that 'some expert judgements are inherently subjective.' If independent experts agree only imperfectly on importance or direction, the agent accuracy numbers may be partially an artifact of label noise rather than a measure of agent capability. Please add a second-expert annotation study on a random subset (with per-label agreement such as Cohen's kappa or percent agreement), including adjudication, and report a human-expert ceiling measured on the final rendered articles. Without this, the central quantitative claim lacks a critical anchor.","section":"§3.1, Table 4, §6"},{"comment":"Labels are assigned to the intended event description before the LLM renders the article and before web-page chrome is applied. There is no validation step showing that the final rendered article, as presented to agents, unambiguously conveys the event and supports the same labels. If rendering introduces ambiguity or accidentally obscures a key fact, agents are being scored against an unobservable target. Please have experts annotate the rendered articles (or a sample) to confirm that the gold labels remain recoverable from the presented text, and report the agreement between labels assigned to the event specification and labels assigned to the rendered article.","section":"§3.1, §3.2"},{"comment":"The scoring allows multiple accepted importance or direction labels, but the frequency and distribution of such secondary labels are never reported. For example, Table 1 notes that a high importance is also accepted for the 'Vera folded into Rubin opportunity' event, but the reader cannot tell how many of the 82 cases have multiple accepted labels or what fraction of the 'correct' predictions use a secondary label. Without this information, the effective difficulty of the target set is unquantified and the reported accuracies may be inflated. Please report, per label and overall, the number of cases with multiple accepted labels and the proportion of agent-correct responses that rely on a secondary label.","section":"§3.1, Table 1 note"},{"comment":"The abstract and discussion describe a spread in 'false-positive rates' from ~1% to ~32%, but these are approximate proxy rates computed on unlabeled distractors: an item is counted as a false positive when a model labels it both new and more important than none. This proxy is not calibrated against expert judgments of whether those distractor articles are actually new and important, so the absolute numbers and the operational trade-off claim rest on an unvalidated assumption. The authors acknowledge the proxy in §3.4 and the Table 5 notes, but the abstract presents the numbers without that context in the first bullet. Please either (a) validate the proxy on a random sample of distractors via expert annotation, or (b) explicitly reframe these as 'escalation rates' throughout the paper and de-emphasise the raw comparison.","section":"§3.4, Table 5, Abstract"}],"minor_comments":[{"comment":"Nemotron 3 Super produced only 28 valid answers out of 82 cases, so its accuracy scores are not comparable with other agents. The paper acknowledges this, but consider reporting Nemotron's results separately or excluding it from the headline ranking to avoid misleading visual comparisons.","section":"§4.1, Table 3"},{"comment":"The abstract says '656 items for assessment,' which is 82 cases × 8 items; making this explicit would help the reader. Also, '82 synthetic events' vs. '656 items' is slightly confusing because each case contains eight items, not one.","section":"Abstract and §3.3"},{"comment":"There is no data availability statement. The conclusion calls the benchmark 'a reproducible foundation,' but the paper does not state whether the benchmark items, expert labels, or agent harness configurations will be released. Please add an explicit availability statement and, if applicable, a link to a public repository.","section":"General"},{"comment":"Figure 1 is informative but a label or legend identifying a few named points would help readability; the log-scale cost axis makes it difficult to distinguish points in the crowded lower-left region.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written and the benchmark design is thoughtful, but for a benchmark paper the lack of any statement about public release of the items and labels is a significant omission. External replication and community adoption will be difficult without a release plan. Additionally, the expert-label reliability issue is the main substantive concern; I would ask for a concrete inter-rater reliability and human-ceiling experiment before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is the benchmark design: expert-designed synthetic events embedded in realistic bundles with a fixed evidence cutoff, jointly scored on newness, importance, and direction, plus a cost/restraint axis. That's a real gap in the AI-finance evaluation literature, and this paper fills it competently. The methodology is described clearly, the prompt and construction details are in the appendix, and the authors are upfront that the false-positive rates are approximate and that sector coverage is narrow. Credit where due: the shared failure examples in Figures 4 and 5 are genuinely instructive, and the cost/accuracy tables are useful for anyone thinking about deployment.\n\nThe soft spots are mostly around the expert labels. There is no inter-rater reliability, no adjudication, and no check that the final LLM-rendered article unambiguously conveys the intended event. The paper itself concedes that \"some expert judgements are inherently subjective.\" That matters because the 52.4% all-label accuracy is only interpretable against a stable target. If two experts would disagree on, say, a third of the cases, the ceiling might be well below 100%, and the headline becomes less striking. The multiple-accepted-labels mechanism helps, but the frequency of secondary labels is never reported. I'd want to see a label-reliability study before taking the absolute accuracy numbers at face value. That said, the relative ranking of agents is probably robust to label noise if the noise is not systematically agent-correlated, so the cost/accuracy tradeoff findings should survive.\n\nMinor quibbles: the benchmark is not released, which limits reproducibility, and the false-positive proxy counts articles flagged as \"new + more important than none\" without verifying they are actually false positives. The authors label it as approximate, so that's more of a scope note than a flaw. Also, 82 cases in one sector is small, but they acknowledge that.\n\nOverall, this is a solid, honest contribution that deserves serious peer review. The authors know their limitations and state them. The main fixes—inter-rater data, released benchmark—are addressable in revision. I'd bring it to a reading group and would consider citing it for the benchmark design, though I'd wait for the reliability data before quoting the 52.4% number.","headline":"A well-built, honest benchmark for agent news-flow filtering, but the 52.4% headline floats on unmeasured expert-label reliability.","tokens_in":16311,"tokens_out":2174,"would_cite":true,"duration_ms":20434,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frontier Financial Judgement shows no evaluated AI agent matches expert analyst news labels on more than 52.4% of cases.","keywords":["financial news evaluation","equity analyst judgement","AI agents","new information detection","materiality","false positives","benchmark","valuation impact"],"falsifier":"Take a random sample of the 82 cases and have several independent professional analysts label each one with the same newness-importance-direction schema, then measure pairwise agreement. If all-three-label agreement among experts is near 52% rather than near 100%, the benchmark ceiling would reflect noisy ground truth instead of agent failure; if expert agreement is high, the 52.4% figure would be confirmed as a genuine capability gap.","tokens_in":15501,"feed_emoji":"📈","tokens_out":3457,"duration_ms":31112,"temperature":0.7,"pith_summary":"This paper introduces Frontier Financial Judgement, a benchmark for testing whether AI research agents can reproduce professional equity analysts' core judgement about company news: whether an item is genuinely new, how much it should matter for valuation, and in which direction. The benchmark uses 82 realistic cases built from expert-designed synthetic events blended with real live articles and historical documents, with a fixed information cutoff to prevent leakage. Evaluating 14 agents, the paper finds that the strongest agent matches all three expert labels in only 52.4% of cases, and that estimated false-positive escalation on surrounding noise ranges from roughly 1% to 32% across agents. The authors argue the failures are not simply retrieval or sentiment errors but reflect missing contextual financial judgement, such as recognizing subtle changes inside repeated disclosures or choosing the financially relevant comparison. They conclude that practical deployment of news-flow filtering requires joint evaluation of accuracy, restraint, output reliability, and cost, not accuracy alone.","feed_headline":"Best AI agent matches analysts on only 52% of news calls","feed_subtitle":"New 82-case benchmark finds frontier agents miss subtle valuation shifts and escalate noise at very different rates.","key_machinery":"The load-bearing device is a three-label judgement schema—information_new, expected importance, and direction—with accepted secondary labels for genuine boundary cases, and an all-label accuracy measure that requires all three classifications to match the expert labels. Synthetic events are designed by professional analysts to be realistic but nonexistent, rendered as web-like articles, and placed in bundles containing five recently collected real articles and two historical documents, all assessed at a fixed evidence cutoff with web-search access. This construction makes novelty detection, materiality calibration, and directional reasoning jointly measurable while reducing the risk that ans","core_discovery":"The central claim is that current frontier agents cannot yet reliably replicate expert equity-analyst judgement on financial news flow. The strongest agent, GPT-5.5, achieves 52.4% all-label accuracy—matching the expert's newness, expected importance, and direction labels for the same event—while its atomic accuracy is 71.1%. On genuinely new events, all-label accuracy drops to 50.0% for the best agent and much lower for others. The paper also demonstrates that target-event accuracy does not determine false-positive restraint: agents with nearly identical target performance differ sharply in how often they escalate irrelevant live articles, from GPT-5.6 Sol's 1.0% to Claude Opus 4.8's 24.9%.","pith_inferences":["If expert labels on these 82 cases are measurably noisy—something the paper does not test—the 52.4% ceiling could understate agent capability rather than define it; an inter-rater reliability study would resolve this.","The shared failure on the sequential-revision case suggests a testable intervention: prompting agents to compute period-over-period deltas before assigning direction may improve accuracy on framing-masked reversals.","The benchmark's current concentration on semiconductor supply-chain companies means sector-specific vocabulary and analyst conventions could inflate or deflate measured difficulty; extending to other sectors would clarify how general the capability gap is.","The observed cost-accuracy frontier implies that deployment choices may be driven less by model intelligence than by engineering around latency and escalation budgets, a direction the paper leaves implicit."],"forward_implications":["If the central claim holds, automated equity-news filtering cannot yet replace professional analyst judgement in consequential financial workflows, because even the best agent misses the complete expert assessment on almost half of events.","Because false-positive rates vary from about 1% to 32% among agents with similar target accuracy, practical deployment must evaluate restraint separately from accuracy, especially in low-base-rate news environments.","The systematic weakness in expected-importance labelling—every agent scores lower on importance than on newness—identifies materiality calibration as a key bottleneck for improving financial judgement agents.","The benchmark design of mixing synthetic expert-designed events with live articles and historical documents offers a contamination-resistant template for repeated evaluation of point-in-time financial judgement.","Cost, latency, and output reliability vary independently of accuracy, so performance comparisons that ignore these operational dimensions are insufficient for real-world deployment decisions."],"fun_headline_variants":["Frontier AI matches analysts on only 52% of news judgments","Top AI agents flunk financial news test: 52% accuracy","AI false alarms vary 1% to 25% on financial news","Best AI can't tell valuable news from noise: 52% match"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Expert-assigned labels are treated as ground truth, but the paper reports no inter-rater reliability or adjudication process, so if independent analysts would disagree on a substantial share of these 82 cases, the headline accuracy ceiling loses its anchor.","fun_headline_variants_meta":{"raw":{"variants":["Frontier AI matches analysts on only 52% of news judgments","Top AI agents flunk financial news test: 52% accuracy","AI false alarms vary 1% to 25% on financial news","Best AI can't tell valuable news from noise: 52% match"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2402,"prompt_tokens":746,"completion_tokens":1656,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":1588}},"tokens_in":490,"tokens_out":1656,"duration_ms":11241,"temperature":1.0,"reasoning_tokens":1588,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:43:52.339279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 82 cases and have several independent professional analysts label each one with the same newness-importance-direction schema, then measure pairwise agreement. If all-three-label agreement among experts is near 52% rather than near 100%, the benchmark ceiling would reflect noisy ground truth instead of agent failure; if expert agreement is high, the 52.4% figure would be confirmed as a genuine capability gap.","supporting_citations":[],"review_version":1}