{"id":"46c05382-fece-4978-a8c2-06a0fa420805","arxiv_id":"2608.11683","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A financial-agent benchmark with 220 open-ended queries and 11,543 source-attributed rubrics finds that the tool harness shapes performance more than the model alone, and that the authors' in-house system leads at 56%.","lead":"FrontierFinance is a new benchmark of 220 investment-research queries and 11,543 expert-written grading rubrics, released openly to test AI finance agents on open-ended analyst work. It claims to be broader and harder than existing finance benchmarks, and reports that the agent harness, not just the model, drives performance, with the authors' own system leading at 56%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM-judge scoring is unvalidated and never verifies rubric claims against their attributed sources; every headline result—qualification rates, difficulty, harness effects—depends on it.","rationale":"The reader's weakest assumption—that LLM-judge scoring is reliable without human validation and without factual verification—is exactly where the paper is most vulnerable. All three headline contributions (benchmark difficulty, harness effect, Samaya leadership) are read off the same rubric qualification rate, so a systematic judge bias or a judge that rewards surface matches would corrupt the central claim. I considered alternative concerns: the in-house harness is unreleased and confounded with Samaya's model, and the public 220 queries are a stratified subset of a larger private pool. These are genuine but secondary: they affect interpretation of one system's result and generalizability, whereas judge validity affects every number in Tables 4, 16, and 17 and the difficulty comparison in Figure 5. The paper does include real mitigating evidence—the benchmark and grading code are released, rubrics are source-attributed, macro-averaging reduces rubric-count skew, and Appendix C shows the difficulty judges are stable to model choice and position order. But those checks do not cover the scoring judges, and source attribution is unused at scoring time. The conditional verdict is therefore correct: accept the dataset as a contribution, but do not treat the model rankings, the harness-effect conclusion, or the 'harder than existing benchmarks' claim as established until the judge pipeline is validated against human experts on a representative sample. My proposed audit would settle the question directly; if LLM majority labels agree with expert labels at high kappa and the corrected rankings are unchanged, the concern is resolved.","tokens_in":25808,"tokens_out":8594,"duration_ms":91073,"concrete_test":"Run a human-validation audit on the released grading code: sample 300 rubric-judgment triples from the existing evaluations, stratified by use case, harness type, judge agreement (unanimous vs 2-1), and rubric category. Have two finance experts, working from the query date and the rubric's source attribution (e.g., the relevant 10-K, transcript, or market-data source), independently (i) decide whether the report actually contains the rubric's claim and (ii) verify that claim against the underlying public source. Compare expert majority labels with LLM majority labels using Cohen's kappa and per-label precision/recall. Then re-score a random 20-query subset of all 220 queries with expert labels and check whether the 56.0 vs 49.2 Samaya-vs-Claude Fable gap and the ordering among GPT 5.6 Sol / Kimi K3 / Gemini 3.6 Flash (46.8/46.4/46.3) survive the correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 defines answer quality as the majority verdict of three LLM judges (GPT 5.4, Gemini 3.1 Pro, Claude Sonnet 4.6), and Appendix F gives the judge prompt. Nowhere is the judge asked to consult the rubric's attributed public source; the prompt only says to decide whether the report 'adequately satisfies' each rubric. Rubrics are phrased as content checks ('States that X,' 'Provides information ... stating ...'), so the scoring pipeline rewards a report for containing the rubric's wording, even if the claim is fabricated or misattributed. The source-attribution field—advertised as enabling objective scoring—is not used by the judge. The paper reports no human validation or agreement statistics for this judge ensemble; the only calibration evidence is an unspecified claim of agreement with a nine-judge LLM committee. The robustness checks in Appendix C apply to the BT difficulty judges, not to the Section 5.1 qualification-rate judges. Because R_all and R_must-have mediate the benchmark's difficulty comparison, the harness-effect claim, and the Samaya-vs-frontier ranking, an uncalibrated judge would invalidate the paper's central measurements, not just its secondary analyses. This is a measurement-validity problem, not a dispute over whether checklist scoring is generally useful.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents FrontierFinance, a benchmark of 220 expert-authored financial research queries and 11,543 binary rubrics spanning six use cases across an investor workflow. Queries are timestamped, rubrics are tagged by essentiality, content category, and expected data source, and answers are scored by a macro-averaged rubric qualification rate computed by majority vote of three LLM judges. The authors evaluate frontier proprietary and open-weight models under three harnesses—a minimal web-search harness, an open-source Finance Agent v2 harness, and Samaya's in-house harness—and report that the harness strongly shapes quality and cost, with Samaya's system at 56.0%, Claude Fable 5 at 49.2%, and Kimi K3 at 46.4%. A trajectory analysis identifies common three-phase tool-use patterns and a link between parametric-knowledge URL recall and higher parse errors.","tokens_in":26042,"tokens_out":4845,"duration_ms":48478,"significance":"If the measurement pipeline is valid, FrontierFinance is a valuable resource: it is among the largest open finance-agent benchmarks to date, covers workflow breadth absent from prior suites, provides expert-authored rubrics with source labels, timestamps queries to reduce temporal leakage, and releases grading code for reproducibility. The trajectory analysis around tool-use efficiency and URL errors is a useful methodological contribution. The central difficulty and ranking conclusions are, however, contingent on the validity of LLM-judge scoring, which the paper does not establish with human validation or source verification.","major_comments":[{"comment":"The headline metric R in Eq. (3) is computed from majority verdicts of three LLM judges, but the paper reports no human agreement study for these judges; the only calibration statement is that the majority 'closely matched a larger committee of nine judges' in preliminary experiments, without statistics. The judge prompt in Appendix F asks whether the report 'adequately satisfies' each rubric and never consults the rubric's source attribution, so a report containing the rubric's wording can be scored as satisfied even if the underlying factual claim is not verified against the attributed public source. Because R_all and R_must-have mediate the difficulty comparison, the harness-effect claim, and the Samaya-versus-frontier ranking in Table 4, this is a measurement-validity issue that affects the paper's central conclusions. A human-annotated validation subset with per-rubric agreement, or a judge protocol that verifies claims against the attributed sources, is needed before the reported numbers can support the claims.","section":"5.1, Appendix F"},{"comment":"The difficulty scale used to claim FrontierFinance is harder than existing benchmarks (Figure 5) is itself derived from LLM pairwise judgments, with robustness checks (Tables 10-11, Figure 12) that only show consistency across LLM judges and not agreement with expert or human difficulty ratings. In addition, the public 220-query set was constructed by stratifying on this difficulty score, so the monotonically decreasing agent performance across easy/medium/hard buckets in Table 2 is at least partly a consequence of the selection procedure rather than an independent validation of the scale. Please provide expert judgment on a sample of pairs or another external anchor, and report an analysis of the public subset that accounts for the stratification.","section":"4.2, Appendix C"},{"comment":"Several headline comparisons are reported without uncertainty estimates. For example, Kimi K3 (46.4%) is only 0.4 pp behind GPT 5.6 Sol (46.8%), and the Screening & Discovery row of Table 16 is based on 17 queries; macro-averaging over so few queries can make differences of this magnitude indistinguishable from noise. The paper should report bootstrap confidence intervals, significance tests, or per-query score distributions before drawing conclusions about the ordering of systems and the claim that open-weight models 'nearly match' proprietary ones.","section":"6, Table 4"},{"comment":"The dataset is described as 'source-attributed,' but the source taxonomy includes 'professional knowledge' as a top-level category, accounting for 19% of rubrics overall and 32% for Sector, Industry & Macro. These rubrics are not tied to a specific public source, which weakens both the objectivity of rubric scoring and the claim of full source attribution; the paper should either reclassify such rubrics or explicitly discuss how they are verified.","section":"4.1, Table 9"}],"minor_comments":[{"comment":"The system prompt for the grading judge ends mid-sentence at 'Likewise against \"7.14%' with no continuation; the full prompt must be included for the claimed reproducibility.","section":"Appendix F"},{"comment":"No details are given for the 'preliminary experiments' with the nine-judge committee; report the agreement statistic and sample size.","section":"5.1"},{"comment":"The statement that the harness ordering 'holds across all six use cases and rubric categories' should be qualified: Table 16 shows small within-harness differences for Screening & Discovery, and the claim would benefit from an explicit statistical test rather than the radar chart alone.","section":"6, Figure 7"},{"comment":"The four-stage pipeline is described qualitatively; reporting inter-annotator agreement for rubric authoring and source labeling would strengthen the reliability claims.","section":"3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper evaluates the authors' own commercial system under an in-house harness, and the benchmark is released by the same company. The manuscript is transparent about this, but an independent audit or a data embargo would increase confidence in the benchmark's adoption. Also, reference [31] uses the same benchmark name 'FrontierFinance' for a different computer-use benchmark; the naming collision is worth flagging to avoid reader confusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Antoine, the useful part of this paper is real. FrontierFinance gives us an open set of 220 expert-crafted finance queries and 11,543 source-attributed rubrics spanning six workflow stages, with careful annotation, timestamped queries, and serious effort on difficulty calibration. The four-stage curation pipeline, the taxonomy, and the released data and grading code are all to the paper's credit. If you work on finance agents, this dataset is worth knowing about.\n\nThe soft spot is exactly where the stress-test note lands: the scoring pipeline in Section 5.1. Rubric qualification rate is the majority verdict of three LLM judges, and the judge prompt in Appendix F never asks the model to consult the rubric's attributed source. Rubrics are phrased as content checks—\"States that X\"—so a report that repeats the right wording gets the point even if the claim is fabricated or misattributed. That is not a quibble about checklist scoring generally; it means the source-attribution field, advertised as enabling objective scoring, is not actually used in grading. The paper reports no human agreement data for the judge ensemble, only an unspecified claim that it matched a nine-judge committee. There are also no confidence intervals around the main qualification rates, which matters when the headline claims are differences of a few points (56.0 vs 49.2, 46.4 vs 46.8).\n\nI would still send this to peer review, because the dataset and the grading code are reproducible and can be re-scored or supplemented by others. But a referee should treat the leaderboard and the harness-effect conclusion as provisional. The Samaya system topping the table while the in-house harness is not released does not help. The authors are transparent about temporal drift and subjectivity, but the missing judge validation is the gap that actually undermines the central measurements.\n\nBottom line: cite it, use the data, run your own evaluation. Do not quote the comparative rankings until the judge is validated against humans or the grading is changed to verify facts against sources.","headline":"A well-curated, genuinely useful finance-agent benchmark whose headline rankings I would not trust yet, because the LLM judges never verify rubrics against their own sources and no human validation is reported.","tokens_in":26621,"tokens_out":2698,"would_cite":true,"duration_ms":28101,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new finance benchmark spans the full investor workflow and shows that the tool harness, not the model alone, drives agent quality and cost.","keywords":["finance agents","benchmark","rubric-based evaluation","LLM-as-a-judge","investor workflow","open-ended research","agent harness"],"falsifier":"Take a random sample of rubric verdicts from the released grading code and have finance experts judge whether each report actually satisfies the rubric and whether the supporting fact is true in the attributed public source; if expert agreement with the LLM majority is substantially below the reported inter-judge consistency, the difficulty comparison and system rankings lose their evidential value.","tokens_in":25606,"feed_emoji":"📊","tokens_out":8759,"duration_ms":82432,"temperature":0.7,"pith_summary":"FrontierFinance is a publicly released benchmark of 220 expert-written investment-research queries, graded by 11,543 binary rubrics, each tied to a public data source. The paper's aim is to measure what existing finance benchmarks do not: whether an AI agent can plan, search, synthesize, and write a long-form research answer across the entire investor workflow rather than extract a single number. The central claim is that the benchmark is both broader and harder than current public finance benchmarks, and that when frontier models are tested under a shared harness, the surrounding tool harness changes quality and cost more than the underlying model does. A sympathetic reader would take away that open-ended financial research remains largely unsolved: the best system reaches 56% overall, and the hardest use cases, screening and discovery and sector/industry/macro, top out near 33% and 39%.","feed_headline":"Tool harness, not model, decides finance agent quality","feed_subtitle":"New 220-query benchmark rates long-form reports; best system hits 56%, hardest tasks stay below 40%.","key_machinery":"The central scoring object is the rubric qualification rate: each query is decomposed into expert-authored binary rubrics, and a system's per-query score is the fraction of those rubrics that a majority of three independent LLM judges finds satisfied in the agent's long-form report. Rubrics carry source attribution to a public data source tier, which turns open-ended answers into checkable criteria. The second load-bearing mechanism is the consensus Bradley–Terry difficulty scale, which places queries and whole benchmark suites on a shared 'how hard is this query' axis and supports the claim that FrontierFinance is harder than its predecessors.","core_discovery":"The paper's central discovery is twofold. First, the agent harness—the suite of tools, prompts, and orchestration surrounding the LLM—strongly shapes both answer quality and cost: the authors' in-house system leads at 56.0%, ahead of the strongest frontier model deployed under an open-source finance harness (49.2%) at roughly 2.2x lower cost, and the best open-weight model reaches 46.4%, nearly matching the best proprietary model while costing about 4.5x less. Second, the two most open-ended use cases—screening and discovery and sector/industry/macro—remain hardest across every system tested, with the best systems scoring only 33% and 39%. The paper further claims, using a consensus Bradley–Terry difficulty model fit over roughly 77,000 pairwise judgments, that FrontierFinance is substantially harder and wider in difficulty than three comparable public benchmarks, which cluster in the easy-to-medium range.","pith_inferences":["If the harness effect is real, raw model-API benchmarks for finance systematically underreport what a well-instrumented agent can do; the field should standardize the harness before comparing models.","The rubric-based design is vulnerable to format gaming: a system that learns to state rubric-like claims without verifying them against sources could inflate its score, so a human audit of judge verdicts against the attributed sources would be a natural validation step.","Because every query is date-anchored, scores will decay as future models memorize post-date data; periodic re-annotation with fresh query dates is the natural maintenance path, and the paper's reserved internal query pool is a ready reservoir for that.","The Bradley–Terry difficulty score could be reused to predict which queries benefit most from extra reasoning effort or tool budget, since the paper finds diminishing returns beyond each model's default effort."],"forward_implications":["Benchmark comparisons that mix model and harness cannot be read as model rankings; the same model will place very differently depending on the tools and orchestration it is given.","Open-weight models at roughly 46% qualification for under $1 per query put strong price-performance pressure on proprietary APIs, and the paper shows the gap closing within about two months of model release.","The two open-ended use cases, screening and discovery and sector/industry/macro, define the clearest headroom: even the best systems leave most of their rubrics unsatisfied.","The three-phase tool-use pattern—data gathering, mid-rollout synthesis, and answer preparation—appears across all systems, suggesting task structure drives agent behavior more than any single model policy.","Agents that navigate directly to remembered financial URLs rather than discovering sources through search incur higher access-error rates and token waste, pointing to a concrete efficiency failure mode."],"supporting_citations":[{"why":"Supplies the simple filing-grounded QA baseline whose queries are all classified as easy on the shared difficulty scale.","marker":"[1]"},{"why":"Supplies the open-source finance harness reused for evaluation and its human solve times for correlating difficulty with effort.","marker":"[8]"},{"why":"Supplies a comparable derivation-focused benchmark used in the difficulty-positioning analysis.","marker":"[9]"},{"why":"Supplies the checklist-based rubric scoring framework that FrontierFinance extends to the full investor workflow.","marker":"[26]"},{"why":"Supplies the Bradley–Terry paired-comparison model used to fit the query difficulty scale.","marker":"[38]"}],"fun_headline_variants":["Tool harness, not model, dictates finance agent quality","New benchmark: harness beats model pick for finance agents","FrontierFinance: hardest finance tasks still below 40%","Open-weight model nearly matches top AI at 4.5x lower cost","Agent harness, not LLM, drives finance answer quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scoring pipeline rests on the assumption that three LLM judges, voting by majority and checking only whether the report contains each rubric's statement, correctly decide which rubrics are satisfied, without any human-verified check of those facts against the attributed public source.","fun_headline_variants_meta":{"raw":{"variants":["Tool harness, not model, dictates finance agent quality","New benchmark: harness beats model pick for finance agents","FrontierFinance: hardest finance tasks still below 40%","Open-weight model nearly matches top AI at 4.5x lower cost","Agent harness, not LLM, drives finance answer quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1568,"prompt_tokens":981,"completion_tokens":587,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":597,"tokens_out":587,"duration_ms":6402,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:30:47.980652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of rubric verdicts from the released grading code and have finance experts judge whether each report actually satisfies the rubric and whether the supporting fact is true in the attributed public source; if expert agreement with the LLM majority is substantially below the reported inter-judge consistency, the difficulty comparison and system rankings lose their evidential value.","supporting_citations":[{"cited_title":"URL https://samaya.ai/blog/criteria-eval","cited_arxiv_id":null,"evidence_quote":"Supplies the checklist-based rubric scoring framework that FrontierFinance extends to the full investor workflow."}],"review_version":1}