{"id":"c8f60f33-a897-4878-a48e-60dde44bf2c5","arxiv_id":"2507.17186","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FinGAIA is a 407-task Chinese financial agent benchmark where the best agent, ChatGPT DeepResearch, scores 48.9%, far below financial experts at 84.7%.","lead":"This paper introduces FinGAIA, a Chinese benchmark with 407 tasks that test AI agents on realistic financial work: reading PDFs and images, browsing websites, and running Python. The best agent scored 48.9%, while financial experts scored 84.7%, showing current agents lag experts by more than 35 points.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FinGAIA's gold-answer key is neither frozen nor internally consistent: dynamic web tasks lack snapshots, and Figure 16's gold answer (13,510) contradicts its own error analysis (15,440), so the 48.9%-vs-84.7% result is not yet reproducible.","rationale":"FinGAIA is a serious resource: 407 expert-validated tasks, a three-tier scenario structure, 10 evaluated agents, human baselines, and a plausible qualitative result. I am not disputing the novelty claim or the existence of a real capability gap. The issue is that the paper's central quantitative claims depend on a gold-answer key that is neither archived nor fully checked. The reader already identified temporal instability as the weakest assumption; the additional internal contradictions (Figure 16's 13,510 vs 15,440; Table 2 RMA 57.1 vs text RMA 60.0; 'six' vs five error types) make the problem concrete rather than hypothetical. These are not stylistic issues—they affect the correctness of individual answers and therefore the aggregate scores. The remedy is straightforward and within the authors' control: release frozen prompts, attachments, archived snapshots, and a complete answer key, and audit the key for internal consistency. Until then, the 48.9% and roughly 36-point expert gap should be treated as preliminary. This supports the reader's CONDITIONAL verdict; I would not escalate to REJECT because the issues are addressable and do not clearly invalidate the benchmark's design.","tokens_in":23685,"tokens_out":4518,"duration_ms":44198,"concrete_test":"Run a full reproducibility audit before public release: (1) Archive every external resource referenced by each of the 407 tasks (webpages, PDFs, audio, market data) at the evaluation date and re-derive each gold answer from the archived snapshot, comparing against the released answer key; report the fraction of gold answers that change or cannot be recovered. (2) As a minimal internal check, recompute the Figure 16 margin-balance reduction from the cited '2024 Margin Trading Parameter List' and determine whether 13,510 or 15,440 is correct; if 15,440 is correct, at least one task's gold answer is wrong. (3) Recompute ChatGPT's weighted average from the official per-scenario scores; if it does not equal 48.9, the reported headline is not self-consistent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that FinGAIA measures agents and that ChatGPT's 48.9% lags experts by roughly 36 points—requires the ground-truth answers to be correct, stable, and recoverable. Many tasks (Figures 5, 8–10) require retrieving live data from official websites and market platforms: fee tables, annual reports, fund holdings, and intraday quotes. The Limitations section concedes that 'certain FinGAIA tasks may inherently involve dynamic elements,' yet the paper provides no snapshots, versioned pages, or frozen answer files. If any referenced page changes, the gold answer changes and all scores inherit that instability. The fragility is concrete, not hypothetical: Figure 16's question asks for the margin-balance reduction on a 19,300 yuan margin purchase; the 'Correct answer' line gives 13,510 yuan, while the same figure's error analysis states the correct value is 15,440 yuan (=19,300 × 0.80). Only one can be the gold answer. Similar inconsistencies appear elsewhere: Table 2 lists ChatGPT RMA = 57.1, while the main text cites 'RMA 60.0' (the value in Table 5 after exclusions), and the error analysis says 'six fundamental limitations' while enumerating five. Because the headline number is an aggregate over this key, even a small rate of unstable or wrong gold answers can shift ChatGPT's 48.9% and the expert gap by several points. This is the load-bearing weakness; the benchmark's novelty and construction effort are not at issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FinGAIA, a 407-task benchmark for evaluating AI agents in Chinese financial workflows, spanning seven sub-domains (securities, funds, banking, insurance, futures, trusts, asset management) and three difficulty tiers (basic business analysis, asset decision support, strategic risk management). The authors evaluate 10 agents in a zero-shot setting, report ChatGPT (DeepResearch) as best with a weighted accuracy of 48.9%, compare against financial experts (84.7%) and non-experts (46.9%), and present an error analysis identifying recurring failure patterns. The benchmark is claimed to be the first end-to-end agent benchmark for the financial domain, with partial data released on GitHub.","tokens_in":23918,"tokens_out":3261,"duration_ms":32471,"significance":"If the benchmark is sound, it is a useful contribution that addresses a real evaluation gap: existing financial benchmarks are largely text- or image-centric, while FinGAIA explicitly targets multi-step, multi-tool agent behavior. The construction pipeline is substantial: four finance professors designed scenarios, six trained annotators created tasks, and four industry experts reviewed every question; human baselines were recruited independently of annotation. The qualitative finding that agents remain far below financial experts is plausible and consistent with prior results. However, the headline numerical claims rest on a ground-truth key that is not frozen and contains at least one internal contradiction, so the exact figures (48.9% vs. 84.7%) are not yet reproducible. The benchmark's design and partial release are valuable, but the evaluation infrastructure needs hardening before the quantitative conclusions can be accepted.","major_comments":[{"comment":"The gold answer for the margin-balance task is given as '13,510 yuan' in the 'Correct answer' line, while the same figure's error analysis states that the correct value is 15,440 yuan (=19,300 × 0.80). Only one of these can be the gold answer; this internal contradiction means the answer key is not self-consistent and the aggregate accuracy numbers cannot be fully trusted until the entire answer key is audited.","section":"Figure 16"},{"comment":"Many tasks require retrieving live data from official websites and market platforms (fee tables, index drawdowns, fund holdings, intraday quotes, and dated market reports). The Limitations section concedes that 'certain FinGAIA tasks may inherently involve dynamic elements,' yet the paper provides no snapshots, versioned pages, or frozen answer files. If any referenced page changes, the gold answers change and the reported 48.9% and expert-comparison numbers become unreproducible. The authors should provide date-stamped snapshots or a public, versioned answer key.","section":"Limitations; Figures 5, 8–10"},{"comment":"Table 2 reports ChatGPT's RMA score as 57.1, while the main text cites 'RMA 60.0' and Table 5 (based on the assessable-task subset after excluding unsupported file formats) reports 60.0. The weighted average 48.9 in Table 2 is not reconciled with these differing RMA values, and the exclusion policy is described only in a table caption. The authors must state explicitly which task set and which RMA value are used to compute the headline weighted average, and whether all subsequent cross-agent comparisons use the same set.","section":"Table 2; Results; Table 5"}],"minor_comments":[{"comment":"The section states that the authors 'identified six fundamental limitations' but then enumerates five (Data Type Handling Error, Financial Terminological Bias, Operational Process Awareness Barrier, Hallucinatory Financial Reasoning, Entity-Causation Misidentification). The Abstract and Conclusion correctly say five; the count should be aligned.","section":"Error Analysis"},{"comment":"The figure contains the typo 'Market Tirend Forecasting'; it should read 'Market Trend Forecasting'.","section":"Figure 1"},{"comment":"Cashcat DeepResearch is excluded with the note that it does not throw errors on unsupported files, but this exclusion rule is stated only in the table caption. It should be part of the main experimental protocol so readers understand how the assessable set was defined.","section":"Table 5"},{"comment":"All reported accuracies are point estimates without confidence intervals or variance measures. Given the 407-item pool and the demonstrated instability of dynamic tasks, the authors should report bootstrap confidence intervals or per-task variance to support cross-agent comparisons.","section":"Results; Evaluation Methods"},{"comment":"The human comparison in Table 3 uses a 50-question subsample, but the manuscript does not describe how this subsample was selected (e.g., random stratified vs. expert-chosen). This matters because the expert vs. agent gap is a headline result, and sampling bias could affect the comparison.","section":"Comparative Analysis"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is timely and the qualitative conclusion is plausible, but the answer-key inconsistency and the lack of frozen snapshots make the precise numerical claims unsuitable for publication as-is. I would also ask the authors to verify the novelty claim 'first agent benchmark closely related to the financial domain' against recent financial agent benchmarks (e.g., FinanceBench, FinAgent evaluation suites) before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about FinGAIA. First, it is a real attempt to fill a gap: existing finance benchmarks are mostly QA, and general agent benchmarks like GAIA aren't finance-specific. The authors built 407 expert-validated tasks across seven financial domains and three difficulty tiers, evaluated ten agents, and compared them with independent human baselines. That construction effort is genuine, and the qualitative result—ChatGPT at 48.9% versus experts at 84.7%—is plausible. Second, the exact numbers are not yet trustworthy because the gold answer key is neither frozen nor internally consistent. Figure 16 lists the correct answer as 13,510 yuan, while the error analysis in the same figure says the correct value is 15,440 yuan (=19,300 × 0.80). Both cannot be right. That is a load-bearing flaw because the headline aggregate is computed over this key.\n\nWhat's actually new: the benchmark artifact itself, the scenario taxonomy, and the zero-shot evaluation of ten agents. The methodology deliberately follows GAIA (web browsing, document parsing, code execution, multi-step reasoning), which is fine—the finance-specific task design is where the value is. The human baseline recruitment is clearly described and independent of annotation, so the agent-vs-expert comparison is not circular.\n\nSoft spots, in proportion. The biggest is answer stability. Many tasks require live data from official websites, fee tables, intraday quotes, and dated market data. The Limitations section concedes dynamic elements, but the paper ships no snapshots, no versioned pages, no frozen answer files. If any of those pages change, the scores change. That alone would be enough for a conditional verdict, but the Figure 16 contradiction shows the fragility is already present. There are smaller internal inconsistencies: the main text says ChatGPT's RMA is 60.0 while Table 2 says 57.1 (the 60.0 comes from Table 5 after exclusions); the error analysis says \"six fundamental limitations\" but lists five. Also, the release is only partial data with no evaluation code and no contamination check, so independent reproduction is not possible today.\n\nThe central qualitative conclusion holds up in spirit: agents are far below financial experts on realistic multi-step financial tasks. But the precise gap, 35+ points, is not yet established to the precision implied.\n\nWho this is for: anyone building or evaluating financial agents, and benchmark designers who care about stability in dynamic domains. It deserves a serious referee—this is not a desk reject. The revision needs to fix the gold-answer inconsistencies, freeze or snapshot the dynamic tasks, release the full dataset and evaluation code, and add a contamination check. I'd want to see that before citing the numbers.\n\nRecommendation: send to peer review, conditional on major revision. The effort and domain value are real; the reproducibility work is unfinished.","headline":"A serious and useful financial agent benchmark whose headline numbers are not yet reproducible because the answer key is neither frozen nor internally consistent.","tokens_in":24576,"tokens_out":2628,"would_cite":false,"duration_ms":24125,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FinGAIA is a new 407-task benchmark for financial AI agents, and its top-scoring agent reaches 48.9 percent accuracy, a gap of more than 35 points behind human experts.","keywords":["FinGAIA","AI agent benchmark","financial domain","multi-tool collaboration","zero-shot evaluation","Chinese financial NLP","agent evaluation","error analysis"],"falsifier":"Randomly sample FinGAIA tasks whose answers depend on a specific official product rate table, annual report, or price history, then re-check the cited page against the published answer on a later date; if a substantial fraction of answers no longer match the current page, the benchmark's accuracy numbers, including the 48.9 percent headline, are time-dependent rather than stable measurements.","tokens_in":23438,"feed_emoji":"💹","tokens_out":11777,"duration_ms":106633,"temperature":0.7,"pith_summary":"FinGAIA is a benchmark for judging whether AI agents can complete real financial work in Chinese from start to finish—identifying a company logo, pulling a fee table from its official website, writing Python to analyze it, and returning a formatted answer. The paper claims it is the first agent benchmark built for the financial domain, with 407 expert-validated tasks across securities, funds, banking, insurance, futures, trusts, and asset management, organized into three tiers of scenario depth. In a zero-shot evaluation of ten agents, the best performer scored 48.9 percent overall, above finance undergraduates but more than 35 points below financial experts' 84.7 percent, a gap the paper attributes to five systematic failure patterns rather than to luck. Readers should care because the benchmark converts a vague worry—'agents are not ready for finance'—into a measurable gap with a named set of weaknesses to work on.","feed_headline":"First finance-agent benchmark puts top agents at 48.9 percent","feed_subtitle":"The best AI agent trails financial experts by over 35 points on real Chinese financial workflows, the paper reports.","key_machinery":"The machinery is the three-tier task architecture that forces end-to-end behavior. FinGAIA's 407 tasks are partitioned into basic business analysis (89 tasks, up to five steps and one or two tools), asset decision support (185 tasks, five to seven steps and more than two tools), and strategic risk management (133 tasks, about ten steps with sequential tool invocation and parameter tuning). Each task pairs a realistic prompt with an expert-validated answer and an explicit solution path, so the benchmark can score any agent purely by whether its final output matches the key. The depth tiers do the causal work: they make rising difficulty correspond to rising demands on tool coordination, which is what separates agentic competence from a language model's ability to answer a question in one shot.","core_discovery":"The paper's central discovery is that current AI agents can pass basic financial-analytic chores but break down when tasks demand the full workflow: multimodal input, live web lookup, document parsing, code execution, and coordinated multi-tool reasoning. FinGAIA makes this visible by constructing tasks that cannot be solved by text QA alone, and the evaluation shows a clear gradient—agents do best on operational analytics and worst on strategic risk, while experts hold roughly 84 percent across all tiers. The authors interpret the persistent expert gap, especially on strategic risk tasks, as evidence that agents lack combined domain comprehension and operational process awareness, and they identify five recurring error types—cross-modal alignment deficiency, financial terminological bias, operational process awareness barrier, hallucinatory financial reasoning, and entity-causation misidentification—that account for the failures. The claim is that this benchmark, and only this benchmark, currently measures an agent's end-to-end financial capability in a way that tracks real business depth.","pith_inferences":["A likely but unstated consequence is that the 48.9 percent figure is closer to a ceiling than a floor: real financial deployments involve ambiguous requests and unvetted web sources, so agents would likely perform worse outside the benchmark's carefully annotated conditions.","The decision not to freeze or version the live sources means FinGAIA is better understood as a methodology for building financial agent benchmarks than as a permanent scoreboard; scores will drift as fee tables, product lists, and market data change.","An extension that would test the paper's main claim directly is to fine-tune an agent specifically on the five error categories and re-run the same 407 tasks; if the expert gap narrows substantially, the error taxonomy is doing real causal work, and if not, the gap may come from something the taxonomy does not capture.","Because the tasks are in Chinese and reference Chinese regulatory and market sources, FinGAIA could double as a probe of an agent's Asia-market data coverage, which may matter more for real deployment than raw reasoning ability."],"forward_implications":["FinGAIA establishes a reproducible yardstick: any new agent can be scored on the same 407 tasks and zero-shot protocol against the 84.7 percent expert baseline.","The 48.9 to 13.1 percent spread across ten agents shows that agent quality in finance is highly stratified, so benchmark results can separate strong general-purpose agents from weaker ones.","The five recurring error patterns give concrete, testable targets—for instance, training on regulatory process rules or financial terminology—that future work can use to close the expert gap.","Because tasks draw on live official websites and real market data, FinGAIA also tests whether an agent has current access to Chinese financial information, not just financial knowledge."],"supporting_citations":[{"why":"It provides the end-to-end, live-web, multi-tool evaluation model that FinGAIA transplants into the financial domain, and it supplies the closest baseline FinGAIA must beat to claim novelty.","marker":"Mialon et al. 2023"},{"why":"It is the existing Chinese financial knowledge benchmark that FinGAIA contrasts with, showing the text-only QA format FinGAIA goes beyond.","marker":"Guo et al. 2024"},{"why":"It represents numerical-reasoning-over-financial-data benchmarks that FinGAIA extends by adding multi-file, multi-tool execution.","marker":"Chen et al. 2021"},{"why":"It is a Chinese financial language-understanding benchmark whose static question format FinGAIA is designed to surpass.","marker":"Zhu et al. 2024"},{"why":"It supplies another comprehensive Chinese financial LLM benchmark, used to justify FinGAIA's agent-specific contribution.","marker":"Nie et al. 2024"},{"why":"It is the multimodal finance benchmark that FinGAIA distinguishes itself from by requiring tool use and multi-step workflows rather than image-text QA.","marker":"Gan et al. 2024"},{"why":"It is a simulated-web agent benchmark whose dependence on fixed DOM structures motivates FinGAIA's use of real financial websites.","marker":"Zhou et al. 2023"},{"why":"It is a system-interaction agent benchmark that FinGAIA contrasts with, since it does not cover finance-specific regulatory and market workflows.","marker":"Liu et al. 2023"},{"why":"It tests API invocation in closed environments, and FinGAIA cites it as an example of benchmarks that allow pattern memorization rather than genuine reasoning.","marker":"Qin et al. 2023"}],"fun_headline_variants":["First finance-agent benchmark: best AI scores 48.9%","AI agents lag financial experts by 35 points on new benchmark","FinGAIA: AI agents fail strategic risk tasks, top score 48.9%","Benchmark shows AI agents fall short in real-world finance","New finance benchmark: top AI agent barely beats non-professionals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's scores assume that the live websites, fee tables, and market data referenced by its tasks will keep matching the published answer keys; nothing in the release freezes or versions those sources, so the reported numbers can decay or become unreproducible as the web changes.","fun_headline_variants_meta":{"raw":{"variants":["First finance-agent benchmark: best AI scores 48.9%","AI agents lag financial experts by 35 points on new benchmark","FinGAIA: AI agents fail strategic risk tasks, top score 48.9%","Benchmark shows AI agents fall short in real-world finance","New finance benchmark: top AI agent barely beats non-professionals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1323,"prompt_tokens":976,"completion_tokens":347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":254}},"tokens_in":592,"tokens_out":347,"duration_ms":4106,"temperature":1.0,"reasoning_tokens":254,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:55:40.930299+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample FinGAIA tasks whose answers depend on a specific official product rate table, annual report, or price history, then re-check the cited page against the published answer on a later date; if a substantial fraction of answers no longer match the current page, the benchmark's accuracy numbers, including the 48.9 percent headline, are time-dependent rather than stable measurements.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the end-to-end, live-web, multi-tool evaluation model that FinGAIA transplants into the financial domain, and it supplies the closest baseline FinGAIA must beat to claim novelty."},{"cited_title":"Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset","cited_arxiv_id":"2405.10542","evidence_quote":"It is a Chinese financial language-understanding benchmark whose static question format FinGAIA is designed to surpass."}],"review_version":1}