{"id":"39cca235-74e3-4e76-98ed-3352f7df66f2","arxiv_id":"2501.18062","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FinanceQA, a new finance-work benchmark, shows top LLMs score below 50% overall and under 5% on assumption-based questions; fine-tuning GPT-4o on generated variants raises total accuracy to 57%.","lead":"FinanceQA is a new benchmark built from real-world financial-analysis tasks, and the paper reports that top AI models answer fewer than half of its questions correctly. It also shows that fine-tuning GPT-4o on FinanceQA-style data raises its score from 39% to 57%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Without an expert-human baseline under the same exact-match rubric and without per-category N, the headline failure rates are uncalibrated; low LLM scores could reflect grading strictness or tiny samples rather than an LLM-specific deficit.","rationale":"I read the paper in good faith. The benchmark is useful and the release is a real asset; the annotator expertise and the fine-tuning result are plausible. The reader’s weakest assumption, the Costco-only tactical set, is candidly acknowledged in Section 8 and would mainly weaken generalization to other industries. I think the more load-bearing gap is internal calibration: the claim that LLMs “fail to meet strict accuracy requirements” rests entirely on scores produced by an exact-match rubric with no human reference point. The paper asserts that prompts and answers are unambiguous, but assumption-based questions by definition involve analyst judgment, so exact-match grading may be too strict. A human baseline is a concrete, inexpensive control that would settle whether the low numbers are a model deficit or an artifact of the evaluation protocol. Because this does not change the overall conditional verdict—the paper is publishable with conditions—I recommend keeping the reader’s CONDITIONAL stance rather than escalating or rejecting.","tokens_in":11326,"tokens_out":9464,"duration_ms":115451,"concrete_test":"Download the released HuggingFace dataset and compute per-category N and unique tickers. Then have at least three experienced financial analysts (matching the §4.1 background) answer a stratified sample of 60 FinanceQA questions (20 basic, 20 assumption, 20 conceptual) using the exact same contexts and exact-match rubric, with responses scored blindly by two annotators. If analyst accuracy is high (e.g., ≥85%) and inter-annotator agreement is high, the low LLM scores stand. If analyst accuracy is within ~15 points of o1, or inter-annotator agreement is poor, the benchmark’s grading and ambiguity are confounds and the headline claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract, Section 5.2, Table 1) interprets FinanceQA scores as evidence that current LLMs fail the accuracy requirements of financial institutions. Two conditions are needed: the benchmark tasks are a fair sample of real analyst work, and the reported failure rates are stable. The paper provides no expert-human baseline under the same exact-match protocol, no inter-annotator reliability, and no per-category N. This is most acute for assumption-based questions (2.2–4.3%, Table 1): an “assumption” question has several defensible answers, so the author-set expected answer plus exact-match grading can manufacture failure. For example, the variable-lease question in Section 4 requires assuming a ratio of variable lease costs to operating lease costs; a different but reasonable assumption yields a different final number and would be marked wrong. If trained analysts also score low under this rubric, the low LLM scores reflect grader strictness, not an LLM-specific shortfall. Additionally, without per-category N we cannot distinguish 1/45 from 1/5 in the assumption column. The Costco-only limitation (Section 8) narrows generalization but does not affect internal validity; the human-baseline/calibration gap affects every number in Table 1.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FinanceQA, a benchmark for evaluating LLMs on financial analysis tasks that the authors claim mirror real-world work by junior investment professionals. The benchmark consists of 'tactical' questions derived from a single company's 10-K (Costco), segmented into basic and assumption-based types, plus 'conceptual' questions. The authors evaluate GPT-4o, o1, Claude-3.5-Sonnet, and Llama-3.3-70B-Instruct, reporting low accuracies, especially on assumption-based questions (2.2–4.3%). They then fine-tune GPT-4o on a synthetic dataset derived from human-written finance questions and report substantial accuracy improvements (39.2% to 56.8% total), along with a paired t-test they claim shows statistical significance. The paper's central claim is that current LLMs fail the strict accuracy requirements of financial institutions, and that fine-tuning on high-quality financial reasoning data can improve performance.","tokens_in":11529,"tokens_out":4508,"duration_ms":46443,"significance":"If the benchmark is well-constructed and the results are reliable, FinanceQA could be a useful contribution to the evaluation of LLMs in professional finance, a domain where existing benchmarks such as FinanceBench are criticized for being too easy. The paper also ships a publicly released dataset and demonstrates a fine-tuning pipeline, which are praiseworthy. However, the current evidence is insufficient to support the strong claims: the benchmark lacks published dataset sizes, confidence intervals, a human expert baseline, and a statistically sound analysis of the fine-tuning gains. The qualitative direction (LLMs struggle with exact-match financial calculation under incomplete information) is plausible, but the quantitative claims need substantial strengthening.","major_comments":[{"comment":"The paper never reports the number of questions in each category (basic, assumption, conceptual) or in total. Without per-category N, the percentages in Table 1 are uninterpretable—for example, the assumption-based accuracies of 2.2–4.3% could represent a single question difference (1/45 vs. 1/23) or a larger effect. This omission also precludes confidence intervals and makes it impossible to evaluate the stability of the central claim that models 'fail less than 5%' on assumption-based questions. Please report exact question counts for every cell, and ideally provide binomial confidence intervals or exact per-question results.","section":"§5.2, Table 1"},{"comment":"The paired t-test is conducted on only three paired aggregates (basic, assumption, conceptual), giving df = 2. With n = 3, the test cannot meaningfully assess statistical significance: the assumptions of the t-test (normality and independence) are unverifiable, and the three observations are not independent outcomes but aggregated proportions from the same benchmark. Moreover, the test ignores question-level variation entirely. The reported p-value of 0.034 therefore does not support the claim that the fine-tuning improvements are 'statistically significant at the 95% confidence level.' Please reanalyze at the question level (e.g., a paired permutation test or a mixed-effects model across individual questions), or clearly limit the significance claim to the three category aggregates and explain why a t-test with df=2 is appropriate.","section":"§6.2, Figure 7"},{"comment":"No expert-human baseline is provided under the same exact-match rubric. The paper's central interpretation—that low LLM scores indicate a failure to meet 'strict accuracy requirements of financial institutions'—requires knowing how well trained analysts perform on the same questions with the same grading rules. The concern is especially acute for assumption-based questions, where the expected answer depends on an author-chosen assumption (e.g., the variable-lease ratio in Section 4). Different but defensible assumptions would yield different final numbers and would be marked wrong under exact match. Without a human baseline, the low LLM scores could reflect grading strictness rather than an LLM-specific deficit. Please add at least a small-scale human expert evaluation under identical conditions, or substantially temper the claim that the results demonstrate an LLM-specific failure.","section":"§5.2, Table 1; §4"},{"comment":"All tactical questions are based on a single company's 10-K (Costco), as the authors acknowledge in Section 8. The benchmark's external validity is therefore extremely narrow: Costco's financial structure is not representative of healthcare, energy, financial services, or other industries, and a single company cannot capture the diversity of 'real-world tasks' the abstract claims to evaluate. The paper should either restrict its generalization claims to the Costco-like retail sector or include additional companies before using the benchmark to draw broad conclusions about LLM performance in professional finance.","section":"§4, §8"}],"minor_comments":[{"comment":"The claim that 'models fail approximately 60% of realistic tasks' conflicts with the best model's performance: o1 fails about 51.3% (accuracy 48.7% per Table 1). The 60% figure appears to be an average across the four models; please state this explicitly or use a more accurate summary.","section":"Abstract"},{"comment":"The paper states that annotators manually verified questions and answers to be unambiguous, but no inter-annotator agreement statistics are reported. Please add a measure of agreement (e.g., Cohen's kappa on a subset) or at least a detailed description of the double-checking procedure.","section":"§4.1, Dataset Annotation"},{"comment":"The sentence 'unless directly prompted, it won’t follow know that it needs to calculate those metrics' contains a typo ('follow know' should be 'know' or 'follow through').","section":"§5.3"},{"comment":"The section heading 'Handing Incomplete Information' should be 'Handling Incomplete Information'.","section":"§3.4"},{"comment":"The phrase 'makes the model suspect to the same dilemma' is awkward; 'susceptible' would be more natural.","section":"§2"},{"comment":"The correlation matrix figure is not described in the text; it is unclear which variables are correlated (presumably accuracies across question types). Please explain the figure in a sentence and provide the numeric correlation values if possible.","section":"Figure 4"},{"comment":"The fine-tuning dataset consists of 9,078 rows, but the paper does not report how many unique human-written seeds were used, how many synthetic variations were generated per seed, or the ratio of tactical to conceptual questions. Please clarify the dataset composition.","section":"§6.1"},{"comment":"The specific model versions are not fully identified (e.g., o1-preview vs. o1-2024-12-17, and the exact release dates for GPT-4o, Claude-3.5-Sonnet, and Llama-3.3-70B-Instruct). Please provide version identifiers and access dates for reproducibility.","section":"§5.1, Setup"},{"comment":"What is labeled 'Figure 7' is in fact a table of percentages and percent differences; please renumber or relabel it as a table.","section":"Figure 7"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful idea and the dataset release is a positive step, but in its current form the evidence for the central claims is thin: missing dataset sizes and confidence intervals, no human baseline, and an inappropriate significance test. These are fixable, so I recommend major revision rather than rejection. The authors should also consider repositioning the benchmark as a pilot study on a narrow set of tasks rather than a comprehensive evaluation of LLMs in finance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth reading and worth citing, but its central quantitative claim — that LLMs fail about 60% of realistic finance tasks — is not yet calibrated. The benchmark itself is a real contribution: the assumption-based tactical questions capture something that FinQA and FinanceBench don't, and the dataset is public. The authors used finance practitioners to generate questions, which shows in the task design, and the fine-tuning experiment is a reasonable sanity check.\n\nWhat's genuinely new is the assumption-based question type. Asking a model to notice missing information, make a defensible assumption, and produce a number is much closer to what analysts actually do than extracting figures from a 10-K. The separation into basic, assumption, and conceptual categories is sensible, and the correlation analysis is a reasonable way to show the categories are not redundant.\n\nSoft spots, in order of importance. First, there is no expert-human baseline under the same exact-match rubric. This matters because assumption questions have multiple defensible answers. The variable-lease example in Section 4 asks the model to assume a ratio; a different but reasonable ratio gives a different number, which would be marked wrong. Without knowing how trained analysts score under this protocol, a 2.2% accuracy could reflect grader strictness rather than an LLM-specific deficit. Second, the paper does not report per-category N, so we cannot tell whether the assumption-based results are 1/45 or 1/5. Third, the abstract says models fail about 60% of tasks, but Table 1 shows o1 at 48.7% and GPT-4o at 39.2%; the 60% figure appears to be a rough average or a misstatement. Fourth, the paired t-test on three category-level aggregates (df=2) is too fragile to carry the claim of statistical significance; a per-item test or bootstrap with the actual data would be far more convincing. Fifth, the tactical questions come from a single company's 10-K, which the authors honestly flag in Section 8 — that limits generalization to 'real-world tasks' across finance, though it does not affect internal validity.\n\nThe citation pattern is fine. The related work discussion of FinQA and FinanceBench is fair. The limitation section is candid.\n\nWho this is for: anyone building or evaluating LLMs for finance, and benchmark designers generally. It deserves a serious referee — the design insight is solid and the dataset release supports reproducibility — but the paper needs revision before the numbers can be accepted. Add a human baseline, report per-category N, fix the abstract overclaim, and strengthen the statistics. My recommendation: send it to review, with a clear request for those additions.","headline":"FinanceQA is a genuinely useful new benchmark for job-realistic finance LLM evaluation, but its headline failure rates need a human baseline and per-category N before they can be taken at face value.","tokens_in":12067,"tokens_out":1777,"would_cite":true,"duration_ms":18991,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FinanceQA is a benchmark built from real analyst work; its central claim is that leading LLMs fail about 60% of realistic financial analysis tasks, with the best model, o1, at 48.7%.","keywords":["financial reasoning","LLM evaluation benchmark","fine-tuning","accounting conventions","incomplete information","10-K analysis","assumption generation","investment analysis"],"falsifier":"Build a held-out set of tactical FinanceQA-style questions from 10-Ks of five companies in different industries, such as healthcare, energy, technology, industrials, and retail, and run the same four models with the same exact-match scoring. If accuracy on this multi-company set is materially higher than 48.7% overall, or well above 5% on assumption-based questions, then the paper's central failure-rate claim is an artifact of the single-company test set rather than a general property of LLMs in finance.","tokens_in":11070,"feed_emoji":"📊","tokens_out":6333,"duration_ms":66183,"temperature":0.7,"pith_summary":"FinanceQA is a benchmark built from the daily work of junior investment professionals: recalculating metrics by hand from primary filings, following accounting and valuation conventions, and forming assumptions when filings omit needed data. The paper's central claim is that current large language models fail this work at an unacceptable rate: the best model tested, OpenAI's o1, answers 48.7% of FinanceQA questions correctly, and every model answers fewer than 5% of assumption-based questions correctly. This matters because the financial industry has a near-zero tolerance for errors, so even a model with 80% accuracy would still be unusable if every figure must be verified line by line. The authors also claim that fine-tuning GPT-4o on 9,078 rows of synthetic FinanceQA-style data raises total accuracy from 39.2% to 56.8%, with the largest relative gain on assumption-based questions. If the claim is right, the bottleneck for professional-grade financial AI is not model scale but the availability of high-quality, convention-compliant training data.","feed_headline":"OpenAI o1 scores 48.7% on real finance tasks","feed_subtitle":"FinanceQA tests what analysts actually do; fine-tuning lifts GPT-4o from 39.2% to 56.8%.","key_machinery":"The central object is the FinanceQA question-answer pair with an expert-written gold answer. Each tactical question carries a context window assembled only from relevant 10-K sections, forcing models to recalculate from primary sources rather than rely on memory; assumption-based questions intentionally omit a datum and require the model to state and apply a plausible assumption, such as scaling variable lease costs to estimate variable lease assets. Questions are scored as exact matches, so a single missed adjustment, such as failing to add back operating lease costs in an EBITDA calculation, fails the item. The fine-tuning arm uses synthetic variations of human-written rows, 9,078 total, that preserve the reasoning template while randomizing line items and numbers, so the model learns the procedure rather than memorizing an answer. The benchmark's design is question-first, mirroring how analysts start with a question and then search documents, in contrast to context-first benchmarks that write questions from facts already present in the text.","core_discovery":"FinanceQA is a testing suite of tactical and conceptual questions modeled on real analyst tasks. Tactical questions are grounded in a company's 10-K, with a context window made from the relevant sections, and split into 'basic' questions and 'assumption-based' questions that require an assumption the filing does not supply; conceptual questions test financial reasoning without a document. Evaluating GPT-4o, o1, Claude-3.5-Sonnet, and Llama-3.3-70B-Instruct with exact-match scoring, the paper finds total accuracies of 39.2%, 48.7%, 39.9%, and 31.1%, respectively, and assumption-question accuracies at or below 4.3%. It interprets the pattern as a specific capability gap: models handle textbook-style conceptual reasoning relatively well but fail precision recalculations, adherence to accounting conventions, and assumption generation under incomplete information. The paper then fine-tunes GPT-4o on synthetic variations of human-written financial reasoning pairs, producing a 44.9% relative improvement in total accuracy and a 690.9% relative improvement on assumption-based questions, with a paired t-test p-value of 0.034. The stated conclusion is that higher-quality, work-aligned training data is the missing ingredient, and that FinanceQA provides a way to measure progress toward it.","pith_inferences":["An obvious extension the authors do not run is a multi-company, multi-industry FinanceQA, which would show whether the near-universal failure on assumption questions is a structural LLM limitation or partly a Costco-specific reporting artifact.","The synthetic fine-tuning recipe, human-written templates with randomized numbers and line items, suggests a cheap scaling path; if the effect generalizes, increasing synthetic diversity could push accuracy past the materiality threshold, but the current 15.2% assumption accuracy is still far too low for unsupervised use.","The same question-first benchmark methodology could transfer to other precision professions, such as legal drafting or clinical documentation, by extracting work-critical tasks from primary documents and expert answers.","The paper's discontinuity argument implies that small accuracy gains around the 50-60% range are commercially meaningless until a much higher threshold is crossed; a useful next test would be measuring where human analysts stop verifying every line."],"forward_implications":["Financial institutions cannot adopt current LLM output as work product: even a model scoring 80% would require line-by-line human verification that takes longer than doing the analysis from scratch.","Assumption-based reasoning is the largest measured gap: with all models below 5%, any useful financial LLM must be trained and evaluated specifically on incomplete-information tasks.","Fine-tuning on synthetic, convention-compliant reasoning data can substantially improve performance on real-world financial tasks, including a 44.9% relative total accuracy gain for GPT-4o.","Benchmarks that start from facts in documents and ask for extraction overstate LLM ability; question-first benchmarks like FinanceQA measure the tasks professionals actually perform."],"supporting_citations":[{"why":"Provides FinQA, the context-first numerical reasoning dataset the paper argues is misaligned with analyst workflow and uses as contrast.","marker":"Chen et al., 2021"},{"why":"Provides FinanceBench, characterized as basic extraction and simple calculation tasks that do not require accounting conventions.","marker":"Islam et al., 2023"},{"why":"Corporate valuation textbook used by annotators as authority for stock-flow averaging and other accounting conventions.","marker":"Holthausen et al., 2014"},{"why":"Valuation textbook used as authority for standards on hand-spread metrics like diluted shares and EBITDA adjustments.","marker":"Koller, et al., 2020"},{"why":"Defines the analyst's question-first workflow used to justify the benchmark's design.","marker":"CFA Institute, 2024"},{"why":"Supports the claim that companies are incentivized to inflate EBITDA and that analysts must reconstruct it.","marker":"S&P Global Ratings, 2024"},{"why":"Establishes that operating leases are capitalized under current standards and should be treated as debt-like, motivating the EBITDA add-back test.","marker":"Ernst & Young, 2021"}],"fun_headline_variants":["LLMs fail 60% of FinanceQA real-world tasks","o1 scores 48.7% on FinanceQA analyst tasks","Fine-tuned GPT-4o jumps 44.9% on FinanceQA","Assumption questions trip LLMs: ≤4.3% on FinanceQA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported failure rates assume that tactical questions built from one company's 10-K (Costco) represent the full range of real-world financial analysis across industries, company types, and reporting styles.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fail 60% of FinanceQA real-world tasks","o1 scores 48.7% on FinanceQA analyst tasks","Fine-tuned GPT-4o jumps 44.9% on FinanceQA","Assumption questions trip LLMs: ≤4.3% on FinanceQA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000321,"raw_usage":{"total_tokens":1829,"prompt_tokens":989,"completion_tokens":840,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":770}},"tokens_in":605,"tokens_out":840,"duration_ms":8827,"temperature":1.0,"reasoning_tokens":770,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:48:22.799084+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a held-out set of tactical FinanceQA-style questions from 10-Ks of five companies in different industries, such as healthcare, energy, technology, industrials, and retail, and run the same four models with the same exact-match scoring. If accuracy on this multi-company set is materially higher than 48.7% overall, or well above 5% on assumption-based questions, then the paper's central failure-rate claim is an artifact of the single-company test set rather than a general property of LLMs in finance.","supporting_citations":[],"review_version":1}