{"id":"095cc863-adc6-4fc1-9d8c-658237a928da","arxiv_id":"2502.06329","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"FailSafeQA, a 220-example financial long-context benchmark, shows no tested LLM can both stay robust to input perturbations and refuse to hallucinate when context is missing or irrelevant.","lead":"FailSafeQA is a new benchmark that stresses financial LLMs with misspelled, incomplete, and out-of-domain questions, plus missing, OCR-corrupted, or irrelevant context. It finds that models that answer robustly under perturbation often hallucinate when the context is unhelpful, leaving large room for improvement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own headline numbers are not reproducible from its tables: the abstract and Section 7 say o3-mini fabricated information in 41% of cases, but Table 2 reports Context Grounding 0.63, implying 37%, and Table 2's Context Grounding columns are internally inconsistent for multiple models.","rationale":"The most load-bearing issue is not a theoretical worry about LLM judges in general; it is that the paper's central numbers do not survive contact with its own tables. The abstract and Section 7 claim o3-mini fabricated information in 41% of tested cases; Table 2's Context Grounding of 0.63 implies 37%. The table's aggregate Context Grounding column is also inconsistent with its reported QA/TG and Irrelevant/No-context subcolumns for multiple models, so a reader cannot reconstruct the headline metrics from the published data. The released dataset and prompts are a real asset, and the perturbation taxonomy is sensible; the central claim is framed as a hypothesis. But a benchmark whose headline percentages cannot be reproduced from its own results should not be accepted as a measurement until the raw ratings are released and the tables reconciled. The unvalidated Qwen judge is a related and deeper concern—no human agreement is reported, and the judge is itself an evaluated model—so the reader's CONDITIONAL verdict is appropriate; this stress-test strengthens that condition rather than overturning it.","tokens_in":20671,"tokens_out":8027,"duration_ms":67295,"concrete_test":"Download the released dataset, extract the judge ratings (or re-run the public judging prompts with Qwen2.5-72B-Instruct at temperature 0), and recompute Context Grounding, Robustness, and the exact o3-mini fabrication rate from the per-condition scores. If the recomputed numbers do not reproduce Table 2 and the abstract's 41% figure, the headline numbers and model ordering are unsupported by the paper's own data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that no model is simultaneously robust and context-grounded, with o3-mini fabricating in 41% of cases and Palmyra-Fin failing robustness in 17%—is only as strong as the scores in Tables 1–2. Those scores are not internally consistent. For o3-mini, Table 2 lists Irrelevant Ctx = 0.67, No Ctx = 0.51, and Ctx Grounding = 0.63; the average of the two context-failure conditions is 0.59, which would imply a 41% fabrication rate, but the table's own Context Grounding value implies 37%. This is not isolated: for Gemini 2.0 Flash Exp, Table 2 lists QA Ctx Grounding = 0.46 and TG Ctx Grounding = 0.74, yet reports aggregate Ctx Grounding = 0.77, while a weighted average over the dataset's 83%/17% QA/TG split would be about 0.51. Similar discrepancies appear across rows. The paper also reports no human rater agreement for Qwen2.5-72B-Instruct's judgments, and that judge is itself one of the evaluated models. The combination means the absolute percentages, the model ordering, and the claimed robustness-versus-grounding trade-off are not currently established from the published tables.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FailSafeQA, a 220-example long-context financial QA benchmark built from truncated 10-K filings, with six human-interface failure modes: misspelled, incomplete, and out-of-domain queries, plus missing, OCR-degraded, and irrelevant contexts. It evaluates 24 long-context LLMs using Qwen2.5-72B-Instruct as an LLM judge on a 1–6 relevance rubric, and summarizes performance through three metrics: Robustness (Eq. 2), Context Grounding (Eq. 3), and a beta-weighted Compliance score (Eq. 4). The headline findings are that the most robust model (OpenAI o3-mini) fails to refuse or fabricates in a large share of unanswerable cases, that the most compliant model (Palmyra-Fin-128k-Instruct) loses robustness in 17% of cases, and that no tested model is simultaneously robust and well-grounded.","tokens_in":20938,"tokens_out":14722,"duration_ms":119280,"significance":"The benchmark targets a real and under-served problem: dependability of financial QA systems under realistic user and document failures. Strengths include a public dataset, a clearly described perturbation taxonomy, citation-based judging to keep the judge's context short, and an explicit trade-off metric for answer-versus-refuse behavior. If the evaluation were independently validated, this would be a useful resource for the long-context and financial NLP communities. However, the current empirical payload—the 41% fabrication rate, the model ordering, and the robustness-versus-grounding trade-off—rests on a single unvalidated judge that is itself a scored model, and it conflicts with the numbers in Table 2. The benchmark may still be valuable, but this draft does not establish the claimed comparative results.","major_comments":[{"comment":"The Context Grounding column in Table 2 is not the value defined by Eq. (3). Under the paired dataset design, Context Grounding should be the equal-weighted average of the Irrelevant Ctx and No Ctx columns. For Gemini 2.0 Flash Exp this gives (0.81+0.66)/2 = 0.735, but the table reports 0.77; for OpenAI o3-mini it gives (0.67+0.51)/2 = 0.59, but the table reports 0.63; for Qwen2.5-72B-Instruct it gives 0.645, but the table reports 0.68. Moreover, the table's QA and TG Context Grounding columns cannot be sub-aggregates of the reported total: for Gemini 2.0 Flash Exp, 0.83×0.46+0.17×0.74 ≈ 0.51, not 0.77. The Compliance column, by contrast, is consistent with Eq. (4) if the equal-weighted average is used (for Palmyra-Fin, (0.95+0.66)/2 = 0.805 together with R=0.83 yields 0.81, matching the table). The paper therefore appears to use two different Context Grounding values in one table. This matters directly for the headline: the Abstract and §7's claim that o3-mini 'fabricated information in 41% of tested cases' corresponds to the equal-weighted average (1−0.59=0.41), whereas the reported Context Grounding of 0.63 implies 37%. The tables, the text, and the metric definitions must be reconciled before any of the comparative conclusions can be assessed.","section":"Table 2; §3.2, Eq. (3)"},{"comment":"All scores in Tables 1 and 2 are produced by Qwen2.5-72B-Instruct, which is itself one of the 24 evaluated models (it has a row in both tables). The paper reports no human-rater agreement, no judge calibration against the 1–6 rubric of Appendix C.3, and no check for judge self-preference. The citation-based argument in §4.2 explains why the judge's context is short, but it does not address systematic grading bias across model families or between answer and refusal responses. Because the entire model ranking and all absolute rates are outputs of this single judge, a modest self-preference or rubric miscalibration would change the reported scores and the claimed trade-off. I would like to see a stratified validation sample (covering all 24 models and both answerable and unanswerable scenarios) double-scored by human annotators, with per-model and per-scenario agreement reported.","section":"§4.2 Judging; Tables 1–2"},{"comment":"The Compliance metric is a free-parameter harmonic score with beta=0.5, and no sensitivity analysis is given. This is load-bearing for the 'most compliant model' claim because Robustness and Context Grounding move in opposite directions for several leading models: OpenAI o3-mini has R=0.90 with Context Grounding around 0.59–0.63, while Palmyra-Fin-128k-Instruct has R=0.83 with Context Grounding around 0.805–0.83. The ranking under Eq. (4) can change as beta moves, so the paper should report Compliance for a range of beta values (e.g., 0.25, 0.5, 1, 2) or explicitly show that the ordering asserted in Section 7 is stable.","section":"§3.2.1, Eq. (4)"}],"minor_comments":[{"comment":"The text says 'The best Context Grounding score of 0.80 is achieved by Palmyra-Fin-128k-Instruct', but Table 2 lists Palmyra-Fin's Context Grounding as 0.83; one of these is wrong.","section":"§5"},{"comment":"Equation (3) is typeset as 'G = 1/nj 2X j=1 nX i=1 ...', which is unreadable; it should be written as a simple average over the two context-failure conditions, e.g., G = (1/2)(G_Missing + G_Irrelevant).","section":"§3.2, Eq. (3)"},{"comment":"The text states that each data point includes 'an irrelevant query', but the perturbation described in §2.3 is an irrelevant context; please clarify the wording and whether the 'irrelevant query' is a question paired with an irrelevant 10-K excerpt.","section":"§2.4"},{"comment":"No confidence intervals or significance tests are reported. With 220 examples, adjacent scores in the tables often differ by 0.01–0.02, which is well within the standard error of a proportion; bootstrap confidence intervals for at least the headline scores would help the reader judge whether the observed ordering is meaningful.","section":"Tables 1–2"},{"comment":"The 10% cap on character error probability is described as 'empirically chosen', but no sensitivity analysis or distribution of injected error rates is reported; please state how Robustness changes under a lower or higher cap, or provide the error-rate distribution.","section":"§2.3, OCR Errors"},{"comment":"The captions of Figures 5 and 6 are close to verbatim copies, and Figure 5's caption contains sentences about Context Grounding even though the title says 'Robustness vs. Query Type'; the captions should be corrected to describe the actual content of each figure.","section":"Figures 5–6"}],"recommendation":"major_revision","confidential_remarks":"The benchmark idea and public dataset are solid enough to warrant a revision. However, because the paper's top-performing model is the authors' own Palmyra-Fin-128k-Instruct and the judge is a Qwen model, the absence of independent judge validation is a particularly acute concern. I would advise the editor to require human-annotated agreement results and corrected tables in the revision. The internal inconsistencies in Table 2 are likely fixable, but they are substantive rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper builds a genuinely useful benchmark artifact for financial long-context QA—public dataset, sensible perturbation taxonomy, and a clean separation between Robustness and Context Grounding. That separation is the right lens, and the Compliance metric captures a real trade-off between answering and refusing. Credit is due for making the dataset and prompts public and for laying out the generation pipeline.\n\nBut the reported numbers do not hold together. The stress-test is correct: Table 2 is internally inconsistent in multiple rows. For Gemini 2.0 Flash Exp, the aggregate Context Grounding is 0.77, but the average of the two context-failure columns (0.81 and 0.66) is 0.74, and the QA/TG columns imply roughly 0.51 under the paper's own 83/17 split. For o3-mini, the abstract says 41% fabrication, while the table's Context Grounding of 0.63 implies 37%; the row's own average is 0.59. These are not footnotes—they are the load-bearing evidence for the headline claims.\n\nTwo more soft spots. The judge is Qwen2.5-72B-Instruct, which is itself one of the 24 scored models, and no human agreement or calibration is reported. Self-preference is a known confound in LLM-as-a-judge, and the paper's argument that citation-based judging is easier does not remove the need for a sanity check. Also, beta=0.5 in the Compliance metric is a free parameter with no sensitivity analysis; the title of 'most compliant model' depends on it.\n\nThe benchmark is a useful contribution, but the empirical conclusions should be treated as provisional. This deserves a serious referee: the question is important, the dataset is public, and the fixes are straightforward—recompute the tables, validate the judge against human raters, and show how the rankings move with beta. I would not cite the current numbers until then.","headline":"A useful benchmark with a promising robustness/refusal split, but the tables contradict themselves and the headline rankings are not reproducible as printed.","tokens_in":21507,"tokens_out":5713,"would_cite":false,"duration_ms":42800,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces FailSafeQA, a 220-item long-context financial QA benchmark, and claims that no tested LLM simultaneously maintains accurate answers under input perturbations and refrains from hallucinating when context is missing or…","keywords":["FailSafeQA","long-context question answering","financial NLP benchmark","LLM robustness","hallucination prevention","context grounding","LLM-as-a-judge","10-K filings"],"falsifier":"Select a random subset of about 50 FailSafeQA items covering all six perturbation types, have two human raters independently apply the paper's 1-6 criteria, and compare their labels to the Qwen2.5-72B-Instruct judge labels, measuring agreement with Cohen's kappa. If agreement falls below 0.6, or if the judge gives Qwen-family answers systematically higher relevance than humans do, every model score and the reported robustness-grounding trade-off would need to be recomputed.","tokens_in":20462,"feed_emoji":"📉","tokens_out":7906,"duration_ms":62108,"temperature":0.7,"pith_summary":"This paper introduces FailSafeQA, a benchmark built from 220 long-context financial question-answer tasks derived from 10-K filings, and uses it to score 24 off-the-shelf LLMs under six realistic input failures: misspelled, incomplete, and out-of-domain queries, alongside missing, OCR-degraded, and irrelevant documents. Its central claim is that dependable deployment requires two behaviours that no tested model exhibits together: answering correctly under messy input (Robustness) and refusing to answer when the context cannot support an answer (Context Grounding). The paper reports a sharp trade-off: the most robust model, o3-mini, fabricates information in 41% of unanswerable cases, while the most grounded model, Palmyra-Fin-128k-Instruct, drops robust predictions in 17% of test cases. A new Compliance score formalizes this trade-off, and the highest score any model achieves is 0.81. If the benchmark is right, single-number accuracy evaluations hide the failure mode that matters most for financial use.","feed_headline":"Most robust finance model still fabricates 41% of unanswerable cases","feed_subtitle":"New 220-item benchmark scores 24 LLMs on messy queries and missing files; none balances accuracy and refusal.","key_machinery":"The central object is the FailSafeQA benchmark itself: 220 examples from 10-K filings, each paired with an original query, three perturbed query variants (misspelled, incomplete, out-of-domain), an OCR-corrupted context, and missing- and irrelevant-context conditions. The mechanism that carries the argument is the pair of scores. Robustness is defined as $R = \\frac{1}{n}\\sum_i \\min_j c_{\\ge4}(\\text{model}(T_j(x_i)), y_i)$, and Context Grounding as $G = \\frac{1}{2n}\\sum_{j=1}^2\\sum_{i=1}^n c_{\\ge4}(\\text{model}(T_j(x_i), Y))$, where $c_{\\ge4}$ maps a 1–6 relevance rating into a binary compliant-versus-fabricated label (rating at least 4 means compliant). The two are combined into a Compliance score $\\mathrm{LLMC}_{\\beta} = \\frac{(1+\\beta^2)RG}{\\beta^2 G + R}$ with $\\beta = 0.5$, which prioritizes refusal over answering. The design makes judging tractable by giving the judge LLM short supporting citations instead of full 25k-token contexts, trading long-context degradation for a simpler verification task.","core_discovery":"FailSafeQA claims that dependable deployment of LLMs in finance requires two separable behaviours—answering correctly under messy user input (Robustness) and refusing to answer when the available context cannot support an answer (Context Grounding)—and that current models cannot do both. Across 24 models, every one loses accuracy under perturbation, with the largest drops from OCR-degraded context and out-of-domain phrasing. More striking, the best-answering models fail the refusal test: o3-mini, the most robust model at 0.90, fabricates answers in 41% of unanswerable cases, and reasoning-focused models fabricate in 41 to 70% of those cases. Palmyra-Fin-128k-Instruct has the best Context Grounding at 0.80 but drops robust predictions in 17% of test cases. The paper's compliance metric formalizes this robustness-versus-grounding trade-off and ranks no tested model above 0.81.","pith_inferences":["If the robustness-grounding trade-off holds beyond these 24 models, interventions that improve refusal, such as a separate abstention classifier, may be more practical than trying to make one LLM excel at both behaviours.","The judge premise is testable independently: a human-agreement audit on a 50-item sample, including answers from Qwen2.5-72B-Instruct itself, would show whether any model is systematically favored.","The same six perturbations could be ported to legal, medical, or technical documentation with minimal change, since the only domain-specific step is generating out-of-domain query paraphrases.","One could also use FailSafeQA-style empty-context items to measure whether retrieval-augmented systems that explicitly signal 'no documents retrieved' reduce hallucination rates below the raw model scores reported here."],"forward_implications":["Robustness and Context Grounding should be reported as separate axes in any financial QA evaluation; averaging them together hides the 41% hallucination rate of the most robust model.","Refusal behavior is a trainable capability: the Compliance score gives a concrete objective, and the 0.81 ceiling shows measurable headroom for models that learn when to say the context is insufficient.","OCR-degraded and out-of-domain inputs are the hardest perturbations, so systems deployed on scanned contracts should detect document quality early rather than rely on the LLM to recover.","Text-generation queries such as 'write a blog post' are more prone to hallucination than question-answering queries, so generation pipelines should retrieve grounded facts first and assemble text later."],"supporting_citations":[{"why":"Supplies the HELM definition of robustness as invariance under semantics-preserving transformations, which the paper adapts into its Robustness score.","marker":"Liang et al. (2023)"},{"why":"Provides the LLM-as-a-Judge methodology and impartial-judge prompting used to rate every answer.","marker":"Zheng et al. (2023)"},{"why":"Documents performance degradation on long contexts, the basis for citation-based judging instead of full-context judging.","marker":"Hsieh et al. (2024)"},{"why":"LongCite extracts the supporting citations that make judging tractable and constitute the ground-truth reference for each query.","marker":"Zhang et al. (2024)"},{"why":"FinTruthQA supplies the 1-6 relevance labeling criteria that define compliant versus fabricated answers.","marker":"Xu et al. (2024b)"},{"why":"Llama 3.1 405B generates and postprocesses the synthetic query-answer pairs and perturbations that form the dataset.","marker":"Dubey et al. (2024)"},{"why":"The Qwen2.5 technical report documents both the judge model and the Qwen2.5 family included in the 24-model evaluation.","marker":"Yang et al. (2024)"}],"fun_headline_variants":["Finance LLMs face trade-off: answer or admit ignorance","Even best finance LLMs hallucinate 41% of unanswerable queries","New benchmark finds finance LLMs can't both answer and refuse","FailSafeQA: finance LLMs fail when context is messy or missing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking assumes the judge model Qwen2.5-72B-Instruct rates answers as accurately as a human rater for all 24 models, an assumption the paper does not test with a human-agreement study.","fun_headline_variants_meta":{"raw":{"variants":["Finance LLMs face trade-off: answer or admit ignorance","Even best finance LLMs hallucinate 41% of unanswerable queries","New benchmark finds finance LLMs can't both answer and refuse","FailSafeQA: finance LLMs fail when context is messy or missing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000587,"raw_usage":{"total_tokens":2791,"prompt_tokens":1014,"completion_tokens":1777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":1702}},"tokens_in":630,"tokens_out":1777,"duration_ms":12556,"temperature":1.0,"reasoning_tokens":1702,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:47:43.341045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random subset of about 50 FailSafeQA items covering all six perturbation types, have two human raters independently apply the paper's 1-6 criteria, and compare their labels to the Qwen2.5-72B-Instruct judge labels, measuring agreement with Cohen's kappa. If agreement falls below 0.6, or if the judge gives Qwen-family answers systematically higher relevance than humans do, every model score and the reported robustness-grounding trade-off would need to be recomputed.","supporting_citations":[],"review_version":1}