{"id":"8b83e519-62a7-4b11-b4d9-cbb8b368f37c","arxiv_id":"2512.09506","paper_version":6,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"CNFinBench finds LLMs lose about 15 points from single modules to full agentic execution chains, and their financial-compliance violations surge roughly 160-170% by the second round of multi-turn adversarial attacks.","lead":"CNFinBench is a 29-task finance benchmark that tests large language models on expertise, agentic execution chains, and compliance under multi-turn adversarial attacks, scored with a new continuous metric (HICS). Across 22 models it reports a 15-point drop from single modules to full workflows, and compliance violations rising roughly 160-170% by the second attack round.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HICS is never given a quantitative definition; without its base scores, severity multipliers, and consistency weights, the headline violation-surge numbers and collapse-rhythm typology are not independently checkable.","rationale":"The reader's weakest_assumption identified exactly the same load-bearing point: HICS weights are never specified, and the human agreement study does not validate them. My reading confirms this is the most critical gap because HICS is the paper's central novel metric and the source of the headline quantitative claims. The Table 2 Fin_Basics copy-paste error is also real, but the appendix contains corrected values, so the Expertise-direction claim can be salvaged by relying on Appendix F. The abstract-vs-intro surge discrepancy (159.05% vs 172.3%) is more serious but is itself a symptom of the underlying HICS/reporting ambiguity. Since the reader already issued CONDITIONAL with the condition of releasing a HICS specification, my recommendation is UNCHANGED: the paper remains conditionally acceptable, with the same condition made more precise. The concrete test directly checks whether the unspecified HICS parameters materially affect the headline results.","tokens_in":60386,"tokens_out":2887,"duration_ms":32867,"concrete_test":"Request the official HICS specification (explicit formulas or machine-readable JSON defining base scores, severity multipliers, cross-turn consistency penalties, and violation-trigger rules) plus the raw judge logs for the three multi-turn tasks. Independently reimplement HICS from that specification and recompute the Round-1→Round-2 increase in average violations for MT_App, MT_Inter, and MT_Cog across at least five models. If the recomputed surge for MT_App is not within, say, 10% of 354.3%, or if the overall 'average violations' increase matches neither 159.05% nor 172.3%, the central Integrity claim is not reproducible. Also check whether the two reported numbers can be reconciled by clarifying whether one refers to HICS deductions and the other to raw violation counts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central Integrity claim—that multi-turn adversarial attacks cause rapid, category-specific compliance collapse, e.g., 'average violations surging by 159.05% in Round 2' (abstract) or '172.3%' (Section 1)—depends entirely on HICS, yet HICS is described only qualitatively in §5.2 and §C.3. The manuscript never specifies the Base Score values, Severity Multiplier schedule, or cross-turn consistency tracking formula. The human agreement study (Appendix D, κ=0.74) validates LLM-as-judge rubrics for open-ended tasks, not the HICS penalty parametrization. Consequently, the 354.3% MT_App Round-2 figure, the 55.4%/32.0% MT_Inter/MT_Cog surges, the Defense Degradation Curve, and the Jaccard-similarity collapse-rhythm analysis are all downstream of an un-auditable scoring function. The discrepancy between 159.05% and 172.3% for the same reported quantity reinforces that the numbers cannot currently be verified. This is not an internal logical contradiction in the benchmark's construction, but it means the paper's flagship contribution—HICS as a quantitative safety metric—lacks the specification needed for independent recomputation. The qualitative direction of findings may survive, but the precise headline magnitudes are not reproducible as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"CNFinBench is a Chinese financial LLM benchmark with 29 subtasks organized into Expertise, Autonomy, and Integrity. It draws on regulatory corpora, financial reports, anonymized loan and fraud data, and simulated API/dialogue workflows; 70% of QA items are LLM-drafted and then filtered by three models plus expert review. The paper evaluates 22 open-/closed-source and finance-tuned models using task-specific rubrics and an LLM-as-judge panel (Cohen's κ = 0.74 against experts). It introduces HICS, a 100-point multi-turn compliance metric, and reports three headline findings: (1) models perform well on applied tasks but poorly on rule understanding; (2) a 15.4-point drop from single modules to full execution chains; and (3) rapid compliance collapse under multi-turn attacks, with average violations surging by 159.05% (abstract) or 172.3% (full abstract) in Round 2 and MT_App collapsing fastest (+354.3% in Round 2).","tokens_in":60604,"tokens_out":7587,"duration_ms":73168,"significance":"If the results hold, CNFinBench would be a valuable community resource: it jointly covers agentic execution, multi-turn adversarial red teaming, and a continuous safety metric, which few prior financial benchmarks do. The strengths are real: expert involvement in task design, use of first-hand financial data, a public evaluation platform, detailed rubrics in Appendix E, and a human-agreement study for the LLM-judge setup. However, the current manuscript does not yet support the flagship Integrity and Autonomy claims as stated. HICS is described only qualitatively, so the headline violation-surge numbers are not independently checkable; Table 2 contains an apparent data duplication; and the 15.4-point drop is not derivable from the reported task scores. These are fixable in revision, but they are load-bearing for the paper's conclusions.","major_comments":[{"comment":"HICS is the basis for the paper's central Integrity findings, but it is never quantitatively defined. §5.2 says HICS 'operates on a 100-point scale' with 'rule triggers, severity-adjusted deduction multipliers, and cross-turn consistency tracking,' and §C.3 introduces Base Score, Severity Multiplier, and cross-turn consistency without any equation, weight table, or worked example. Consequently, the Round-2 surge figures (159.05% in the first abstract, 172.3% in the full-text abstract; MT_App +354.3%, §5.3.3) and the Jaccard-similarity collapse-rhythm analysis cannot be recomputed independently. The Appendix D κ=0.74 study validates LLM rubric judgments against experts; it does not validate the HICS penalty parametrization. Please provide the full HICS formula, base-score schedule, multiplier values, consistency term, and reconcile the 159.05/172.3 discrepancy.","section":"§5.2, §C.3, Abstract/§1"},{"comment":"In Table 2, the Fin_Basics (BK) column is identical to the Fin_Report_Parse (RP) column for all 22 models (e.g., Doubao 76.7, GPT-4o 83.1, Qwen3-14B 77.0), while Appendix F Table 10 reports different Fin_Basics values for the same models (e.g., GPT-4o 35.0±0.6, Qwen3-14B 3.9±0.3). The §5.3.1 claim that models score only 38.50 on Fin_Basics and the Expertise-vs-Autonomy contrast depend on this column. The duplication appears to be a data error; it must be corrected and the affected analyses rerun before the Expertise conclusions can be evaluated.","section":"Table 2, §5.3.1"},{"comment":"The headline Autonomy result, 'a 15.4-point drop from single-step to multi-step tasks,' is not supported by the numbers in the same paragraph. It cites Path_Plan 63.3, Ret_API 53.9, and Multi_App 56.65; the pairwise differences are 9.4 and 6.65, not 15.4. Please define explicitly which tasks are aggregated as 'single modules' and which as 'execution chains,' report the per-model computation, and show how 15.4 is obtained. Without this, the central module-to-chain degradation claim is unverifiable.","section":"§5.3.1"},{"comment":"The difficulty filter discards any item that two of Qwen3-235B, DeepSeek-V3, and GPT-4o answer correctly, and the same model families are also used as LLM judges and are among the 22 evaluated models. This is an ad hoc selection rule that can systematically disadvantage those families and their derivatives, so cross-model rankings may partly reflect benchmark construction rather than capability. I am not claiming fitted-parameter circularity, but this is a validity risk. Please report results separately for the 30% expert-authored subset, or provide a stability analysis (e.g., rank correlation with and without the filtered items).","section":"§4.3"}],"minor_comments":[{"comment":"The abbreviation key lists 'MI = MT_Iner'; this should be 'MT_Inter'.","section":"Table 1"},{"comment":"Minor naming inconsistencies: 'Itnet_ID' should be 'Intent_ID', and Appendix F uses 'Qwen3-235B-A22B' while the main text uses 'Qwen3-235B'.","section":"Table 2 / Appendix F"},{"comment":"All models are evaluated with a 2,048-token context window, yet the benchmark includes Long_QA and Long_Conv tasks over long financial documents. Please clarify how documents are chunked or selected; otherwise these tasks measure truncated-context behavior rather than long-context reasoning.","section":"§5.1"},{"comment":"The statement that online leaderboard scores may differ from the paper because the platform uses 'a curated challenging subset' makes the reported numbers hard to compare with the live system. Please archive and version the exact evaluation set used for the paper's tables.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and relevant topic, and the benchmark has obvious community value. My recommendation is driven by three fixable but load-bearing issues: the missing HICS specification, the duplicated Fin_Basics column, and the unexplained 15.4-point drop. The 159.05/172.3 abstract discrepancy and the Table 2 duplication should be checked carefully before acceptance; they may indicate a broader data-handling problem. I would support acceptance after a revision that makes HICS fully specified, corrects the tables, and reconciles the headline numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: CNFinBench is a real artifact—29 tasks, 11,947 single-turn items, 321 four-turn adversarial dialogues, an agentic execution pipeline, and a continuous safety metric—and the qualitative findings (weak regulatory precision, ~15-point module-to-chain drop, rapid multi-turn compliance collapse) hold up in the appendix tables. But as printed it has three problems that block acceptance.\n\nFirst, Table 2's Fin_Basics column is byte-for-byte the same as the Fin_Report_Parse column across all 22 models. The text's reported median for Fin_Basics (38.50) matches Appendix F, not Table 2, so this looks like a rendering error rather than a fabricated result, but it erodes trust in the main table and must be fixed. Second, the abstract says average violations surge 159.05% in Round 2; Section 1 says 172.3%. Both describe the same quantity, so at least one is wrong. That's a sloppy inconsistency in a headline number. Third, and most importantly, HICS—the paper's flagship contribution—is never given a quantitative definition. The base scores, severity multipliers, and cross-turn consistency weights are described qualitatively in §5.2 and §C.3, with no equations or tables. The human agreement study (κ=0.74) validates the LLM-as-judge rubrics for open-ended tasks, not HICS's penalty parametrization. So the 354.3% MT_App Round-2 figure, the collapse-rhythm Jaccard similarities, and the Defense Degradation Curve are not independently checkable.\n\nWhat's good: the construction is careful. Expert-panel curation, dual-path QA generation, difficulty filtering, standard deviations in the appendix, and a genuine human-agreement check put this above many benchmark papers. The attack personas and strategies are detailed and plausible. The qualitative conclusions—weak regulatory understanding despite applied competence, multi-step autonomy degradation, faster application-layer collapse—survive scrutiny of the appendix data. The paper is also honest about the leaderboard using a different subset and about 70% of gold labels being LLM-drafted with expert review.\n\nThese flaws are localized but load-bearing for the quantitative claims. The qualitative direction may be right, but the precise magnitudes should not be cited as printed. I'd send this to peer review with a request for a corrected Table 2, reconciled headline numbers, and a released HICS specification (equations or deduction tables). With those, the benchmark becomes a usable measurement instrument. Without them, it's a promising template whose central metric is un-auditable.\n\nThis paper is for anyone building or deploying financial LLM agents, and for benchmark researchers working on compliance evaluation. It deserves a serious referee, and the revision path is clear.","headline":"CNFinBench is a genuine new benchmark with careful construction and plausible qualitative findings, but three concrete reporting/auditability problems—a copy-paste column, inconsistent headline numbers, and an unspecified HICS metric—block acceptance until fixed.","tokens_in":61316,"tokens_out":2287,"would_cite":true,"duration_ms":24550,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Financial LLM agents that ace applied tasks still lack rule understanding, lose 15 points when chained into workflows, and break under two-round adversarial attacks.","keywords":["financial LLM agents","agentic autonomy","regulatory compliance","multi-turn adversarial attack","HICS","LLM-as-judge","compliance drift","benchmark"],"falsifier":"Re-run the MT_App dialogues and have three independent compliance experts score each response with a fixed rubric on a 1–5 risk scale, blinded to HICS. If expert-rated risk does not increase from Round 1 to Round 2 (or increases far less than the reported 354.3%), the escalation is an artifact of the unpublished HICS weights rather than a property of model behavior.","tokens_in":60097,"feed_emoji":"⚠️","tokens_out":4749,"duration_ms":48211,"temperature":0.7,"pith_summary":"This paper argues that evaluating financial large language models on textbook-style question answering hides the risks that appear when they act as autonomous agents. To close that gap, it builds CNFinBench, a benchmark of 29 subtasks spanning expertise, autonomy, and integrity. Running 22 models through it, the paper establishes three failure modes: models handle applied loan and credit tasks well but score poorly on regulatory exam-style items; performance drops 15.4 points when tasks are chained into end-to-end execution paths; and in multi-turn adversarial dialogues compliance violations surge by roughly 172% by the second round, with impersonation attacks degrading fastest (+354%). The central contribution is a new metric, HICS, which claims to quantify violation escalation and behavioral consistency that binary refusal rates miss.","feed_headline":"Finance LLMs ace tasks, flunk the rules","feed_subtitle":"Benchmark of 22 models shows chained workflows and two-round attacks expose compliance collapse.","key_machinery":"CNFinBench's Expertise–Autonomy–Integrity triad of 29 subtasks, and the Harmful Instruction Compliance Score (HICS), a 100-point multi-turn safety metric. HICS works by decomposing each adversarial response into atomic and sequential violations, assigning a base score by risk type, scaling by a severity multiplier, and tracking consistency across dialogue turns. Its role is to replace binary refusal rates with a continuous signal of compliance erosion; the paper credits it with revealing the divergent collapse rhythms across the three attack families.","core_discovery":"On the paper's own terms, the discovery is that the three capabilities that matter for safe financial deployment — expertise, autonomy, and integrity — are separable and measurably out of balance in current LLMs. Closed and open models both show an 'illusion of regulatory competence': they choose correct answers on applied credit and loan tasks while failing Fin_Basics and Fin_Cert_Exams that require exact rule interpretation. In autonomy, models lose 15.4 points going from isolated steps to full execution chains, with the largest drops in parameter precision and inter-agent coordination rather than API selection. In integrity, single-turn defenses look strong (several models score above 98","pith_inferences":["Editorial inference: if the unspecified HICS penalty weights are arbitrary, the headline escalation percentages (172%, 354%) should be read as ordinal signals — 'violations clearly grow' — rather than as calibrated measures of real-world risk magnitude.","Editorial inference: the Round-2 collapse in application-layer attacks suggests that identity impersonation exploits a distinct mechanism — the model's willingness to trust a claimed insider role — which may be addressable by explicit authorization checks in the agent scaffold rather than by further safety fine-tuning.","Editorial inference: the same Expertise–Autonomy–Integrity measurement design is portable to other high-privilege agent domains (health records, legal document handling, cloud administration), where the same pattern of good single-turn knowledge but weak chained compliance is plausible.","Editorial inference: a direct test of whether HICS adds information over refusal rates would be to re-score the same MT_App dialogues with binary 'compliant/not' labels; if the binary signal already peaks in Round 2, HICS's contribution is diagnostic granularity rather than new detection."],"forward_implications":["Models that look safe in single-turn tests (refusal rates above 98%) can nonetheless comply with harmful requests by the second round of a dialogue, so deployment-time safety checks need multi-turn evaluation.","High performance on applied financial tasks is not evidence of rule understanding; certification-style regulatory items expose a capability gap untouched by loan and credit benchmark accuracy.","Autonomy failures are procedural, not conceptual: models pick the right APIs and roles, then fail at parameter precision, unit handling, and inter-agent data flow.","The three attack families collapse on different timelines (application early, internal gradual, cognitive late), implying that defenses need to be tailored to trust logic rather than uniformly hardened.","Because HICS tracks specific violation rule types, the benchmark offers traceable, per-turn deduction logs that can guide targeted model refinement."],"fun_headline_variants":["Finance LLMs: rule-smart on paper, rule-blind in practice","AI finance agents fail compliance when workflows get real","Chained tasks and attack rounds expose LLM compliance collapse","Model scores plummet 15.4 points in full agent workflows","Two-round attacks nearly triple compliance violations in finance LLMs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole integrity measurement rests on the assumption that HICS's base scores, severity multipliers, and cross-turn deductions — whose exact values are never stated — reflect the real compliance risk of a financial response.","fun_headline_variants_meta":{"raw":{"variants":["Finance LLMs: rule-smart on paper, rule-blind in practice","AI finance agents fail compliance when workflows get real","Chained tasks and attack rounds expose LLM compliance collapse","Model scores plummet 15.4 points in full agent workflows","Two-round attacks nearly triple compliance violations in finance LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000317,"raw_usage":{"total_tokens":1645,"prompt_tokens":778,"completion_tokens":867,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":798}},"tokens_in":522,"tokens_out":867,"duration_ms":9183,"temperature":1.0,"reasoning_tokens":798,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T17:25:50.442446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the MT_App dialogues and have three independent compliance experts score each response with a fixed rubric on a 1–5 risk scale, blinded to HICS. If expert-rated risk does not increase from Round 1 to Round 2 (or increases far less than the reported 354.3%), the escalation is an artifact of the unpublished HICS weights rather than a property of model behavior.","supporting_citations":[],"review_version":1}