REVIEW 3 major objections 9 minor 28 references
Frontier LLMs ace actuarial exams on knowledge but still lag humans on real case work and spreadsheet/R workflows.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 19:26 UTC pith:3JIVIHMO
load-bearing objection Solid, citable actuarial LLM benchmark with a real Know→Case→Practice gap; the absolute model pattern is robust, the “much weaker than humans” framing rests on a thin baseline. the 3 major comments →
INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors claim a clear capability boundary: frontier LLMs perform strongly on standardized actuarial knowledge, but remain much weaker than actuarially trained humans on long-context case reasoning, tool-based spreadsheet and R workflows, and jurisdiction-sensitive practice. In their main results, leading models reach roughly the low-to-mid 90s on knowledge while human experts sit near the mid-80s, yet humans lead by a wide margin on cases and practice. That gap, not a single leaderboard score, is the central finding they want the field to act on.
What carries the argument
INS-ActBench’s three-level progressive structure: INS-Act-Know (standardized knowledge across six subfields), INS-Act-Case (long-context multi-select insurance cases, original contexts averaging tens of thousands of tokens), and INS-Act-Practice (spreadsheet and R tasks scored by executing model-produced code against numerical ground truth with 0.2% relative tolerance). The machinery is the joint, auditable pipeline from credentialing materials to comparable scores across knowledge, context, and executable tools.
Load-bearing premise
The load-bearing premise is that turning open exam and case material into fixed multiple-choice and numerical spreadsheet/R tasks still measures real professional actuarial work, not just exam-format skill under automated scoring.
What would settle it
Re-run the same models and human panel on a held-out set of unaltered open-ended actuarial case write-ups and live spreadsheet/R engagements scored by independent actuaries; if the knowledge–case–practice gap disappears or reverses under that scoring, the claimed professional-capability boundary would not hold.
If this is right
- Actuarial LLM development should prioritize long-context evidence integration and correct tool-mediated calculation over further gains on standard knowledge MCQs.
- Benchmarking and training need multi-jurisdiction coverage, because performance varies across SOA, IFoA, and other association sources more once tasks leave pure knowledge.
- Practice failures that are mostly modeling/calculation rather than pure tool-call failures imply models need actuarial procedure skill, not only better tool invocation.
- INS-ActBench becomes a shared, reproducible target for assistants meant for pricing, reserving, and related insurance analysis with verifiable outputs.
Where Pith is reading between the lines
- The same knowledge-vs-workflow split likely appears in neighboring regulated professions (tax, audit, clinical protocols) that mix closed-book facts with long documents and executable tools.
- Partial-credit gains on multi-select cases suggest under-selection of evidence is a dominant error mode; retrieval-plus-checklist agents may close more of the case gap than pure scale.
- If firms adopt this style of eval, hiring and model procurement may shift from “passes actuarial MCQs” toward measured case and spreadsheet/R reliability under jurisdiction tags.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces INS-ActBench, a 12,050-item actuarial benchmark built from public examination materials of 16 actuarial associations, organized into three progressive subsets: INS-Act-Know (single-answer MCQ knowledge), INS-Act-Case (multi-select questions over long case contexts averaging 58K tokens), and INS-Act-Practice (spreadsheet and R-code tasks scored by executing generated code against reference numerical outputs with 0.2% relative tolerance). Nine LLMs plus a five-person actuarially trained human baseline are evaluated. The central claim is a capability boundary: frontier models reach ~93–95% on Know (at or above the human baseline of 84.4%) but drop to ~44–56% on Case and ~58–69% on Practice, versus human scores of ~80% and ~85%. Supporting analyses cover numerical vs. non-numerical items, an error taxonomy for Practice, cross-jurisdiction variation, partial-credit re-scoring, and contamination diagnostics. The benchmark construction pipeline, expert verification (9.65% of items revised), contamination checks, and released code are genuine strengths. My main concerns center on the human baseline protocol, which under-specifies the conditions that make the headline human–LLM comparison interpretable.
Significance. If the human comparison is repaired, this is a substantial contribution. It is, to my knowledge, the first large-scale actuarial benchmark integrating knowledge, long-context case reasoning, and executable tool use, and the construction effort is real: 16 associations, multi-stage expert verification with a named credential (SOA China Committee member / ASA), a documented selection funnel, and per-item context-length statistics. The Practice subset with execution-based scoring against ground-truth workbooks and Docker-run R code is a genuine advance over QA-style finance benchmarks, and the three-way contamination diagnostics (exposure, option reconstruction, pre/post-cutoff) plus the partial-credit re-scoring analysis are more rigorous than most benchmark papers provide. The capability-boundary finding (strong Know, weak Case/Practice) is plausible, internally supported, and useful to the community even if the human comparison only bounds its magnitude.
major comments (3)
- [§4.1, Human Baseline; Table 3] The human baseline is under-specified in ways that are load-bearing for the central 'much weaker than actuarially trained humans' claim (Table 3). INS-Act-Case items average 58K tokens (max 175K), yet five participants completed 300 items in ~20 hours, i.e., ~4 minutes per item, which is incompatible with actually reading the long contexts. The paper states (§4.1) that Qwen3-14B and Gemma-3-12B-IT use simplified contexts 'due to context-length limits', but never states which context version the humans received, and logs no per-item timing. If humans used the ≤32K simplified contexts (or skimmed), the human CA score of 79.6 is not comparable to LLM scores on the long context, and the magnitude of the headline Case gap (~56 vs ~80) could change substantially. Please state the context version, report per-item timing, and if necessary re-run or re-analyze the human Case subset under matched
- [§4.1, Tables 3 and 8] Two further comparability issues with the human study. (a) No uncertainty quantification: with n=100 items per subset, the standard error on a human accuracy near 84% is ~4 points, so the Know comparison (LLM 93–95 vs human 84.4) is only marginally significant; the Case/Practice gaps are large enough to survive, but this should be shown with bootstrap CIs rather than asserted. Table 8 shows one participant (Expert 4) scoring 70.0 on Know total vs 92.0 for Expert 1, so between-participant variance is non-trivial. (b) The sampling of the 300 human items is not described (random? stratified by source association and subcategory?), and the instrument differs across arms: humans work natively in Excel/R while LLMs must generate Python/openpyxl code to fill answer cells — part of the Practice gap may measure code-generation friction rather than actuarial workflow competence. Please describe th
- [§4.5, Table 5] The cross-jurisdiction claim ('actuarial reasoning by LLMs remains sensitive to regulatory context', also asserted in the abstract) is confounded with item difficulty and language. The 'Others' category pools 13 heterogeneous associations (CAA, IAJ, DAV, IBA, ...), and models consistently score *higher* on Others-Case than on SOA/IFoA-Case (e.g., GPT-5.5: 74.0 Others vs 55.0 SOA/43.4 IFoA; Gemini-3.1-Pro: 80.2 vs 42.3/57.6). A pure 'jurisdiction sensitivity' account does not predict that the residual category is easiest; this pattern is at least as consistent with SOA/IFoA case items being intrinsically harder (longer contexts, multi-statement verification) or with language-mix effects. Please either add a difficulty-controlled comparison (e.g., matched context length, or per-association breakdown for the largest 'Others' sources) or soften the claim to variation across sources.
minor comments (9)
- [§4.1, Inference Settings] Typo: 'models are are evaluated'.
- [§4.2 / Appendix C.2, Table 9] Exact-match all-and-only scoring for multi-select Case items is stringent; Table 9's partial-credit re-scoring (e.g., Qwen3.6-Plus 44.6→62.1) shows much of the Case deficit is under-selection rather than wrong selection. Since this materially qualifies the main result, consider reporting partial-credit scores alongside exact-match in Table 3 rather than only in the appendix.
- [Appendix D.1 and D.3] The 13-gram overlap analysis uses Google top-10 appearance as a proxy for training-data exposure; search-engine indexability is not equivalent to inclusion in pretraining corpora, and the May-2026 'post-cutoff' split assumes knowledge cutoffs the authors cannot verify for proprietary models. Please acknowledge these proxies explicitly as limitations rather than 'no clear evidence of contamination'.
- [Appendix B, context compression statistics] The aggregate token-level compression rate (54.70%) versus mean per-case compression rate (25.41%) discrepancy is confusing; please clarify the definitions or reconcile them.
- [Appendix B, Table 7] The selection funnel percentages (40% + 35% + 5% + 5% + 5% removed, 10% retained) appear to be relative to the initial collection; please state whether the stages are sequential and mutually exclusive, and why 'more than five years old' alone removes 40%.
- [Appendix C.1] Human compensation is stated as 'a rate equivalent to one day of their regular pay' for ~20 hours of work; please clarify, as incentive adequacy bears on effort on the long-context items. Also state whether the 300 human items overlap with the contamination-diagnostic samples.
- [§4.1, Evaluation item (3)] For Spreadsheet tasks, 'If multiple answer cells are required, we take the average accuracy of each cell' — please report how many tasks have multiple answer cells and whether averaging (vs all-cells-correct) materially changes scores.
- [Figure 1; Table 10] Figure 1 ('Evaluation dataset') adds little information; consider replacing it with the progressive capability structure or removing it. Table 10 (execution times) could be trimmed to a sentence since efficiency is explicitly not a scoring criterion.
- [§4.1, Inference Settings] INS-Act-Practice uses zero-shot while Know/Case use two-shot; please justify the asymmetry, since few-shot examples could plausibly help tool-use formatting and narrow the Practice gap.
Circularity Check
No significant circularity: empirical benchmark evaluation, not a fitted-input-as-prediction or self-definitional derivation chain.
full rationale
INS-ActBench is a construction-and-evaluation paper. Its load-bearing claims are empirical accuracies of nine LLMs and a five-person human panel on exam-derived items (Tables 3–5, 8–12), plus contamination diagnostics (Appendix D). There is no derivation in which a quantity labeled “prediction” or “first-principles result” is algebraically or statistically identical to a fitted input: Know/Case/Practice scores are direct task accuracies under stated scoring rules (exact-match MCQ; 0.2% relative numerical tolerance), not parameters refit and re-reported. Self-citations (e.g., INS-MMBench, INSEva) appear only as related-work context for insurance benchmarks and do not supply uniqueness theorems or forced ansätze that underwrite the main gap claim. Contamination checks (13-gram exposure, option reconstruction, pre/post May 2026 split) treat residual public-exam exposure as a limited risk rather than as the mechanism producing the Know–Case–Practice boundary. Construct-validity choices (subjective→MCQ conversion, numerical-only Practice) affect what is measured but are not circular reductions of outputs to inputs. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (3)
- numerical_relative_error_tolerance =
0.2%
- subset_size_and_mix_targets =
12050 items; ~8:1:1 mix
- human_baseline_sample_size =
300 questions; 5 humans
axioms (4)
- domain assumption Public actuarial association exam/sample items, after MCQ and numerical-tool standardization, are a valid proxy for professional actuarial capability along knowledge, case reasoning, and tool execution.
- ad hoc to paper Exact-match multi-select scoring (all-and-only correct options) is an appropriate primary metric for case competence.
- domain assumption SOA/CAS, IFoA, and 13 other associations sufficiently represent global jurisdictional diversity for cross-jurisdiction claims.
- domain assumption Automated execution of model-written Python(openpyxl) and R code with printed numerical outputs fairly grades tool-based actuarial practice.
invented entities (1)
-
INS-ActBench (Know/Case/Practice progressive suite)
independent evidence
read the original abstract
Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.
Figures
Reference graph
Works this paper leans on
-
[1]
TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , pages=
-
[2]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data , author=. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[3]
Advances in Neural Information Processing Systems , volume=
Pixiu: A comprehensive benchmark, instruction dataset and large language model for finance , author=. Advances in Neural Information Processing Systems , volume=
-
[4]
Advances in Neural Information Processing Systems , volume=
Finben: A holistic financial benchmark for large language models , author=. Advances in Neural Information Processing Systems , volume=
-
[5]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Investorbench: A benchmark for financial decision-making tasks with llm-based agent , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[6]
arXiv preprint arXiv:2501.10943 , year=
InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models , author=. arXiv preprint arXiv:2501.10943 , year=
-
[7]
arXiv preprint arXiv:2509.04455 , year=
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance , author=. arXiv preprint arXiv:2509.04455 , year=
-
[8]
arXiv preprint arXiv:2511.07794 , year=
Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark , author=. arXiv preprint arXiv:2511.07794 , year=
-
[9]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[10]
British Actuarial Journal , volume=
ActuaryGPT: Applications of large language models to insurance and actuarial work , author=. British Actuarial Journal , volume=. 2024 , publisher=
2024
-
[11]
Cfinbench: A comprehensive chinese financial benchmark for large language models , author=. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2025
-
[12]
Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V
Can LLM-based financial investing strategies outperform the market in long run? , author=. Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1 , pages=
-
[13]
Proceedings of the 33rd ACM International Conference on Multimedia , pages=
Mme-finance: A multimodal finance benchmark for expert-level understanding and reasoning , author=. Proceedings of the 33rd ACM International Conference on Multimedia , pages=
-
[14]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Financereasoning: Benchmarking financial numerical reasoning more credible, comprehensive and challenging , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[15]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
FinMathBench: A Formula-Driven Benchmark for Evaluating LLMs’ Math Reasoning Capabilities in Finance , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[16]
2017 , publisher=
Computational Actuarial Science With R , author=. 2017 , publisher=
2017
-
[17]
arXiv preprint arXiv:2311.11944 , year=
Financebench: A new benchmark for financial question answering , author=. arXiv preprint arXiv:2311.11944 , year=
-
[18]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=
Docfinqa: A long-context financial reasoning dataset , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=
-
[19]
arXiv preprint arXiv:2603.07316 , year=
FinSheet-Bench: From Simple Lookups to Complex Reasoning, Where LLMs Break on Financial Spreadsheets , author=. arXiv preprint arXiv:2603.07316 , year=
-
[20]
arXiv preprint arXiv:2512.13168 , year=
Finch: Benchmarking finance & accounting across spreadsheet-centric enterprise workflows , author=. arXiv preprint arXiv:2512.13168 , year=
-
[21]
Scientific Data , year=
Statllm: A dataset for evaluating the performance of large language models in statistical analysis , author=. Scientific Data , year=
-
[22]
Proceedings of the 29th symposium on operating systems principles , pages=
Efficient memory management for large language model serving with pagedattention , author=. Proceedings of the 29th symposium on operating systems principles , pages=
-
[23]
arXiv preprint arXiv:2601.23048 , year=
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics , author=. arXiv preprint arXiv:2601.23048 , year=
-
[24]
NYU Stern School of Business , year=
On the financial regulation of insurance companies , author=. NYU Stern School of Business , year=
-
[25]
Journal of Artificial Societies and Social Simulation , volume=
The insurance industry as a complex social system: Competition, cycles, and crises , author=. Journal of Artificial Societies and Social Simulation , volume=. 2018 , publisher=
2018
-
[26]
British Actuarial Journal , volume=
The importance of actuarial management in insurance business decision-making in the twenty-first century , author=. British Actuarial Journal , volume=. 2021 , publisher=
2021
-
[27]
Actuarial Practice Forum , pages=
Spreadsheet issues: pitfalls, best practices, and practical tips , author=. Actuarial Practice Forum , pages=
-
[28]
Journal of Statistical software , volume=
actuar: An R package for actuarial science , author=. Journal of Statistical software , volume=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.