REVIEW 4 major objections 5 minor 30 references
A new seven-dimension reliability score for capital-markets LLM outputs finds frontier closed-source models clustered within 0.22 points, with the open-weights baseline ranked last by all judges.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:43 UTC pith:WSXXFGXJ
load-bearing objection A genuinely workflow-level reliability rubric for capital-markets LLM outputs, transparently demonstrated, but the empirical ordering claims rest on uncalibrated LLM judges and some internal numbers don't reconcile. the 4 major comments →
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is a measurement result: on the CM-LRS rubric, the three frontier closed-source models are effectively indistinguishable as a group — Sonnet 4.6 at 4.31, Opus 4.7 at 4.30, GPT-5.5 at 4.09 on four-judge averaged aggregates — and every judge places the open-weights baseline (Llama 3.3 70B, 3.15) last, a gap of 0.94–1.16 points. The gap is not uniform: it is 2.23 points on precedent retrieval and 2.15 on issuer-profile synthesis but only 0.84 on single-document debt-terms extraction. Decision Usefulness (D6) shows the largest cross-model dispersion in the study (a 4.0-point spread on the issuer-profile workflow) while sitting in the top tier of inter-judge agreement (mean
What carries the argument
The key machinery is CM-LRS itself: a seven-dimension, 0–5 rubric with anchor descriptions per score level (0 = unusable, 5 = production-grade), aggregated as a weighted sum with workflow-class default weights; scoring is executed by four LLM judges (two in-panel, two out-of-panel, spanning three model families) using the same prompt and rubric, with pairwise Spearman correlations and per-dimension Pearson correlations reported to characterise judge agreement. The rubric is paired with a workflow taxonomy — extraction, retrieval, synthesis, comparison & reasoning, drafting — that sets default weights and directs which dimensions dominate; the paper's working principle, 'predict with the LLM,
Load-bearing premise
The load-bearing premise is that four LLM judges' rubric scores are a valid proxy for expert human review of bankability; the paper explicitly states that this is 'an automated-judging baseline rather than a human-rater study' and that its primary judge (Claude Sonnet 4.6) is itself one of the models being scored — if LLM judges do not track what human reviewers would reject, the numeric scores and the model ordering lose their stated meaning.
What would settle it
Score the same 104 outputs with practising bankers and compliance reviewers using the same 0–5 rubric and record their accept/reject decisions; if human reviewers' model ranking differs from the four-judge LLM ordering — for instance, if they do not place Llama 3.3 70B last, or they find a wide gap among Sonnet 4.6, Opus 4.7, and GPT-5.5 — the claim that CM-LRS measures bankability would be refuted.
If this is right
- For capital-markets buyers, headline reliability on these five workflows does not differentiate the three frontier closed-source models; cost, latency, and workflow-class fit become the deciding factors.
- Open-weights LLMs may suffice for single-document structured extraction (gap 0.84) but are not yet bankable for multi-document retrieval and synthesis, where traceability failures dominate.
- A deployment gate requiring CM-LRS ≥ 4.0, no dimension below 3, and no 0 or 1 scores gives a concrete, tunable proxy for reviewer readiness in regulated settings.
- The gap between CM-LRS and QA-pair scores is itself a risk indicator: surface-correct but untraceable or incomplete outputs will systematically score lower on CM-LRS, which is the intended behaviour.
- Decision Usefulness (D6) is the most reliable dimension-level separator, and combining it with numerical consistency (D3) and workflow completeness (D4) predicts whether the rest of an output is safe to admit.
Where Pith is reading between the lines
- If the near-zero agreement between the two most independent judges (GPT-5.5 vs Gemini, Spearman rho = 0.03) reflects genuine judge-level noise, then the robust conclusions may be limited to the frontier-vs-open-weights gap and to D6's dispersion, not to finer distinctions among the closed-source models.
- A testable extension is to apply CM-LRS to drafting workflows (pitch evidence, compliance memos), which the paper leaves untested; if D6 remains the separator there, the metric generalises beyond the five demonstrated classes.
- Because all documents were truncated to 6,000 tokens, the demonstrated scores bias toward front-matter and cover-page content; a full-document chunked pipeline is the natural stress test of whether the frontier cluster survives at depth.
- If the 'predict with the LLM, calculate with code, decide with the human' principle is adopted architecturally, numerical consistency and reviewability become build-time constraints rather than evaluation-only dimensions, turning CM-LRS from a benchmark into a compliance instrument.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CM-LRS, a seven-dimension 0-5 rubric for scoring LLM workflow outputs in capital-markets settings at the workflow-output layer rather than the QA-pair layer. It defines a workflow taxonomy, an aggregate formula with tunable weights, and demonstrates the framework on five public/synthetic workflows, scoring four models (Claude Opus 4.7, GPT-5.5, Claude Sonnet 4.6, Llama 3.3 70B) with four LLM judges. Headline findings are that frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.31, Opus 4.30, GPT-5.5 4.09), Llama is last at 3.15, the open-weights gap concentrates on retrieval/synthesis rather than extraction, and Decision Usefulness (D6) is the cleanest dimension-level separator. The paper releases rubrics, prompts, outputs, and a deterministic verification script.
Significance. If the ordering claims held, CM-LRS would be a useful deployment-screening instrument. The strengths are real: the workflow-output framing addresses a genuine gap; the corpus is public or synthetic; the four-judge protocol is a serious attempt to control self- and family-bias; and the release of rubrics, prompts, and scoring outputs with a verification script is exemplary for reproducibility. However, the empirical headline currently rests on unvalidated LLM-as-judge scores: there is no human-rater calibration, no statistical significance testing, and at least one judge pair shows effectively independent rankings (GPT-5.5 vs Gemini rho = 0.03). These issues are load-bearing for the paper's central deployment-relevance claim, not presentation details. The framework is promising, but the evidence as presented does not yet establish 'bankability'.
major comments (4)
- [§5.4, §10.1] The load-bearing assumption that LLM-judge scores proxy human 'bankability' decisions is untested. The paper explicitly concedes this is "an automated-judging baseline rather than a human-rater study." The 0.22-point frontier cluster and the Llama-last ordering are deployment-meaningful only if judges track the human review bar. A stratified human pilot on a sample of the 104 outputs (same rubric, banker/compliance raters, inter-rater reliability, agreement with each judge) is needed to anchor the headline. Without it the central ordering is an unvalidated proxy.
- [§5.4, Table 5, bottom line] The abstract and bottom line claim the three frontier models are "statistically indistinguishable," but no significance test, confidence interval, or standard error is reported for the 0.22-point cluster or the 1.16-point gap. Pairwise judge Spearman correlations range from 0.03 (GPT-5.5 vs Gemini) to 0.94 (Sonnet vs Haiku); the four-judge mean is an average of discordant rankings. Report per-judge aggregate tables, judge×model interactions, and a bootstrap or mixed-effects test of the cluster/gap before making the indistinguishability claim.
- [§7, Table 5, Conclusion] The conclusion states Sonnet "wins or ties every workflow," but Table 5 is the primary judge's scoring and Sonnet is that judge. Section 5.4 itself says the in-cluster ranking is judge-dependent. This is a self-judged claim and should be removed or explicitly labelled as the primary judge's self-assessment. The four-judge protocol mitigates self-bias for the overall ordering, but not for this specific sentence about Table 5.
- [§4.3] The claim that the frontier-vs-Llama gap is "robust to any reasonable weight choice" is unsupported. Eq. (1)'s weights and the deployment-readiness threshold are free parameters. A weight sweep over Table 3 defaults and plausible alternatives should be reported to demonstrate that the ordering and cluster membership do not flip; otherwise the claim should be softened to a conjecture.
minor comments (5)
- [Appendix C] The subsections are labeled B.1–B.4 inside Appendix C; renumber them to C.1–C.4.
- [§5.4] "The strongest single empirical signal ... that family-bias matters" overstates a single Spearman correlation; present a formal test of family-bias or soften the wording.
- [§5.2] "Four independent LLM judges" is imprecise since GPT-5.5 is both a judge and an evaluated model; suggest "four LLM judges, two in-panel and two out-of-panel."
- [Table 4] Cost figures would be easier to verify if the list-price retrieval date and source URL are reported.
- [§10.5] The acknowledged 6,000-token truncation is the largest external-validity threat for long documents; consider a small full-document pilot to quantify truncation sensitivity.
Circularity Check
No significant circularity: the central reliability ordering is an empirical scoring result, with the self-judge bias explicitly disclosed, controlled by out-of-panel judges, and not load-bearing for the headline claims.
full rationale
CM-LRS is an operational metric defined as a weighted sum of seven rubric dimensions (Eq. 1), scored by LLM judges. The headline findings—frontier-cluster within 0.22 points and Llama last—are descriptive results from applying that metric, not predictions derived from fitted inputs. No parameter is fitted to the reported scores, no equation reduces to itself, and no uniqueness theorem or ansatz is imported from self-citation. The primary judge (Claude Sonnet 4.6) is one of the four scored models, a potential self-judge bias the paper explicitly acknowledges in Section 5.4 ('The most obvious bias in this setup is that the primary judge (Claude Sonnet 4.6) is itself one of the four models in the panel.') and in Section 10.1. However, the paper controls for this with three additional judges including two fully out-of-panel (Haiku 4.5, Gemini 2.5 Pro) and reports that the substantive findings 'survive all four judges' and that the in-cluster ranking is 'judge-dependent' and 'reported with that caveat throughout.' The central ordering (Llama last, frontier cluster) is thus not forced by the self-judge; it is reproduced by independent judges. The absence of human-rater calibration (Section 10.1: 'this remains an automated-judging baseline rather than a human-rater study') is a validity limitation, not a circular step: the metric is transparently defined as an LLM-judge score, and its construct validity is a separate empirical question. The paper's self-citation (Ahuja, 2026) is only a code/data release link, not evidence in the derivation. No circular reduction can be quoted from the equations or method; the scoring protocol is an open, reproducible procedure rather than a tautology. Therefore the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (2)
- Aggregate dimension weights =
Equal weights in headline; Table 3 workflow-class defaults
- Deployment-readiness threshold =
CM-LRS ≥ 4.0 and no dimension below 3, with 0 or 1 disqualifying
axioms (4)
- domain assumption LLM-judge scores are a valid proxy for expert human reviewer judgments under the rubric
- domain assumption The seven chosen dimensions and 0–5 anchors capture what makes an output bankable
- domain assumption 6,000-token document truncation preserves the evidence needed for scoring long filings
- domain assumption Ordinal 0–5 rubric scores can be averaged and Pearson-correlated as though interval-valued
read the original abstract
In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defensible in front of a counter-party or a regulator, with the documents in hand. Existing methods address parts of that gap: open-domain QA benchmarks reward surface accuracy, and finance benchmarks (FinanceBench, FinQA, ConvFinQA) advance document-grounded and numerical QA but evaluate at the question-answer layer rather than the workflow outputs practitioners defend. We introduce CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability/auditability. Each is scored 0-5 against a rubric anchored on signals reviewers in regulated settings use; the aggregate is tunable to the workflow. We demonstrate CM-LRS on five workflows (DCM transaction-terms extraction, precedent retrieval, issuer profile synthesis, M&A transaction-comparable reasoning, ECM transaction-terms extraction) over public SEC EDGAR filings, a public UK takeover release, and fictional synthetic supplements, scoring four models against four independent LLM judges spanning three model families. Three findings. First, the frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.6 = 4.31, Opus 4.7 = 4.30, GPT-5.5 = 4.09); all four judges place the open-weights baseline (Llama 3.3 70B = 3.15) last. Second, that gap concentrates on retrieval (2.23) and synthesis (2.15), not extraction (0.84). Third, Decision Usefulness shows the widest cross-model dispersion of any dimension (4.0 points on issuer profiling) and top-tier inter-judge agreement (mean r = 0.52). Plausibility is cheap. Bankability is the bar.
Reference graph
Works this paper leans on
-
[1]
Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2023). Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. arXiv:2310.11511
Pith/arXiv arXiv 2023
-
[2]
Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-H., Routledge, B., & Wang, W. Y. (2021). FinQA: A Dataset of Numerical Reasoning over Financial Data. EMNLP. arXiv:2109.00122
Pith/arXiv arXiv 2021
-
[3]
Chen, Z., Li, S., Smiley, C., Ma, Z., Shah, S., & Wang, W. Y. (2022). ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering. EMNLP. arXiv:2210.03849
Pith/arXiv arXiv 2022
-
[4]
Ahuja, P. (2026). cm-lrs: companion code release for ``CM-LRS'' [source code]. https://github.com/dsauce/cm-lrs
2026
-
[5]
Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217
Pith/arXiv arXiv 2023
-
[6]
European Union. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European Union, L series, 12 July 2024
2024
-
[7]
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring Massive Multitask Language Understanding. ICLR. arXiv:2009.03300
Pith/arXiv arXiv 2021
-
[8]
ISO/IEC. (2023). ISO/IEC 42001:2023 - Information technology - Artificial intelligence - Management system. International Organization for Standardization
2023
-
[9]
Islam, P., Kannappan, A., Kiela, D., Qian, R., Scherrer, N., & Vidgen, B. (2023). FinanceBench: A New Benchmark for Financial Question Answering. arXiv:2311.11944
Pith/arXiv arXiv 2023
-
[10]
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A., & Fung, P. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1--38
2023
-
[11]
uttler, H., Lewis, M., Yih, W.-t., Rockt\
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K\"uttler, H., Lewis, M., Yih, W.-t., Rockt\"aschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS. arXiv:2005.11401
Pith/arXiv arXiv 2020
-
[12]
Liang, P., Bommasani, R., Lee, T., et al. (2022). Holistic Evaluation of Language Models (HELM). Transactions on Machine Learning Research. arXiv:2211.09110
Pith/arXiv arXiv 2022
-
[13]
Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. ACL
2020
-
[14]
W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. EMNLP. arXiv:2305.14251
Pith/arXiv arXiv 2023
-
[15]
National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1
2023
-
[16]
Huang, Y., Sun, L., Wang, H., et al. (2024). TrustLLM: Trustworthiness in Large Language Models. arXiv:2401.05561
Pith/arXiv arXiv 2024
-
[17]
Suzgun, M., Scales, N., Sch\"arli, N., et al. (2022). Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. arXiv:2210.09261
Pith/arXiv arXiv 2022
-
[18]
Grattafiori, A., et al. (2024). The Llama 3 Herd of Models. arXiv:2407.21783
Pith/arXiv arXiv 2024
-
[19]
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A Large Language Model for Finance. arXiv:2303.17564
Pith/arXiv arXiv 2023
-
[20]
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks Track. arXiv:2306.05685
Pith/arXiv arXiv 2023
-
[21]
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP. arXiv:2303.16634
Pith/arXiv arXiv 2023
-
[22]
Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., & Sui, Z. (2023). Large Language Models are not Fair Evaluators. arXiv:2305.17926
Pith/arXiv arXiv 2023
-
[23]
Chiang, C.-H., & Lee, H. (2023). Can Large Language Models Be an Alternative to Human Evaluations? ACL. arXiv:2305.01937
Pith/arXiv arXiv 2023
-
[24]
Saaty, T. L. (1980). The Analytic Hierarchy Process: Planning, Priority Setting, Resource Allocation. McGraw-Hill
1980
-
[25]
Rezaei, J. (2015). Best-Worst Multi-Criteria Decision-Making Method. Omega, 53, 49--57
2015
-
[26]
Li, H., Cao, Y., Yu, Y., Javaji, S. R., Suchow, J. W., et al. (2025). InvestorBench: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent. ACL. arXiv:2412.18174
Pith/arXiv arXiv 2025
-
[27]
Mohsin, M. T. (2025). Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis. arXiv:2507.22936
arXiv 2025
-
[28]
H., Scardigli, A., Tang, L., Chen, W., Levkin, D., et al
Wang, S. H., Scardigli, A., Tang, L., Chen, W., Levkin, D., et al. (2023). MAUD: An Expert-Annotated Legal NLP Dataset for Merger Agreement Understanding. EMNLP. arXiv:2301.00876
Pith/arXiv arXiv 2023
-
[29]
Liu, S., Li, Z., Ma, R., Zhao, H., & Du, M. (2025). ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Contracts. arXiv:2508.03080
Pith/arXiv arXiv 2025
-
[30]
Kulkarni, S., & Kulkarni, Y. (2026). Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies. arXiv:2603.22651
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.