{"id":"44e0cec2-5f58-4753-b8de-18cab634027d","arxiv_id":"2607.21340","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.","lead":"This paper introduces CM-LRS, a seven-dimension 0–5 rubric that rates whether an LLM's capital-markets output is 'bankable' — traceable, numerically consistent, and reviewable. It reports that four frontier-style models cluster within 0.22 points while an open-weights model lags by about a point, but all scores come from LLM judges with no human-rater calibration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human-rater calibration leaves the central ordering claim unanchored: if LLM judges do not track human bankability decisions, the 0.22-point frontier cluster and Llama-last finding lose their deployment meaning; a stratified human pilot on the 104 outputs would settle it.","rationale":"The reader's weakest assumption identifies exactly the same concern: LLM-judge scores are used as a proxy for expert human bankability review, with no human calibration. I agree that this is the single most load-bearing point. The paper's own Section 10.1 concedes the absence of a human-rater study, and Section 5.4 reports a near-zero cross-family judge correlation (GPT-5.5 vs Gemini rho = 0.03). Both facts make the deployment-meaning of the numeric ordering conditional on an untested alignment between LLM judges and human reviewers. The rest of the paper—the rubric design, the workflow taxonomy, and the cross-judge robustness attempt—is solid and would remain useful even if the ordering claim were weakened to 'under LLM judging, these models order as...' But the abstract and bottom-line box go further, asserting a bankability-relevant cluster and open-weights lag. That stronger reading requires the human-calibration check. The reader's CONDITIONAL verdict is therefore appropriate; I would not change it. If the proposed pilot is conducted and supports the LLM-judge ordering, the verdict could be upgraded. If it does not, the central claim would need to be substantially softened. Since I am not proposing a different verdict than the reader's, I mark this UNCHANGED.","tokens_in":20006,"tokens_out":4081,"duration_ms":46840,"concrete_test":"Run a human-rater calibration pilot on a stratified sample of the 104 scored outputs. Select 30–40 outputs covering every workflow × model cell, with emphasis on W3 (largest D6 dispersion) and W4 (cross-judge disagreement). Have 3–5 practicing investment-banking analysts or compliance reviewers independently score each output using the Section 4 rubric, blind to model identity and to the LLM-judge scores. Pre-register agreement thresholds (e.g., Gwet's AC1 ≥ 0.6 per dimension and Spearman rho ≥ 0.7 on aggregate scores). Then recompute per-model CM-LRS from the human-consensus scores and compare with each LLM judge and the four-judge average. If the human-derived ordering places a different model last, or if the frontier-cluster band widens beyond 0.22, the deployment-relevant ordering claim fails. If human–LLM correlation is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.4 and Section 10.1 explicitly acknowledge that CM-LRS is 'an automated-judging baseline rather than a human-rater study.' Yet the paper's central claim—that the frontier closed-source models cluster within 0.22 points (Sonnet 4.31, Opus 4.30, GPT-5.5 4.09) and that all four judges place Llama 3.3 70B last—is a claim about deployment-relevant reliability. That claim is only meaningful if the four LLM judges' rubric scores approximate what a banker or compliance reviewer would accept or reject. This condition is untested. The protocol compounds the risk: the primary judge (Sonnet 4.6) is itself one of the evaluated models, and GPT-5.5 is both a panel model and a judge. The two out-of-panel judges that share neither family nor panel overlap disagree sharply on per-cell aggregates (GPT-5.5 vs Gemini Spearman rho = 0.03, Section 5.4), so the 'four-judge average' is not a converged measurement; it is an average of discordant rankings. If human raters applying the same rubric produced materially different per-cell scores—for example, ranking Llama above GPT-5.5 on W4 or failing to reproduce the 0.22-point cluster—then the headline ordering and the D6 'cleanest signal' conclusion would not survive. Because every reported finding flows through uncalibrated LLM-judge scores, this is the load-bearing assumption. The paper's transparency about the limitation is creditworthy, but transparency does not supply the missing empirical support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CM-LRS, a seven-dimension 0-5 rubric for scoring LLM workflow outputs in capital-markets settings at the workflow-output layer rather than the QA-pair layer. It defines a workflow taxonomy, an aggregate formula with tunable weights, and demonstrates the framework on five public/synthetic workflows, scoring four models (Claude Opus 4.7, GPT-5.5, Claude Sonnet 4.6, Llama 3.3 70B) with four LLM judges. Headline findings are that frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.31, Opus 4.30, GPT-5.5 4.09), Llama is last at 3.15, the open-weights gap concentrates on retrieval/synthesis rather than extraction, and Decision Usefulness (D6) is the cleanest dimension-level separator. The paper releases rubrics, prompts, outputs, and a deterministic verification script.","tokens_in":20352,"tokens_out":6888,"duration_ms":74088,"significance":"If the ordering claims held, CM-LRS would be a useful deployment-screening instrument. The strengths are real: the workflow-output framing addresses a genuine gap; the corpus is public or synthetic; the four-judge protocol is a serious attempt to control self- and family-bias; and the release of rubrics, prompts, and scoring outputs with a verification script is exemplary for reproducibility. However, the empirical headline currently rests on unvalidated LLM-as-judge scores: there is no human-rater calibration, no statistical significance testing, and at least one judge pair shows effectively independent rankings (GPT-5.5 vs Gemini rho = 0.03). These issues are load-bearing for the paper's central deployment-relevance claim, not presentation details. The framework is promising, but the evidence as presented does not yet establish 'bankability'.","major_comments":[{"comment":"The load-bearing assumption that LLM-judge scores proxy human 'bankability' decisions is untested. The paper explicitly concedes this is \"an automated-judging baseline rather than a human-rater study.\" The 0.22-point frontier cluster and the Llama-last ordering are deployment-meaningful only if judges track the human review bar. A stratified human pilot on a sample of the 104 outputs (same rubric, banker/compliance raters, inter-rater reliability, agreement with each judge) is needed to anchor the headline. Without it the central ordering is an unvalidated proxy.","section":"§5.4, §10.1"},{"comment":"The abstract and bottom line claim the three frontier models are \"statistically indistinguishable,\" but no significance test, confidence interval, or standard error is reported for the 0.22-point cluster or the 1.16-point gap. Pairwise judge Spearman correlations range from 0.03 (GPT-5.5 vs Gemini) to 0.94 (Sonnet vs Haiku); the four-judge mean is an average of discordant rankings. Report per-judge aggregate tables, judge×model interactions, and a bootstrap or mixed-effects test of the cluster/gap before making the indistinguishability claim.","section":"§5.4, Table 5, bottom line"},{"comment":"The conclusion states Sonnet \"wins or ties every workflow,\" but Table 5 is the primary judge's scoring and Sonnet is that judge. Section 5.4 itself says the in-cluster ranking is judge-dependent. This is a self-judged claim and should be removed or explicitly labelled as the primary judge's self-assessment. The four-judge protocol mitigates self-bias for the overall ordering, but not for this specific sentence about Table 5.","section":"§7, Table 5, Conclusion"},{"comment":"The claim that the frontier-vs-Llama gap is \"robust to any reasonable weight choice\" is unsupported. Eq. (1)'s weights and the deployment-readiness threshold are free parameters. A weight sweep over Table 3 defaults and plausible alternatives should be reported to demonstrate that the ordering and cluster membership do not flip; otherwise the claim should be softened to a conjecture.","section":"§4.3"}],"minor_comments":[{"comment":"The subsections are labeled B.1–B.4 inside Appendix C; renumber them to C.1–C.4.","section":"Appendix C"},{"comment":"\"The strongest single empirical signal ... that family-bias matters\" overstates a single Spearman correlation; present a formal test of family-bias or soften the wording.","section":"§5.4"},{"comment":"\"Four independent LLM judges\" is imprecise since GPT-5.5 is both a judge and an evaluated model; suggest \"four LLM judges, two in-panel and two out-of-panel.\"","section":"§5.2"},{"comment":"Cost figures would be easier to verify if the list-price retrieval date and source URL are reported.","section":"Table 4"},{"comment":"The acknowledged 6,000-token truncation is the largest external-validity threat for long documents; consider a small full-document pilot to quantify truncation sensitivity.","section":"§10.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent, and the artifact release is exemplary. The main risk is overclaiming deployment relevance from uncalibrated LLM-as-judge scores: the 'bankability' framing is not yet supported. A human-rater pilot and significance testing are feasible within a revision and would materially change the verdict. If those cannot be added, the paper should be reframed as a position/demonstration of a structured checklist rather than a validated reliability ordering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth reading for the rubric and the transparency, not for the model rankings. The rankings are only as solid as the LLM judges, and the paper admits it doesn't know how those judges track human reviewers.\n\nWhat's genuinely new: CM-LRS evaluates at the workflow-output layer, not QA-pair or component level. The seven dimensions are well-specified with practical anchors, and the paper ships rubrics, prompts, scoring files, and a verification script. That's a real gift to the community. The worked example in Appendix C is a good model of how to report rubric scoring.\n\nThe four-judge cross-validation is a serious attempt at controlling self-bias, and the paper is honest that the primary judge is also a scored model. The finding that all four judges place Llama last is the most robust result. The D6 decision-usefulness analysis is interesting, though the 4-point spread on one workflow is a single cell and should be treated as exploratory.\n\nSoft spots, in order of severity:\n\nFirst, no human-rater calibration. The paper calls itself an automated-judging baseline in Section 10, which is creditworthy, but the abstract and bottom line talk about 'bankable' and 'deployment-readiness' as if the score means something to a banker. It might, but it's untested. The 'statistically indistinguishable' phrase in the bottom line is not supported: no significance test is reported, and the 0.22-point cluster could easily arise from judge noise. The near-zero GPT-5.5–Gemini correlation on per-cell aggregates is a red flag that the four-judge average is averaging disagreeing views. The paper should either run a human pilot on a stratified sample of the outputs or clearly reframe all findings as LLM-judge-assessed reliability.\n\nSecond, the numbers don't fully reconcile. The paper says judges score the same 104 outputs, but summing the documented per-workflow inputs gives 41 outputs per model, 164 for four models. Maybe the repo explains this, but as written the reader can't tell. That's a fixable problem, but it undermines trust in the empirical claims.\n\nThird, the per-cell scores in Table 5 are primary-judge only, while the headline uses four-judge averages. That's fine, but the paper should be clearer that Table 5 is not the headline.\n\nWho it's for: people building evaluation frameworks for financial NLP, practitioners choosing LLMs for regulated workflows, and anyone studying LLM-as-judge reliability. The framework itself is a solid contribution and deserves peer review. The empirical section needs revision, not rejection.\n\nRecommendation: send to peer review. Ask for the human pilot or explicit reframing, reconciliation of the sample counts, and removal of the 'statistically indistinguishable' phrase unless a test is supplied.","headline":"A genuinely workflow-level reliability rubric for capital-markets LLM outputs, transparently demonstrated, but the empirical ordering claims rest on uncalibrated LLM judges and some internal numbers don't reconcile.","tokens_in":20883,"tokens_out":3770,"would_cite":true,"duration_ms":39074,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new seven-dimension reliability score for capital-markets LLM outputs finds frontier closed-source models clustered within 0.22 points, with the open-weights baseline ranked last by all judges.","keywords":["large language models","evaluation","reliability","capital markets","investment banking","AI governance","retrieval-augmented generation","financial NLP"],"falsifier":"Score the same 104 outputs with practising bankers and compliance reviewers using the same 0–5 rubric and record their accept/reject decisions; if human reviewers' model ranking differs from the four-judge LLM ordering — for instance, if they do not place Llama 3.3 70B last, or they find a wide gap among Sonnet 4.6, Opus 4.7, and GPT-5.5 — the claim that CM-LRS measures bankability would be refuted.","tokens_in":19824,"feed_emoji":"📊","tokens_out":8334,"duration_ms":74818,"temperature":0.7,"pith_summary":"CM-LRS is a proposed reliability metric for LLM outputs in regulated capital-markets workflows, scoring each output on seven 0–5 dimensions — factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability — with a weighted aggregate tunable to the workflow class. The paper's claim is that this workflow-output-layer score, unlike QA-pair benchmarks, captures whether a draft would survive review by a banker, analyst, or compliance officer — 'bankability'. Applying CM-LRS to five workflows with four LLM judges, it finds the three frontier closed-source models (Sonnet 4.6, Opus 4.7, GPT-5.5) clustered within 0.22 points, while the open-weights Llama 3.3 70B trails by about one point, with the gap concentrated in retrieval and synthesis rather than extraction. The intended consequence: deployment choices at the frontier can turn on cost, latency, and workflow fit, while open-weights models need targeted work on multi-document traceability before they can be bankable.","feed_headline":"Frontier LLMs cluster on bankability score; open-weights last","feed_subtitle":"Seven dimensions — traceability, numbers, completeness, usefulness — separate drafts a banker can defend from cheap plausibility.","key_machinery":"The key machinery is CM-LRS itself: a seven-dimension, 0–5 rubric with anchor descriptions per score level (0 = unusable, 5 = production-grade), aggregated as a weighted sum with workflow-class default weights; scoring is executed by four LLM judges (two in-panel, two out-of-panel, spanning three model families) using the same prompt and rubric, with pairwise Spearman correlations and per-dimension Pearson correlations reported to characterise judge agreement. The rubric is paired with a workflow taxonomy — extraction, retrieval, synthesis, comparison & reasoning, drafting — that sets default weights and directs which dimensions dominate; the paper's working principle, 'predict with the LLM,","core_discovery":"The central discovery is a measurement result: on the CM-LRS rubric, the three frontier closed-source models are effectively indistinguishable as a group — Sonnet 4.6 at 4.31, Opus 4.7 at 4.30, GPT-5.5 at 4.09 on four-judge averaged aggregates — and every judge places the open-weights baseline (Llama 3.3 70B, 3.15) last, a gap of 0.94–1.16 points. The gap is not uniform: it is 2.23 points on precedent retrieval and 2.15 on issuer-profile synthesis but only 0.84 on single-document debt-terms extraction. Decision Usefulness (D6) shows the largest cross-model dispersion in the study (a 4.0-point spread on the issuer-profile workflow) while sitting in the top tier of inter-judge agreement (mean","pith_inferences":["If the near-zero agreement between the two most independent judges (GPT-5.5 vs Gemini, Spearman rho = 0.03) reflects genuine judge-level noise, then the robust conclusions may be limited to the frontier-vs-open-weights gap and to D6's dispersion, not to finer distinctions among the closed-source models.","A testable extension is to apply CM-LRS to drafting workflows (pitch evidence, compliance memos), which the paper leaves untested; if D6 remains the separator there, the metric generalises beyond the five demonstrated classes.","Because all documents were truncated to 6,000 tokens, the demonstrated scores bias toward front-matter and cover-page content; a full-document chunked pipeline is the natural stress test of whether the frontier cluster survives at depth.","If the 'predict with the LLM, calculate with code, decide with the human' principle is adopted architecturally, numerical consistency and reviewability become build-time constraints rather than evaluation-only dimensions, turning CM-LRS from a benchmark into a compliance instrument."],"forward_implications":["For capital-markets buyers, headline reliability on these five workflows does not differentiate the three frontier closed-source models; cost, latency, and workflow-class fit become the deciding factors.","Open-weights LLMs may suffice for single-document structured extraction (gap 0.84) but are not yet bankable for multi-document retrieval and synthesis, where traceability failures dominate.","A deployment gate requiring CM-LRS ≥ 4.0, no dimension below 3, and no 0 or 1 scores gives a concrete, tunable proxy for reviewer readiness in regulated settings.","The gap between CM-LRS and QA-pair scores is itself a risk indicator: surface-correct but untraceable or incomplete outputs will systematically score lower on CM-LRS, which is the intended behaviour.","Decision Usefulness (D6) is the most reliable dimension-level separator, and combining it with numerical consistency (D3) and workflow completeness (D4) predicts whether the rest of an output is safe to admit."],"fun_headline_variants":["Bankable AI: Frontier LLMs tie, open-weights lag","LLM bankability: closed models cluster, open falls behind","Plausible vs bankable: frontier LLMs tie on 7-dim score","Retrieval, not extraction, splits AI drafts from bankable","Decision usefulness widest gap among LLM bankability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that four LLM judges' rubric scores are a valid proxy for expert human review of bankability; the paper explicitly states that this is 'an automated-judging baseline rather than a human-rater study' and that its primary judge (Claude Sonnet 4.6) is itself one of the models being scored — if LLM judges do not track what human reviewers would reject, the numeric scores and the model ordering lose their stated meaning.","fun_headline_variants_meta":{"raw":{"variants":["Bankable AI: Frontier LLMs tie, open-weights lag","LLM bankability: closed models cluster, open falls behind","Plausible vs bankable: frontier LLMs tie on 7-dim score","Retrieval, not extraction, splits AI drafts from bankable","Decision usefulness widest gap among LLM bankability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1322,"prompt_tokens":975,"completion_tokens":347,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":719,"completion_tokens_details":{"reasoning_tokens":257}},"tokens_in":719,"tokens_out":347,"duration_ms":4278,"temperature":1.0,"reasoning_tokens":257,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:43:39.321777+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score the same 104 outputs with practising bankers and compliance reviewers using the same 0–5 rubric and record their accept/reject decisions; if human reviewers' model ranking differs from the four-judge LLM ordering — for instance, if they do not place Llama 3.3 70B last, or they find a wide gap among Sonnet 4.6, Opus 4.7, and GPT-5.5 — the claim that CM-LRS measures bankability would be refuted.","supporting_citations":[],"review_version":1}