Pith. sign in

REVIEW 4 major objections 4 minor 50 references

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper argues that rubrics for judging financial AI agents should be built from authentic practitioner deliverables, and reports a 21.2-point held-out coverage advantage over prompt-only rubrics for role-specialized roles.

desk verdict A genuinely new source of rubric evidence, with a plausible regime split that is not yet independently verified because the held-out standards are LLM-extracted from the same document genre. read the letter →

arxiv 2608.04077 v1 pith:WKK4Z6EM submitted 2026-08-04 cs.AI cs.CL

classification cs.AIcs.CL
keywords financialAIagentsrubricconstructionprofessionaldeliverablesLLM-as-a-judgebenchmarkrole-groundedrubricstacitstandardsNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FinProBench and its Role-Grounded Rubric Construction (RGRC) pipeline argue that the right way to judge an AI agent's professional financial work is with criteria extracted from real practitioner deliverables, not from task prompts or model outputs. The paper's central evidence is a regime split: across 30 conventional roles whose genres are common online, prompt-only rubrics nearly match RGRC (89.2% versus 90.7% held-out coverage), but across 27 role-specialized roles with sparse public priors, RGRC reaches 99.1% versus 78.0%. That 21.2-point gap is the paper's proof that tacit professional standards exist beyond what prompts can recover. If correct, deliverable corpora become a reusable evaluation asset: role-level rubrics transfer across tasks within an occupation, cutting estimated per-task construction effort by 6.7 times.

What carries the argument

The load-bearing object is the role-grounded rubric: a two-layer scoring instrument whose normative layer is distilled once per occupation from a corpus of authentic practitioner deliverables and whose thin task-specific layer is added per task. The RGRC pipeline carries the argument through four stages — Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation — using an authenticity criterion (the author's role must match the target role), a 60% cross-document frequency threshold, four design principles (evaluability, unambiguity, non-hackability, orthogonality), and a scoring schema with weights in $\{1,2,3,5\}$ plus negative catastrophic-failure penalties. Because criteria come from deliverables rather than prompts, the rubric can name standards that are never stated in the task description.

What would settle it

Audit the held-out professional-standard extraction: check whether the held-out documents overlap the RGRC source corpus and whether the same LLM family and prompt style produced both the rubrics and the held-out standards. If either condition holds, the 99.1% versus 78.0% gap could reflect contamination or shared-bias artifacts rather than deliverable grounding; re-running the coverage comparison with independently practitioner-written rubrics as ground truth would settle it.

Watch

Extended reading notes

Core claim

The paper's central claim is that assessment criteria for open-ended financial work should be grounded in authentic work products produced by practitioners in the same occupational role, because those artifacts embody tacit quality standards that task descriptions and model outputs do not. RGRC operationalizes this in four stages: collect at least 20 authentic deliverables per role, extract competencies per document and then synthesize them across documents with a 60% frequency threshold, turn competencies into scored criteria obeying evaluability, unambiguity, non-hackability, and orthogonality, and validate via prompt-coverage, discriminative review, and a cross-evaluator agreement gate requiring $\kappa \geq 0.75$. The decisive result is the pre-classified split of 57 occupations into 30 prior-rich conventional and 27 prior-sparse role-specialized roles: the coverage advantage of RGRC over prompt-only rubrics is small in the conventional regime (90.7% versus 89.2%) and large in the role-specialized regime (99.1% versus 78.0%). Under a gold-held-out role-level rubric that privileges no system by construction, authentic human deliverables score highest on average (73.7) against three agent systems (70.3, 70.2, and 69.6), with overlapping 95% confidence intervals.

Load-bearing premise

The central result assumes that the professional standards extracted from held-out deliverables are a fair, unbiased representation of real professional quality and that those held-out documents are genuinely separate from the corpus RGRC learned from; the paper does not detail the extraction procedure or the overlap audit.

Editorial extensions

If this is right

  • If RGRC is right, professional-work benchmarks can be built from role-structured evidence corpora instead of hand-authored rubrics per task, cutting estimated per-task construction effort by 6.7 times through role-level reuse.
  • For occupations with well-represented public conventions, prompt-only rubric generation is nearly as good, so the method's value concentrates where public priors are thin.
  • Role-level rubrics transfer across tasks within a role, with 60-70% of criteria inherited, meaning that adding a new task to the benchmark requires mostly task-specific criteria.
  • The reported judge-panel stability (Fleiss' kappa = 0.76 and no material same-provider bias) suggests deliverable-grounded criteria can be scored reliably by automated judges.
  • Human deliverables leading on average while confidence intervals overlap implies the benchmark is suited to profile comparisons across quality dimensions rather than a single leaderboard winner.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the size of the RGRC advantage should shrink as public corpora for specialized roles grow; if regulatory and statutory documents become more widely available to model training, the prior-rich/prior-sparse boundary will move and the 21.2-point gap is a moving target rather than a fixed property of the method.
  • Editorial inference: a natural stress test is to replace the LLM-extracted held-out standards with rubrics written independently by practitioners; the paper's own limitations section concedes the current validation lacks practitioner consensus, which could narrow the 99.1% figure.
  • Editorial inference: deliverable-grounded rubrics could serve as reward-model training signal, since criteria tied to concrete observable artifacts may be less gameable than prompt-derived rubrics, though the paper does not test this.
  • Editorial inference: cross-jurisdiction transfer is the clearest next experiment; because the corpus is China-sourced, one would predict role-level rubrics transfer worse for jurisdiction-bound genres such as compliance and statutory formats than for more international genres like equity research.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces FinProBench, a benchmark for financial AI agents built from 1,723 authentic practitioner deliverables across 57 occupations, and RGRC, a four-stage pipeline that derives evaluation rubrics from those deliverables rather than from task prompts or model outputs. The central reported result is a coverage split: under a budget-matched comparison, Prompt@78 reaches 89.2% coverage on 30 'conventional' roles versus RGRC's 90.7%, but only 78.0% on 27 'role-specialized' roles versus RGRC's 99.1% (Table 1). The paper also reports a human-agent evaluation under a gold-held-out role-level rubric, with human deliverables ranking first on average (73.7 versus 70.3, 70.2, and 69.6) while all 95% confidence intervals overlap, and provides judge-reliability and self-preference analyses.

Significance. If the coverage result is valid, RGRC is a genuinely useful contribution: it moves rubric construction to authentic professional artifacts, introduces a reusable role-level rubric layer, and provides a concrete benchmark instantiation with real financial deliverables. The paper is careful on several internal-validity fronts: the conventional/role-specialized split is pre-analytic, the headline human-agent comparison uses a held-out rubric rather than the task rubric, confidence intervals are reported with no overclaiming of significance, and a self-preference check is included. The main weakness is that the coverage ground truth is itself LLM-extracted from the same deliverable genre and is not human-validated; combined with the absence of released artifacts, the central headline claim is currently a reproducibility-risk claim about automated coverage rather than a fully verified claim about professional standards.

major comments (4)
  1. [Professional-Standard Coverage] The ground truth for the headline coverage comparison is not independent in the way the interpretation requires. The text states that 'evaluation standards are independently extracted from held-out documents of the same type,' but it does not report the extraction model, prompt, number of held-out documents per role, inter-extractor agreement, or any blinding of the extractor to RGRC criteria. Because RGRC criteria are also LLM-generated from the same deliverable genre, the 99.1% versus 78.0% split could reflect shared LLM extraction behavior rather than authentic professional standards. The Limitations section correctly concedes that the work 'lacks independent practitioner annotation'; this concession means the central claim currently validates automated coverage only. Please add a human-validated subset of held-out standards, or otherwise show that the held-out extractor is not aligned with RGRC outputs, before the coverage split can be interpreted as evidence about tacit professional standards.
  2. [Professional-Standard Coverage; Table 1] The baseline Prompt@78 is not described anywhere in the Methods. The table reports an average of 79.2 criteria for Prompt@78, but the text never defines how the prompt-only rubric is generated, which LLM is used, what prompt template is used, or how the nominal 78-criterion budget is enforced. Since the central conclusion is a comparison between Prompt@78 and RGRC, the baseline construction is load-bearing. Please specify the generation procedure and budget-matching algorithm, and release the exact prompts and baseline rubrics.
  3. [Abstract; The FinProBench Benchmark] The abstract and benchmark section state that FinProBench 'releases an initial evaluation set of 20 complete tasks,' but the manuscript provides no repository URL, dataset download, or code release. Without the task prompts, rubrics, gold deliverables, held-out documents, and scoring code, none of the reported numbers—Table 1 coverage, Table 2 scores, and the Fleiss κ values—can be independently checked. For a benchmark contribution, artifact release is part of the claim; please provide a persistent repository with the full protocol and data.
  4. [Metrics; Eq. (5)] The relationship between the scoring schema and the reported binary judgments is ambiguous. Eq. (5) uses a binary indicator for each criterion, and the Metrics paragraph says 'a positive criterion counts as met when its awarded score reaches at least half its weight,' but Stage 3 describes graded criteria such as 'lists volatility, drawdown, Sharpe ratio, and Beta, scored 0–4 by the number covered.' It is unclear whether the LLM judges first assign a graded score and then threshold it, or whether they directly output a binary decision. Please clarify the scoring protocol and, if graded scores are used, report them or explain the thresholding explicitly.
minor comments (4)
  1. [Professional-Standard Coverage] The notation 'Prompt@78' is introduced only in Table 1; please define it at first use and explain what the '@78' refers to.
  2. [Abstract] The submitted text contains many formatting artifacts from PDF extraction (for example, missing spaces between words in the Abstract); please provide a clean camera-ready version.
  3. [From Role-Level to Task-Level Rubrics] The estimated 6.7× per-task effort reduction is stated in the Abstract and Conclusion, but the measurement basis is not described; please specify how construction times were measured or soften the claim to an internal estimate.
  4. [Method Overview; Figure 2] Figure 2 caption mentions a 'bounded synthesis-retry loop,' but the retry bound and the failure criteria for returning to Stage 3 are not specified in the text; please state them.

Circularity Check

1 steps flagged · score 6.0 of 10

Coverage ground truth is LLM-extracted from the same deliverable genre, so the headline 21.2-point RGRC advantage partly measures self-consistency rather than external professional standards; otherwise the derivation is independent.

  1. self definitional [Experiments > Professional-Standard Coverage; Limitations]
    "Evaluation standards are independently extracted from held-out documents of the same type, with all held-out documents audited against the RGRC source pool."

    RGRC Stage 2 builds rubrics by prompting an LLM to extract competencies and standards from authentic deliverables (Eq. 2: Ci = LLM extract(di, role_context(r))). The coverage ground truth is also standards 'extracted from held-out documents of the same type,' with no procedure, prompt, or model specified, and with no human validation. The target variable and the method's output are therefore the same operation applied to the same genre, so the 99.1% vs. 78.0% headline gap partly measures LLM self-consistency across two document samples rather than agreement with an independent professional standard. The Limitations passage confirms this: 'Our validation combines held-out professional deliverables with a heterogeneous LLM judge panel but lacks independent practitioner annotation.

full rationale

The derivation chain is mostly self-contained, and the authors are unusually careful about one obvious tautology: they explicitly call a human win under the task-specific rubric 'close to tautological' and report headline results under a gold-held-out role-level rubric from a disjoint corpus. That is a genuine mitigation, and I do not count it as circular. The circularity I do count is the operational definition of the evaluation standard in the headline coverage experiment. RGRC is, by construction, an LLM-based extraction of standards from professional deliverables (Eq. 2). The held-out ground truth used to measure coverage is also 'evaluation standards... extracted from held-out documents of the same type,' with no description of a different extractor, no human validation, and no extractor-blinding details. Consequently, the central 21.2-point advantage (99.1% vs. 78.0%) measures agreement between two extractions of the same genre by the same kind of pipeline, rather than correspondence between RGRC and an independent professional standard. The paper's own Limitations section concedes this, stating the results establish 'automated coverage, discriminative validity, and evaluator consistency rather than practitioner consensus.' This makes the central claim partially circular: the target and the method share the same definitional source. Other elements, including role-level rubric reuse, judge-panel reliability, and the human-agent comparison under the gold-held-out role-level rubric, are independent and non-circular. The score is 6 rather than higher because the limitation is disclosed, the held-out documents are audited against the source pool, and the transfer and reliability claims do not reduce to the same tautology.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

The central claims rest on assumptions about the authenticity of the corpus, the sharedness of professional standards, the fidelity of LLM extraction, the independence of held-out documents, and the validity of coverage as a quality proxy. These are reasonable domain assumptions but not independently verified.

free parameters (4)
  • τ_freq (frequency threshold) = 60%
    Stage 2 requires competencies to appear in at least 60% of analyzed deliverables; this hand-chosen threshold shapes which standards enter the role-level rubric.
  • Criterion budget for Prompt@78 = 78 criteria
    The prompt-only baseline is budget-matched to roughly 78 criteria to control for rubric size; the choice of 78 affects the comparison.
  • Criterion count formula weights = 10 per difficulty level
    |C_task| ≈ 10*difficulty + ε; difficulty is subjectively assigned per task and changes rubric length.
  • Weight distribution targets = 50% weight-1, 30% weight-2, 20% weight-3+
    Gate 1 enforces this distribution on positive weights, influencing the scoring outcome.
assumptions (6)
  • domain assumption Authenticity criterion: role(author(d)) ≈ r (Equation 1)
    Assumes author roles can be accurately determined from institutional authorship or content, which gates the whole corpus.
  • domain assumption Professionals in the same role share implicit quality standards embodied in work products
    Core premise of RGRC; if false, role-level rubrics have no stable target.
  • domain assumption LLM extraction from deliverables faithfully captures latent quality dimensions
    Stage 2 uses LLM analysis to extract competencies; fidelity is not separately verified.
  • domain assumption Held-out documents are independent of the RGRC source pool
    The coverage evaluation claims held-out documents were audited against the source pool, but the audit procedure is not described.
  • domain assumption Coverage of held-out standards is a valid proxy for rubric quality
    The central finding uses coverage as the metric; no evidence connects coverage to downstream evaluation quality.
  • domain assumption LLM judges can reliably apply rubric criteria
    Supported by κ=0.76, but this is still an LLM-based rather than practitioner-validated measure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables." pith.science (2026). https://pith.science/paper/WKK4Z6EM

@misc{pith2026260804077,
  author       = {Pith},
  title        = {Pith review of: FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WKK4Z6EM}},
  note         = {Machine review of arXiv:2608.04077}
}
read the original abstract

Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.

Figures

Figures reproduced from arXiv: 2608.04077 by the authors.

Figure 1
Figure 1. FinProBench coverage: 1,723 professional deliv [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The Role-Grounded Rubric Construction (RGRC) pipeline. Starting from a corpus of [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Per-dimension positive-hit rate (%) under the gold [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

50 extracted references · 24 canonical work pages

  1. [1]

    Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022. Constitutional AI : Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073

  2. [2]

    A.; and Terry, M

    Bradley, R. A.; and Terry, M. E. 1952. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4): 324--345

  3. [3]

    Chen, W.; Wang, Q.; Long, Z.; Zhang, X.; Lu, Z.; et al. 2023. DISC-FinLLM : A C hinese Financial Large Language Model based on Multiple Experts Fine-tuning. arXiv preprint arXiv:2310.15205

  4. [4]

    N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M

    Chiang, W.-L.; Zheng, L.; Sheng, Y.; Angelopoulos, A. N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M. I.; Gonzalez, J. E.; and Stoica, I. 2024. Chatbot Arena: An Open Platform for Evaluating LLM s by Human Preference. In Proceedings of the 41st International Conference on Machine Learning

  5. [5]

    Cohen, J. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1): 37--46

  6. [6]

    Fan, Z.; Chen, R.; Hu, T.; Peng, R.; Huang, Z.; Xu, H.; Chen, Y.; Wu, J.; Zhao, J.; and Liu, Z. 2026. OptimSyn : Influence-Guided Rubrics Optimization for Synthetic Data Generation. arXiv preprint arXiv:2604.00536

  7. [7]

    E.; R \'e , C.; Chilton, A.; Narayana, A.; Chohlas-Wood, A.; Peters, A.; Waldon, B.; Rockmore, D

    Guha, N.; Nyarko, J.; Ho, D. E.; R \'e , C.; Chilton, A.; Narayana, A.; Chohlas-Wood, A.; Peters, A.; Waldon, B.; Rockmore, D. N.; et al. 2023. LegalBench : A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In Advances in Neural Information Processing Systems

  8. [8]

    Guo, X.; Xia, H.; Liu, Z.; Cao, H.; Yang, Z.; et al. 2023. FinEval : A C hinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models. arXiv preprint arXiv:2308.09975

Show all 50 references
  1. [9]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations

  2. [10]

    Islam, P.; Kannappan, A.; Kiber, D.; Decena, H.; Kulkarni, S.; Hantous, N.; and Scott, A. 2023. FinanceBench : A New Benchmark for Financial Question Answering. arXiv preprint arXiv:2311.11944

  3. [11]

    E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K

    Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. R. 2024. SWE-bench : Can Language Models Resolve Real-world G it H ub Issues? In International Conference on Learning Representations

  4. [12]

    Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2020. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. arXiv preprint arXiv:2009.13081

  5. [13]

    Kim, S.; Shin, J.; Cho, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, S.; Kim, S.; Thorne, J.; and Seo, M. 2024. Prometheus : Inducing Fine-grained Evaluation Capability in Language Models. In International Conference on Learning Representations

  6. [14]

    Y.; Ramamurthy, R.; Sheng, Y.; Coste, T.; Nandwani, Y.; et al

    Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Ramamurthy, R.; Sheng, Y.; Coste, T.; Nandwani, Y.; et al. 2024. RewardBench : Evaluating Reward Models for Language Modeling. arXiv preprint arXiv:2403.13787

  7. [15]

    Lei, Y.; Li, J.; Cheng, D.; Ding, Z.; and Jiang, C. 2023. Chinese Financial Assistant Benchmark for Large Language Model. arXiv preprint arXiv:2311.05812

  8. [16]

    Li, S.; Zhao, J.; Wei, M.; Ren, H.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Chen, W. 2026 a . RubricHub : A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. arXiv preprint arXiv:2601.08430

  9. [17]

    Li, X.; Bao, K.; Li, M.; Ma, Y.; Zhang, Y.; Wang, W.; Feng, F.; and Liu, D. 2026 b . ARES : Automated Rubric Synthesis for Scalable LLM Reinforcement Learning. arXiv preprint arXiv:2605.23454

  10. [18]

    W.; Ramasubramanian, B.; Niu, L.; Yue, X.; and Poovendran, R

    Li, Y.; Feng, Y.; Xu, Z.; Ma, Z.; Zheng, K.; Jiang, F.; Sun, X.; Shao, R.; Chen, Z.; Huang, Y.; Han, X.; Lee, B.; Xu, K.; Zeng, S.; Hua, H.; Zhang, X.; Alomair, B.; Krishna, R.; Zettlemoyer, L.; Koh, P. W.; Ramasubramanian, B.; Niu, L.; Yue, X.; and Poovendran, R. 2026 c . Job...

  11. [19]

    Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. 2023. Holistic Evaluation of Language Models. Transactions on Machine Learning Research

  12. [20]

    Y.; Deng, Y.; Chandu, K.; Brahman, F.; Bhagavatula, C.; and Choi, Y

    Lin, B. Y.; Deng, Y.; Chandu, K.; Brahman, F.; Bhagavatula, C.; and Choi, Y. 2025. WildBench : Benchmarking LLMs with Challenging Tasks from Real Users in the Wild. In International Conference on Learning Representations

  13. [21]

    Liu, T.; Xu, R.; Yu, T.; Hong, I.; Yang, C.; Zhao, T.; and Wang, H. 2025. OpenRubrics : Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment. arXiv preprint arXiv:2510.07743

  14. [22]

    Liu, W.; Jin, J.; Huang, Z.; Wen, T.; Dong, G.; Zhao, Z.; Zhu, Y.; Dou, Z.; and Wen, J.-R. 2026. The Rules of the Game: A Survey of Rubrics for Large Language Models. https://openreview.net/forum?id=FnSimngGYk. OpenReview preprint

  15. [23]

    Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. 2023 a . AgentBench : Evaluating LLM s as Agents. arXiv preprint arXiv:2308.03688

  16. [24]

    Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023 b . G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

  17. [25]

    Luan, B.; Sun, R.; Wang, S.; Gu, Y.; Li, C.; Xiong, Z.; Li, J.; and Bai, Z. 2026. FinResearchBench II : A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality. arXiv preprint arXiv:2607.12252

  18. [26]

    Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; and Scialom, T. 2024. GAIA : a benchmark for General AI Assistants. In International Conference on Learning Representations

  19. [27]

    Nie, Y.; Yan, B.; Guo, T.; Liu, H.; Wang, H.; He, W.; Zheng, B.; Wang, W.; Li, Q.; Sun, W.; Wang, Y.; and Tao, D. 2024. CFinBench : A Comprehensive C hinese Financial Benchmark for Large Language Models. In Advances in Neural Information Processing Systems

  20. [28]

    OpenAI . 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774

  21. [29]

    L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training Language Models to Follow Instructions with Human Feedback. Advances in Neural Information Processing Systems, 35: 27730--27744

  22. [30]

    P.; Aljubeh, M.; Thacker, P.; Fauconnet, L.; Kim, N

    Patwardhan, T.; Dias, R.; Proehl, E.; Kim, G.; Wang, M.; Watkins, O.; Fishman, S. P.; Aljubeh, M.; Thacker, P.; Fauconnet, L.; Kim, N. S.; Chao, P.; Miserendino, S.; Chabot, G.; Li, D.; Sharman, M.; Barr, A.; Glaese, A.; and Tworek, J. 2025. GDPval : Evaluating AI Model Perfor...

  23. [31]

    D.; Ermon, S.; and Finn, C

    Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems

  24. [32]

    L.; Stickland, A

    Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; and Bowman, S. R. 2024. GPQA : A Graduate-Level G oogle-Proof Q & A Benchmark. In Conference on Language Modeling

  25. [33]

    S.; Chawla, K.; Eidnani, D.; Shah, A.; Du, W.; Chava, S.; Raman, N.; Smiley, C.; Chen, J.; and Yang, D

    Shah, R. S.; Chawla, K.; Eidnani, D.; Shah, A.; Du, W.; Chava, S.; Raman, N.; Smiley, C.; Chen, J.; and Yang, D. 2022. When FLUE Meets FLANG : Benchmarks and Large Pretrained Language Model for Financial Domain. In Proceedings of the 2022 Conference on Empirical Methods in Nat...

  26. [34]

    F.; Qiu, X.; Whitehouse, C.; Alazraki, L.; Goel, S.; Barbieri, F.; Willi, T.; Mathur, A.; and Leontiadis, I

    Shen, W. F.; Qiu, X.; Whitehouse, C.; Alazraki, L.; Goel, S.; Barbieri, F.; Willi, T.; Mathur, A.; and Leontiadis, I. 2026. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks. arXiv preprint arXiv:2602.05125

  27. [35]

    Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2022. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. arXiv preprint arXiv:2206.04615

  28. [36]

    Sun, Y.; Han, X.; Zhang, W.; Pang, Y.; Wang, T.; Cao, Y.; Huang, Y.; Duroiu, C.; Zhang, H.; Lin, J.; et al. 2026. Agents' Last Exam. arXiv preprint arXiv:2606.05405

  29. [37]

    R.; Zhang, S.; Sun, Y.; and Wang, W

    Wang, X.; Hu, Z.; Lu, P.; Zhu, Y.; Zhang, J.; Subramaniam, S.; Loomba, A. R.; Zhang, S.; Sun, Y.; and Wang, W. 2024 a . SciBench : Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models. In Proceedings of the 41st International Conference on Mac...

  30. [38]

    Wang, Y.; Ma, X.; Zhang, G.; Ni, Y.; Chandra, A.; Guo, S.; Ren, W.; Arulraj, A.; He, X.; Jiang, Z.; Li, T.; Ku, M.; Wang, K.; Zhuang, A.; Fan, R.; Yue, X.; and Chen, W. 2024 b . MMLU-Pro : A More Robust and Challenging Multi-Task Language Understanding Benchmark. In Advances i...

  31. [39]

    Wu, S.; Irsoy, O.; Lu, S.; Daber \'e ri, V.; Dredze, M.; Gehrmann, S.; Kambadur, P.; Rosenberg, D.; and Mann, G. 2023. BloombergGPT : A Large Language Model for Finance. arXiv preprint arXiv:2303.17564

  32. [40]

    Xie, L.; Huang, S.; Zhang, Z.; Zou, A.; Zhai, Y.; Ren, D.; Zhang, K.; Hu, H.; Liu, B.; Chen, H.; Liu, Z.; and Ding, B. 2025. Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling. arXiv preprint arXiv:2510.17314

  33. [41]

    Xie, Q.; Han, W.; Chen, Z.; Xia, R.; Zhang, X.; He, Y.; Xiao, M.; Li, D.; Dai, Y.; Feng, D.; et al. 2024 a . FinBen : A Holistic Financial Benchmark for Large Language Models. In Advances in Neural Information Processing Systems

  34. [42]

    Xie, Q.; Han, W.; Zhang, X.; Lai, Y.; Peng, M.; Lopez-Lira, A.; and Huang, J. 2023. PIXIU : A Large Language Model, Instruction Data and Evaluation Benchmark for Finance. In Advances in Neural Information Processing Systems

  35. [43]

    J.; Cheng, Z.; Shin, D.; Lei, F.; et al

    Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; et al. 2024 b . OSWorld : Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Advances in Neural Information Processing Systems

  36. [44]

    F.; Li, Y.; Zhu, B.; Bisk, Y.; and Neubig, G

    Xu, F. F.; Li, Y.; Zhu, B.; Bisk, Y.; and Neubig, G. 2024. TheAgentCompany : Benchmarking LLM Agents on Consequential Real World Tasks. arXiv preprint arXiv:2412.14161

  37. [45]

    Xu, Y.; Potje, G.; Shandilya, S.; Yuan, T.; de Oliveira Nunes, L.; Agarwal, R.; Asgari, S.; Atkinson, A.; K c man, E.; Lu, S.; Chandra, R.; and Chakraborty, T. 2026. SibylSense : Adaptive Rubric Learning via Memory Tuning and Adversarial Probing. arXiv preprint arXiv:2602.20751

  38. [46]

    Yang, H.; Liu, X.-Y.; and Wang, C. D. 2023. FinGPT : Open-Source Financial Large Language Models. arXiv preprint arXiv:2306.06031

  39. [47]

    Zhang, Q.; Zhou, J.; Wang, Y.; Lyu, F.; Ming, Y.; Xu, C.; Sun, Q.; Zheng, K.; Kang, P.; Liu, X.; and Ma, C. 2026. RubricBench : Aligning Model-Generated Rubrics with Human Standards. arXiv preprint arXiv:2603.01562

  40. [48]

    P.; Zhang, H.; Gonzalez, J

    Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023. Judging LLM -as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems

  41. [49]

    Zhong, J.; Zhang, H.; Southern, C.; Yang, J.; Wang, T.; Jung, K.; Zhang, S.; Yarats, D.; Ho, J.; and Ma, J. 2026. DRACO : A Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity. arXiv preprint arXiv:2602.11685

  42. [50]

    F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; Alon, U.; and Neubig, G

    Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; Alon, U.; and Neubig, G. 2024. WebArena : A Realistic Web Environment for Building Autonomous Agents. In International Conference on Learning Representations

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.