REVIEW 4 major objections 4 minor 50 references
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper argues that rubrics for judging financial AI agents should be built from authentic practitioner deliverables, and reports a 21.2-point held-out coverage advantage over prompt-only rubrics for role-specialized roles.
desk verdict A genuinely new source of rubric evidence, with a plausible regime split that is not yet independently verified because the held-out standards are LLM-extracted from the same document genre. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the role-grounded rubric: a two-layer scoring instrument whose normative layer is distilled once per occupation from a corpus of authentic practitioner deliverables and whose thin task-specific layer is added per task. The RGRC pipeline carries the argument through four stages — Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation — using an authenticity criterion (the author's role must match the target role), a 60% cross-document frequency threshold, four design principles (evaluability, unambiguity, non-hackability, orthogonality), and a scoring schema with weights in $\{1,2,3,5\}$ plus negative catastrophic-failure penalties. Because criteria come from deliverables rather than prompts, the rubric can name standards that are never stated in the task description.
What would settle it
Audit the held-out professional-standard extraction: check whether the held-out documents overlap the RGRC source corpus and whether the same LLM family and prompt style produced both the rubrics and the held-out standards. If either condition holds, the 99.1% versus 78.0% gap could reflect contamination or shared-bias artifacts rather than deliverable grounding; re-running the coverage comparison with independently practitioner-written rubrics as ground truth would settle it.
Extended reading notes
Core claim
The paper's central claim is that assessment criteria for open-ended financial work should be grounded in authentic work products produced by practitioners in the same occupational role, because those artifacts embody tacit quality standards that task descriptions and model outputs do not. RGRC operationalizes this in four stages: collect at least 20 authentic deliverables per role, extract competencies per document and then synthesize them across documents with a 60% frequency threshold, turn competencies into scored criteria obeying evaluability, unambiguity, non-hackability, and orthogonality, and validate via prompt-coverage, discriminative review, and a cross-evaluator agreement gate requiring $\kappa \geq 0.75$. The decisive result is the pre-classified split of 57 occupations into 30 prior-rich conventional and 27 prior-sparse role-specialized roles: the coverage advantage of RGRC over prompt-only rubrics is small in the conventional regime (90.7% versus 89.2%) and large in the role-specialized regime (99.1% versus 78.0%). Under a gold-held-out role-level rubric that privileges no system by construction, authentic human deliverables score highest on average (73.7) against three agent systems (70.3, 70.2, and 69.6), with overlapping 95% confidence intervals.
Load-bearing premise
The central result assumes that the professional standards extracted from held-out deliverables are a fair, unbiased representation of real professional quality and that those held-out documents are genuinely separate from the corpus RGRC learned from; the paper does not detail the extraction procedure or the overlap audit.
Editorial extensions
If this is right
- If RGRC is right, professional-work benchmarks can be built from role-structured evidence corpora instead of hand-authored rubrics per task, cutting estimated per-task construction effort by 6.7 times through role-level reuse.
- For occupations with well-represented public conventions, prompt-only rubric generation is nearly as good, so the method's value concentrates where public priors are thin.
- Role-level rubrics transfer across tasks within a role, with 60-70% of criteria inherited, meaning that adding a new task to the benchmark requires mostly task-specific criteria.
- The reported judge-panel stability (Fleiss' kappa = 0.76 and no material same-provider bias) suggests deliverable-grounded criteria can be scored reliably by automated judges.
- Human deliverables leading on average while confidence intervals overlap implies the benchmark is suited to profile comparisons across quality dimensions rather than a single leaderboard winner.
Reading between the lines
- Editorial inference: the size of the RGRC advantage should shrink as public corpora for specialized roles grow; if regulatory and statutory documents become more widely available to model training, the prior-rich/prior-sparse boundary will move and the 21.2-point gap is a moving target rather than a fixed property of the method.
- Editorial inference: a natural stress test is to replace the LLM-extracted held-out standards with rubrics written independently by practitioners; the paper's own limitations section concedes the current validation lacks practitioner consensus, which could narrow the 99.1% figure.
- Editorial inference: deliverable-grounded rubrics could serve as reward-model training signal, since criteria tied to concrete observable artifacts may be less gameable than prompt-derived rubrics, though the paper does not test this.
- Editorial inference: cross-jurisdiction transfer is the clearest next experiment; because the corpus is China-sourced, one would predict role-level rubrics transfer worse for jurisdiction-bound genres such as compliance and statutory formats than for more international genres like equity research.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FinProBench, a benchmark for financial AI agents built from 1,723 authentic practitioner deliverables across 57 occupations, and RGRC, a four-stage pipeline that derives evaluation rubrics from those deliverables rather than from task prompts or model outputs. The central reported result is a coverage split: under a budget-matched comparison, Prompt@78 reaches 89.2% coverage on 30 'conventional' roles versus RGRC's 90.7%, but only 78.0% on 27 'role-specialized' roles versus RGRC's 99.1% (Table 1). The paper also reports a human-agent evaluation under a gold-held-out role-level rubric, with human deliverables ranking first on average (73.7 versus 70.3, 70.2, and 69.6) while all 95% confidence intervals overlap, and provides judge-reliability and self-preference analyses.
Significance. If the coverage result is valid, RGRC is a genuinely useful contribution: it moves rubric construction to authentic professional artifacts, introduces a reusable role-level rubric layer, and provides a concrete benchmark instantiation with real financial deliverables. The paper is careful on several internal-validity fronts: the conventional/role-specialized split is pre-analytic, the headline human-agent comparison uses a held-out rubric rather than the task rubric, confidence intervals are reported with no overclaiming of significance, and a self-preference check is included. The main weakness is that the coverage ground truth is itself LLM-extracted from the same deliverable genre and is not human-validated; combined with the absence of released artifacts, the central headline claim is currently a reproducibility-risk claim about automated coverage rather than a fully verified claim about professional standards.
major comments (4)
- [Professional-Standard Coverage] The ground truth for the headline coverage comparison is not independent in the way the interpretation requires. The text states that 'evaluation standards are independently extracted from held-out documents of the same type,' but it does not report the extraction model, prompt, number of held-out documents per role, inter-extractor agreement, or any blinding of the extractor to RGRC criteria. Because RGRC criteria are also LLM-generated from the same deliverable genre, the 99.1% versus 78.0% split could reflect shared LLM extraction behavior rather than authentic professional standards. The Limitations section correctly concedes that the work 'lacks independent practitioner annotation'; this concession means the central claim currently validates automated coverage only. Please add a human-validated subset of held-out standards, or otherwise show that the held-out extractor is not aligned with RGRC outputs, before the coverage split can be interpreted as evidence about tacit professional standards.
- [Professional-Standard Coverage; Table 1] The baseline Prompt@78 is not described anywhere in the Methods. The table reports an average of 79.2 criteria for Prompt@78, but the text never defines how the prompt-only rubric is generated, which LLM is used, what prompt template is used, or how the nominal 78-criterion budget is enforced. Since the central conclusion is a comparison between Prompt@78 and RGRC, the baseline construction is load-bearing. Please specify the generation procedure and budget-matching algorithm, and release the exact prompts and baseline rubrics.
- [Abstract; The FinProBench Benchmark] The abstract and benchmark section state that FinProBench 'releases an initial evaluation set of 20 complete tasks,' but the manuscript provides no repository URL, dataset download, or code release. Without the task prompts, rubrics, gold deliverables, held-out documents, and scoring code, none of the reported numbers—Table 1 coverage, Table 2 scores, and the Fleiss κ values—can be independently checked. For a benchmark contribution, artifact release is part of the claim; please provide a persistent repository with the full protocol and data.
- [Metrics; Eq. (5)] The relationship between the scoring schema and the reported binary judgments is ambiguous. Eq. (5) uses a binary indicator for each criterion, and the Metrics paragraph says 'a positive criterion counts as met when its awarded score reaches at least half its weight,' but Stage 3 describes graded criteria such as 'lists volatility, drawdown, Sharpe ratio, and Beta, scored 0–4 by the number covered.' It is unclear whether the LLM judges first assign a graded score and then threshold it, or whether they directly output a binary decision. Please clarify the scoring protocol and, if graded scores are used, report them or explain the thresholding explicitly.
minor comments (4)
- [Professional-Standard Coverage] The notation 'Prompt@78' is introduced only in Table 1; please define it at first use and explain what the '@78' refers to.
- [Abstract] The submitted text contains many formatting artifacts from PDF extraction (for example, missing spaces between words in the Abstract); please provide a clean camera-ready version.
- [From Role-Level to Task-Level Rubrics] The estimated 6.7× per-task effort reduction is stated in the Abstract and Conclusion, but the measurement basis is not described; please specify how construction times were measured or soften the claim to an internal estimate.
- [Method Overview; Figure 2] Figure 2 caption mentions a 'bounded synthesis-retry loop,' but the retry bound and the failure criteria for returning to Stage 3 are not specified in the text; please state them.
Circularity Check
Coverage ground truth is LLM-extracted from the same deliverable genre, so the headline 21.2-point RGRC advantage partly measures self-consistency rather than external professional standards; otherwise the derivation is independent.
-
self definitional
[Experiments > Professional-Standard Coverage; Limitations]
"Evaluation standards are independently extracted from held-out documents of the same type, with all held-out documents audited against the RGRC source pool."
RGRC Stage 2 builds rubrics by prompting an LLM to extract competencies and standards from authentic deliverables (Eq. 2: Ci = LLM extract(di, role_context(r))). The coverage ground truth is also standards 'extracted from held-out documents of the same type,' with no procedure, prompt, or model specified, and with no human validation. The target variable and the method's output are therefore the same operation applied to the same genre, so the 99.1% vs. 78.0% headline gap partly measures LLM self-consistency across two document samples rather than agreement with an independent professional standard. The Limitations passage confirms this: 'Our validation combines held-out professional deliverables with a heterogeneous LLM judge panel but lacks independent practitioner annotation.
full rationale
The derivation chain is mostly self-contained, and the authors are unusually careful about one obvious tautology: they explicitly call a human win under the task-specific rubric 'close to tautological' and report headline results under a gold-held-out role-level rubric from a disjoint corpus. That is a genuine mitigation, and I do not count it as circular. The circularity I do count is the operational definition of the evaluation standard in the headline coverage experiment. RGRC is, by construction, an LLM-based extraction of standards from professional deliverables (Eq. 2). The held-out ground truth used to measure coverage is also 'evaluation standards... extracted from held-out documents of the same type,' with no description of a different extractor, no human validation, and no extractor-blinding details. Consequently, the central 21.2-point advantage (99.1% vs. 78.0%) measures agreement between two extractions of the same genre by the same kind of pipeline, rather than correspondence between RGRC and an independent professional standard. The paper's own Limitations section concedes this, stating the results establish 'automated coverage, discriminative validity, and evaluator consistency rather than practitioner consensus.' This makes the central claim partially circular: the target and the method share the same definitional source. Other elements, including role-level rubric reuse, judge-panel reliability, and the human-agent comparison under the gold-held-out role-level rubric, are independent and non-circular. The score is 6 rather than higher because the limitation is disclosed, the held-out documents are audited against the source pool, and the transfer and reliability claims do not reduce to the same tautology.
Assumptions & free parameters
free parameters (4)
- τ_freq (frequency threshold) =
60%
- Criterion budget for Prompt@78 =
78 criteria
- Criterion count formula weights =
10 per difficulty level
- Weight distribution targets =
50% weight-1, 30% weight-2, 20% weight-3+
assumptions (6)
- domain assumption Authenticity criterion: role(author(d)) ≈ r (Equation 1)
- domain assumption Professionals in the same role share implicit quality standards embodied in work products
- domain assumption LLM extraction from deliverables faithfully captures latent quality dimensions
- domain assumption Held-out documents are independent of the RGRC source pool
- domain assumption Coverage of held-out standards is a valid proxy for rubric quality
- domain assumption LLM judges can reliably apply rubric criteria
Cite this review
Pith. "Pith review of FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables." pith.science (2026). https://pith.science/paper/WKK4Z6EM
@misc{pith2026260804077,
author = {Pith},
title = {Pith review of: FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables},
year = {2026},
howpublished = {\url{https://pith.science/paper/WKK4Z6EM}},
note = {Machine review of arXiv:2608.04077}
}
read the original abstract
Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.
Figures
Reference graph
Works this paper leans on
-
[1]
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022. Constitutional AI : Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073
arXiv 2022
-
[2]
Bradley, R. A.; and Terry, M. E. 1952. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4): 324--345
work page 1952
-
[3]
Chen, W.; Wang, Q.; Long, Z.; Zhang, X.; Lu, Z.; et al. 2023. DISC-FinLLM : A C hinese Financial Large Language Model based on Multiple Experts Fine-tuning. arXiv preprint arXiv:2310.15205
arXiv 2023
-
[4]
N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M
Chiang, W.-L.; Zheng, L.; Sheng, Y.; Angelopoulos, A. N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M. I.; Gonzalez, J. E.; and Stoica, I. 2024. Chatbot Arena: An Open Platform for Evaluating LLM s by Human Preference. In Proceedings of the 41st International Conference on Machine Learning
work page 2024
-
[5]
Cohen, J. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1): 37--46
work page 1960
- [6]
-
[7]
E.; R \'e , C.; Chilton, A.; Narayana, A.; Chohlas-Wood, A.; Peters, A.; Waldon, B.; Rockmore, D
Guha, N.; Nyarko, J.; Ho, D. E.; R \'e , C.; Chilton, A.; Narayana, A.; Chohlas-Wood, A.; Peters, A.; Waldon, B.; Rockmore, D. N.; et al. 2023. LegalBench : A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In Advances in Neural Information Processing Systems
work page 2023
-
[8]
Guo, X.; Xia, H.; Liu, Z.; Cao, H.; Yang, Z.; et al. 2023. FinEval : A C hinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models. arXiv preprint arXiv:2308.09975
arXiv 2023
Show all 50 references
-
[9]
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations
2021
-
[10]
Islam, P.; Kannappan, A.; Kiber, D.; Decena, H.; Kulkarni, S.; Hantous, N.; and Scott, A. 2023. FinanceBench : A New Benchmark for Financial Question Answering. arXiv preprint arXiv:2311.11944
2023 arXiv
-
[11]
E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K
Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. R. 2024. SWE-bench : Can Language Models Resolve Real-world G it H ub Issues? In International Conference on Learning Representations
2024
-
[12]
Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2020. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. arXiv preprint arXiv:2009.13081
2020 arXiv
-
[13]
Kim, S.; Shin, J.; Cho, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, S.; Kim, S.; Thorne, J.; and Seo, M. 2024. Prometheus : Inducing Fine-grained Evaluation Capability in Language Models. In International Conference on Learning Representations
2024
-
[14]
Y.; Ramamurthy, R.; Sheng, Y.; Coste, T.; Nandwani, Y.; et al
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Ramamurthy, R.; Sheng, Y.; Coste, T.; Nandwani, Y.; et al. 2024. RewardBench : Evaluating Reward Models for Language Modeling. arXiv preprint arXiv:2403.13787
2024 arXiv
-
[15]
Lei, Y.; Li, J.; Cheng, D.; Ding, Z.; and Jiang, C. 2023. Chinese Financial Assistant Benchmark for Large Language Model. arXiv preprint arXiv:2311.05812
2023 arXiv
-
[16]
Li, S.; Zhao, J.; Wei, M.; Ren, H.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Chen, W. 2026 a . RubricHub : A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. arXiv preprint arXiv:2601.08430
2026
-
[17]
Li, X.; Bao, K.; Li, M.; Ma, Y.; Zhang, Y.; Wang, W.; Feng, F.; and Liu, D. 2026 b . ARES : Automated Rubric Synthesis for Scalable LLM Reinforcement Learning. arXiv preprint arXiv:2605.23454
2026 arXiv
-
[18]
W.; Ramasubramanian, B.; Niu, L.; Yue, X.; and Poovendran, R
Li, Y.; Feng, Y.; Xu, Z.; Ma, Z.; Zheng, K.; Jiang, F.; Sun, X.; Shao, R.; Chen, Z.; Huang, Y.; Han, X.; Lee, B.; Xu, K.; Zeng, S.; Hua, H.; Zhang, X.; Alomair, B.; Krishna, R.; Zettlemoyer, L.; Koh, P. W.; Ramasubramanian, B.; Niu, L.; Yue, X.; and Poovendran, R. 2026 c . Job...
2026 arXiv
-
[19]
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. 2023. Holistic Evaluation of Language Models. Transactions on Machine Learning Research
2023
-
[20]
Y.; Deng, Y.; Chandu, K.; Brahman, F.; Bhagavatula, C.; and Choi, Y
Lin, B. Y.; Deng, Y.; Chandu, K.; Brahman, F.; Bhagavatula, C.; and Choi, Y. 2025. WildBench : Benchmarking LLMs with Challenging Tasks from Real Users in the Wild. In International Conference on Learning Representations
2025
-
[21]
Liu, T.; Xu, R.; Yu, T.; Hong, I.; Yang, C.; Zhao, T.; and Wang, H. 2025. OpenRubrics : Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment. arXiv preprint arXiv:2510.07743
2025
-
[22]
Liu, W.; Jin, J.; Huang, Z.; Wen, T.; Dong, G.; Zhao, Z.; Zhu, Y.; Dou, Z.; and Wen, J.-R. 2026. The Rules of the Game: A Survey of Rubrics for Large Language Models. https://openreview.net/forum?id=FnSimngGYk. OpenReview preprint
2026
-
[23]
Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. 2023 a . AgentBench : Evaluating LLM s as Agents. arXiv preprint arXiv:2308.03688
2023 arXiv
-
[24]
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023 b . G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[25]
Luan, B.; Sun, R.; Wang, S.; Gu, Y.; Li, C.; Xiong, Z.; Li, J.; and Bai, Z. 2026. FinResearchBench II : A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality. arXiv preprint arXiv:2607.12252
2026 arXiv
-
[26]
Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; and Scialom, T. 2024. GAIA : a benchmark for General AI Assistants. In International Conference on Learning Representations
2024
-
[27]
Nie, Y.; Yan, B.; Guo, T.; Liu, H.; Wang, H.; He, W.; Zheng, B.; Wang, W.; Li, Q.; Sun, W.; Wang, Y.; and Tao, D. 2024. CFinBench : A Comprehensive C hinese Financial Benchmark for Large Language Models. In Advances in Neural Information Processing Systems
2024
-
[28]
OpenAI . 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774
2023 arXiv
-
[29]
L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training Language Models to Follow Instructions with Human Feedback. Advances in Neural Information Processing Systems, 35: 27730--27744
2022
-
[30]
P.; Aljubeh, M.; Thacker, P.; Fauconnet, L.; Kim, N
Patwardhan, T.; Dias, R.; Proehl, E.; Kim, G.; Wang, M.; Watkins, O.; Fishman, S. P.; Aljubeh, M.; Thacker, P.; Fauconnet, L.; Kim, N. S.; Chao, P.; Miserendino, S.; Chabot, G.; Li, D.; Sharman, M.; Barr, A.; Glaese, A.; and Tworek, J. 2025. GDPval : Evaluating AI Model Perfor...
2025 arXiv
-
[31]
D.; Ermon, S.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems
2023
-
[32]
L.; Stickland, A
Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; and Bowman, S. R. 2024. GPQA : A Graduate-Level G oogle-Proof Q & A Benchmark. In Conference on Language Modeling
2024
-
[33]
S.; Chawla, K.; Eidnani, D.; Shah, A.; Du, W.; Chava, S.; Raman, N.; Smiley, C.; Chen, J.; and Yang, D
Shah, R. S.; Chawla, K.; Eidnani, D.; Shah, A.; Du, W.; Chava, S.; Raman, N.; Smiley, C.; Chen, J.; and Yang, D. 2022. When FLUE Meets FLANG : Benchmarks and Large Pretrained Language Model for Financial Domain. In Proceedings of the 2022 Conference on Empirical Methods in Nat...
2022
-
[34]
F.; Qiu, X.; Whitehouse, C.; Alazraki, L.; Goel, S.; Barbieri, F.; Willi, T.; Mathur, A.; and Leontiadis, I
Shen, W. F.; Qiu, X.; Whitehouse, C.; Alazraki, L.; Goel, S.; Barbieri, F.; Willi, T.; Mathur, A.; and Leontiadis, I. 2026. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks. arXiv preprint arXiv:2602.05125
2026
-
[35]
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2022. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. arXiv preprint arXiv:2206.04615
2022 arXiv
-
[36]
Sun, Y.; Han, X.; Zhang, W.; Pang, Y.; Wang, T.; Cao, Y.; Huang, Y.; Duroiu, C.; Zhang, H.; Lin, J.; et al. 2026. Agents' Last Exam. arXiv preprint arXiv:2606.05405
2026
-
[37]
R.; Zhang, S.; Sun, Y.; and Wang, W
Wang, X.; Hu, Z.; Lu, P.; Zhu, Y.; Zhang, J.; Subramaniam, S.; Loomba, A. R.; Zhang, S.; Sun, Y.; and Wang, W. 2024 a . SciBench : Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models. In Proceedings of the 41st International Conference on Mac...
2024
-
[38]
Wang, Y.; Ma, X.; Zhang, G.; Ni, Y.; Chandra, A.; Guo, S.; Ren, W.; Arulraj, A.; He, X.; Jiang, Z.; Li, T.; Ku, M.; Wang, K.; Zhuang, A.; Fan, R.; Yue, X.; and Chen, W. 2024 b . MMLU-Pro : A More Robust and Challenging Multi-Task Language Understanding Benchmark. In Advances i...
2024
-
[39]
Wu, S.; Irsoy, O.; Lu, S.; Daber \'e ri, V.; Dredze, M.; Gehrmann, S.; Kambadur, P.; Rosenberg, D.; and Mann, G. 2023. BloombergGPT : A Large Language Model for Finance. arXiv preprint arXiv:2303.17564
2023 arXiv
-
[40]
Xie, L.; Huang, S.; Zhang, Z.; Zou, A.; Zhai, Y.; Ren, D.; Zhang, K.; Hu, H.; Liu, B.; Chen, H.; Liu, Z.; and Ding, B. 2025. Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling. arXiv preprint arXiv:2510.17314
2025
-
[41]
Xie, Q.; Han, W.; Chen, Z.; Xia, R.; Zhang, X.; He, Y.; Xiao, M.; Li, D.; Dai, Y.; Feng, D.; et al. 2024 a . FinBen : A Holistic Financial Benchmark for Large Language Models. In Advances in Neural Information Processing Systems
2024
-
[42]
Xie, Q.; Han, W.; Zhang, X.; Lai, Y.; Peng, M.; Lopez-Lira, A.; and Huang, J. 2023. PIXIU : A Large Language Model, Instruction Data and Evaluation Benchmark for Finance. In Advances in Neural Information Processing Systems
2023
-
[43]
J.; Cheng, Z.; Shin, D.; Lei, F.; et al
Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; et al. 2024 b . OSWorld : Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Advances in Neural Information Processing Systems
2024
-
[44]
F.; Li, Y.; Zhu, B.; Bisk, Y.; and Neubig, G
Xu, F. F.; Li, Y.; Zhu, B.; Bisk, Y.; and Neubig, G. 2024. TheAgentCompany : Benchmarking LLM Agents on Consequential Real World Tasks. arXiv preprint arXiv:2412.14161
2024 arXiv
-
[45]
Xu, Y.; Potje, G.; Shandilya, S.; Yuan, T.; de Oliveira Nunes, L.; Agarwal, R.; Asgari, S.; Atkinson, A.; K c man, E.; Lu, S.; Chandra, R.; and Chakraborty, T. 2026. SibylSense : Adaptive Rubric Learning via Memory Tuning and Adversarial Probing. arXiv preprint arXiv:2602.20751
2026
-
[46]
Yang, H.; Liu, X.-Y.; and Wang, C. D. 2023. FinGPT : Open-Source Financial Large Language Models. arXiv preprint arXiv:2306.06031
2023
-
[47]
Zhang, Q.; Zhou, J.; Wang, Y.; Lyu, F.; Ming, Y.; Xu, C.; Sun, Q.; Zheng, K.; Kang, P.; Liu, X.; and Ma, C. 2026. RubricBench : Aligning Model-Generated Rubrics with Human Standards. arXiv preprint arXiv:2603.01562
2026
-
[48]
P.; Zhang, H.; Gonzalez, J
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023. Judging LLM -as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems
2023
-
[49]
Zhong, J.; Zhang, H.; Southern, C.; Yang, J.; Wang, T.; Jung, K.; Zhang, S.; Yarats, D.; Ho, J.; and Ma, J. 2026. DRACO : A Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity. arXiv preprint arXiv:2602.11685
2026
-
[50]
F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; Alon, U.; and Neubig, G
Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; Alon, U.; and Neubig, G. 2024. WebArena : A Realistic Web Environment for Building Autonomous Agents. In International Conference on Learning Representations
2024
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.