DataPRM is an environment-aware generative process reward model that improves LLM data analysis agents by 7-11% on benchmarks via active verification and reflection-aware ternary rewards.
Tablebench: a comprehensive and complex benchmark for table question answering
2 Pith papers cite this work, alongside 20 external citations. Polarity classification is still indexing.
2
Pith papers citing it
20
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 2years
2026 2roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
DataPRM is an environment-aware generative process reward model that improves LLM data analysis agents by 7-11% on benchmarks via active verification and reflection-aware ternary rewards.
- InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis