DataPRM is an environment-aware generative process reward model that improves LLM data analysis agents by 7-11% on benchmarks via active verification and reflection-aware ternary rewards.
Pan, Guilin Qi, Haofen Wang, and Huajun Chen
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
Ctx2Skill automatically produces natural-language skill files via self-play between a challenger, a reasoner, and a judge, improving LLM performance on context-learning tasks.
citing papers explorer
-
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
DataPRM is an environment-aware generative process reward model that improves LLM data analysis agents by 7-11% on benchmarks via active verification and reflection-aware ternary rewards.
-
From Context to Skills: Can Language Models Learn from Context Skillfully?
Ctx2Skill automatically produces natural-language skill files via self-play between a challenger, a reasoner, and a judge, improving LLM performance on context-learning tasks.