REFORM uses reward-guided controlled decoding to generate preference-class-consistent responses that the reward model mis-scores, then retrains the reward model on these failure modes to improve robustness.
Zero-shot llm-guided counterfactual generation: A case study on nlp model evaluation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
representative citing papers
Comp-MCTS is a training-free tree-search framework that maximizes yield of unique oracle-validated counterfactuals under fixed LLM budgets on tabular datasets, outperforming LATS-style baselines in experiments.
citing papers explorer
-
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling
REFORM uses reward-guided controlled decoding to generate preference-class-consistent responses that the reward model mis-scores, then retrains the reward model on these failure modes to improve robustness.
-
Agentic Search for Counterfactual Recourse under Fixed LLM Budgets
Comp-MCTS is a training-free tree-search framework that maximizes yield of unique oracle-validated counterfactuals under fixed LLM budgets on tabular datasets, outperforming LATS-style baselines in experiments.