Introduces MOOD benchmark for OOD LLM alignment failures and shows guard models plus Mahalanobis and perplexity OOD detectors improve recall from 39% to 45% with positive scaling.
Title resolution pending
4 Pith papers cite this work, alongside 7 external citations. Polarity classification is still indexing.
representative citing papers
CoAct synergistically merges self-rewarding and active learning via self-consistency to select reliable AI labels and oracle-needed samples, delivering 8-13% gains on GSM8K, MATH, and WebInstruct.
Synthetic measurements show that gold-standard performance degrades according to distinct functional forms when optimizing proxy reward models via RL or best-of-n, with coefficients scaling smoothly by reward model parameter count.
Bayesian visual transformers with ensemble and sampling methods achieve a 7.4 percentage point gain on weighted F-beta score for affordance instance segmentation on the IIT-Aff dataset while providing calibrated epistemic and aleatoric uncertainty maps.
citing papers explorer
-
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Introduces MOOD benchmark for OOD LLM alignment failures and shows guard models plus Mahalanobis and perplexity OOD detectors improve recall from 39% to 45% with positive scaling.
-
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
CoAct synergistically merges self-rewarding and active learning via self-consistency to select reliable AI labels and oracle-needed samples, delivering 8-13% gains on GSM8K, MATH, and WebInstruct.
-
Scaling Laws for Reward Model Overoptimization
Synthetic measurements show that gold-standard performance degrades according to distinct functional forms when optimizing proxy reward models via RL or best-of-n, with coefficients scaling smoothly by reward model parameter count.
-
Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers
Bayesian visual transformers with ensemble and sampling methods achieve a 7.4 percentage point gain on weighted F-beta score for affordance instance segmentation on the IIT-Aff dataset while providing calibrated epistemic and aleatoric uncertainty maps.