Decentralized poly-time algorithms achieve (1-1/e)-approximate assistance regret Õ(T^{3/4}) (or Õ(√T) with shared randomness) for online assistance games, and better approximation is intractable.
Cooperative inverse reinforcement learning
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
HyRe personalizes reward models at test time by reweighting an ensemble of heads trained on aggregate preferences, using few target examples to outperform uniform averaging and prior methods on RewardBench and 32 tasks.
TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.
citing papers explorer
-
Provably Optimal Learning Algorithms for Assistance Games
Decentralized poly-time algorithms achieve (1-1/e)-approximate assistance regret Õ(T^{3/4}) (or Õ(√T) with shared randomness) for online assistance games, and better approximation is intractable.
-
Test-Time Alignment via Hypothesis Reweighting
HyRe personalizes reward models at test time by reweighting an ensemble of heads trained on aggregate preferences, using few target examples to outperform uniform averaging and prior methods on RewardBench and 32 tasks.
-
TrustLLM: Trustworthiness in Large Language Models
TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.