ReLU networks' division of input space into convex polytopal regions permits direct extraction of causal rules that exactly match the original network's linear behavior in each region.
Aligning Robot and Human Representations
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
The paper formalizes three types of pluralistic AI models and three benchmark classes, arguing that current alignment techniques may reduce rather than increase distributional pluralism.
A POMDP tree search technique estimates user reward weights from action discrepancies to reconcile and explain differences between algorithm and human decisions.
citing papers explorer
-
Causal Explanations from the Geometric Properties of ReLU Neural Networks
ReLU networks' division of input space into convex polytopal regions permits direct extraction of causal rules that exactly match the original network's linear behavior in each region.
-
A Roadmap to Pluralistic Alignment
The paper formalizes three types of pluralistic AI models and three benchmark classes, arguing that current alignment techniques may reduce rather than increase distributional pluralism.
-
Explanation through Reward Model Reconciliation using POMDP Tree Search
A POMDP tree search technique estimates user reward weights from action discrepancies to reconcile and explain differences between algorithm and human decisions.