MORA breaks the safety-helpfulness ceiling in LLMs by pre-sampling single-reward prompts and rewriting them to incorporate multi-dimensional intents, delivering 5-12.4% gains in sequential alignment and 4.6% overall improvement in simultaneous alignment.
Towards friendly ai: A comprehensive review and new perspectives on human-ai alignment
3 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
SIGMA applies post-hoc XAI saliency maps to define reusable sparse masks for magnitude-bounded perturbations on self-supervised speech features, evaluated on IEMOCAP and TESS for competitive attack success with explanation consistency trade-offs.
In real human subjects, AI transparency impacts imperfectly cooperative interactions far more than personality traits, unlike simulations where both are comparably influential.
citing papers explorer
-
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
MORA breaks the safety-helpfulness ceiling in LLMs by pre-sampling single-reward prompts and rewriting them to incorporate multi-dimensional intents, delivering 5-12.4% gains in sequential alignment and 4.6% overall improvement in simultaneous alignment.
-
SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition
SIGMA applies post-hoc XAI saliency maps to define reusable sparse masks for magnitude-bounded perturbations on self-supervised speech features, evaluated on IEMOCAP and TESS for competitive attack success with explanation consistency trade-offs.
-
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
In real human subjects, AI transparency impacts imperfectly cooperative interactions far more than personality traits, unlike simulations where both are comparably influential.