A technique identifies minimal convergence-divergence points in LLM transformer blocks and calibrates residual-stream directions to achieve targeted ethical-framework control at inference time.
Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
A framework maps Reddit moral judgments to weighted logical predicates and optimizes for consistency via Weighted MaxSAT, producing verdicts that differ from majority labels 62% of the time while matching human evaluators 86% of the time.
citing papers explorer
-
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
A technique identifies minimal convergence-divergence points in LLM transformer blocks and calibrates residual-stream directions to achieve targeted ethical-framework control at inference time.
-
Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework
A framework maps Reddit moral judgments to weighted logical predicates and optimizes for consistency via Weighted MaxSAT, producing verdicts that differ from majority labels 62% of the time while matching human evaluators 86% of the time.