REVIEW 3 cited by
Contextual Moral Value Alignment Through Context-Based Aggregation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capability to consolidate multiple independently trained dialogue agents, each aligned with a distinct moral value, into a unified system that can adapt to and be aligned with multiple moral values is of paramount importance. In this paper, we propose a system that does contextual moral value alignment based on contextual aggregation. Here, aggregation is defined as the process of integrating a subset of LLM responses that are best suited to respond to a user input, taking into account features extracted from the user's input. The proposed system shows better results in term of alignment to human value compared to the state of the art.
Forward citations
Cited by 3 Pith papers
-
User identity conditions moral wrongness ratings in non-reasoning large language models
Implicitly conveying a user's professional role in multi-turn LLM conversations shifts moral wrongness ratings across ten common-morality rules in two non-reasoning models.
-
Position: Theory of Mind Benchmarks are Broken for Large Language Models
The paper proposes that LLM theory-of-mind evaluation should measure functional adaptation to partners, not just literal prediction of their behavior, and shows the two can diverge sharply in simple games.
-
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
Staggered asynchronous inference lets reinforcement learning agents with large, slow models act at every time step in realtime environments, at the cost of delay regret that grows with environment stochasticity.
Discussion (0). Continue with ORCID to comment.