Pith. sign in

Best-of-venom: Attacking RLHF by injecting poisoned preference data

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Self-Consuming Generative Models with Adversarially Curated Data

cs.LG · 2025-05-14 · reject · novelty 5.0

Under adversarial data curation, a self-consuming generative model's alignment with user preferences is claimed to depend on the covariance between true and malicious reward functions; the proposed attack algorithms aim to make that covariance negative.

citing papers explorer

Showing 1 of 1 citing paper.

  • Self-Consuming Generative Models with Adversarially Curated Data cs.LG · 2025-05-14 · reject · none · ref 5

    Under adversarial data curation, a self-consuming generative model's alignment with user preferences is claimed to depend on the covariance between true and malicious reward functions; the proposed attack algorithms aim to make that covariance negative.