Pith. sign in

Investigating adversarial trigger transfer in large language models.Transactions of the Association for Computational Linguistics, 13:953–979

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

On the Hardness of Junking LLMs

cs.LG · 2026-05-06 · unverdicted · novelty 7.0

Greedy random search recovers token sequences that elicit harmful response prefixes from LLMs without meaningful instructions, showing natural backdoors are present yet require more effort than semantic attacks.

citing papers explorer

Showing 1 of 1 citing paper.

  • On the Hardness of Junking LLMs cs.LG · 2026-05-06 · unverdicted · none · ref 35

    Greedy random search recovers token sequences that elicit harmful response prefixes from LLMs without meaningful instructions, showing natural backdoors are present yet require more effort than semantic attacks.