Pith. sign in

Poisoning attacks on llms require a near-constant number of poison samples

13 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.

13 Pith papers citing it
5 external citations · external index

citation-role summary

background 3

citation-polarity summary

years

2026 13

roles

background 3

polarities

background 2 support 1

representative citing papers

When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks

cs.LG · 2026-05-21 · unverdicted · novelty 8.0

In the proportional high-dimensional regime, stronger backdoor training triggers improve clean accuracy and make attack success non-monotonic for regularized GLMs on Gaussian mixtures, with closed-form proofs for squared loss and fixed-point extensions to convex losses.

Narrow Secret Loyalty Dodges Black-Box Audits

cs.CR · 2026-05-07 · unverdicted · novelty 7.0 · 3 refs

First model organisms of narrow secret loyalties in LLMs evade black-box audits without principal knowledge and persist even at low poison fractions in training data.

The Power of Backdoor Absorption in Community Training

cs.CR · 2026-07-07 · conditional · novelty 6.0

A DTMC of injection versus absorption shows that natural decay plus 10% lazy verification and weight penalties forces any bounded adversary's success probability to zero with no utility loss.

citing papers explorer

Showing 13 of 13 citing papers.