RedAct redacts agent traces to drop normalized skill transfer below the no-skill baseline on CapTraceBench while preserving audit evidence and adding detectable behavioral watermarks.
DOGe: Defensive output generation for LLM protection against knowledge distillation
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
A learned transformation matrix minimizes CMI in teacher logits to degrade distillation performance while preserving task accuracy.
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.
citing papers explorer
-
RedAct: Redacting Agent Capability Traces for Procedural Skill Protection
RedAct redacts agent traces to drop normalized skill transfer below the no-skill baseline on CapTraceBench while preserving audit evidence and adding detectable behavioral watermarks.
-
Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
A learned transformation matrix minimizes CMI in teacher logits to degrade distillation performance while preserving task accuracy.
-
Asking Back: Interaction-Layer Antidistillation Watermarks
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.