Pith. sign in

arXiv preprint arXiv:2506.22806 , year=

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

fields

cs.CV 4 cs.AI 1

years

2026 4 2025 1

verdicts

UNVERDICTED 5

representative citing papers

Concept Removal for Frontier Image Generative Models

cs.CV · 2026-06-24 · unverdicted · novelty 6.0

A transcoder-based in-place replacement of the bottleneck layer enables selective concept removal in modern diffusion and autoregressive image models without degrading output quality.

SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training

cs.CV · 2026-05-18 · unverdicted · novelty 6.0

SafeDiffusion-R1 uses online GRPO with CLIP embedding steering to cut inappropriate content from 48.9% to 18.07% and nudity detections from 646 to 15 in diffusion models while raising GenEval scores from 42.08% to 47.83% and generalizing across seven harm categories without supervised pairs or extra

Orthogonal Concept Erasure for Diffusion Models

cs.AI · 2026-05-27 · unverdicted · novelty 5.0

OCE reformulates editing-based concept erasure in diffusion models as multiplicative orthogonal transformations on neuron parameters to decouple concept direction from magnitude and angular geometry.

BARRIER: Bounded Activation Regions for Robust Information Erasure

cs.CV · 2026-05-15 · unverdicted · novelty 5.0

BARRIER applies interval arithmetic to SVD-based activation projections to create bounded forget regions that enable aggressive unlearning while providing formal protection for retain distributions via tail bounds on functional drift.

citing papers explorer

Showing 5 of 5 citing papers.

  • Concept Removal for Frontier Image Generative Models cs.CV · 2026-06-24 · unverdicted · none · ref 64

    A transcoder-based in-place replacement of the bottleneck layer enables selective concept removal in modern diffusion and autoregressive image models without degrading output quality.

  • SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training cs.CV · 2026-05-18 · unverdicted · none · ref 78

    SafeDiffusion-R1 uses online GRPO with CLIP embedding steering to cut inappropriate content from 48.9% to 18.07% and nudity detections from 646 to 15 in diffusion models while raising GenEval scores from 42.08% to 47.83% and generalizing across seven harm categories without supervised pairs or extra

  • Orthogonal Concept Erasure for Diffusion Models cs.AI · 2026-05-27 · unverdicted · none · ref 2

    OCE reformulates editing-based concept erasure in diffusion models as multiplicative orthogonal transformations on neuron parameters to decouple concept direction from magnitude and angular geometry.

  • BARRIER: Bounded Activation Regions for Robust Information Erasure cs.CV · 2026-05-15 · unverdicted · none · ref 35

    BARRIER applies interval arithmetic to SVD-based activation projections to create bounded forget regions that enable aggressive unlearning while providing formal protection for retain distributions via tail bounds on functional drift.

  • GrOCE:Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models cs.CV · 2025-11-17 · unverdicted · none · ref 17

    GrOCE uses dynamic semantic graphs for online, training-free erasure of target concepts from diffusion model prompts via cluster identification and selective severing.