Pith. sign in

REVIEW 2 cited by

A Reproducible Extraction of Training Images from Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08694 v1 pith:FPVVW24J submitted 2023-05-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords trainingattackdiffusionextractiontheyimagesmodelregurgitate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, Carlini et al. demonstrated the widely used model Stable Diffusion can regurgitate real training samples, which is troublesome from a copyright perspective. In this work, we provide an efficient extraction attack on par with the recent attack, with several order of magnitudes less network evaluations. In the process, we expose a new phenomena, which we dub template verbatims, wherein a diffusion model will regurgitate a training sample largely in tact. Template verbatims are harder to detect as they require retrieval and masking to correctly label. Furthermore, they are still generated by newer systems, even those which de-duplicate their training set, and we give insight into why they still appear during generation. We extract training images from several state of the art systems, including Stable Diffusion 2.0, Deep Image Floyd, and finally Midjourney v4. We release code to verify our extraction attack, perform the attack, as well as all extracted prompts at \url{https://github.com/ryanwebster90/onestep-extraction}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models

    cs.CV 2026-02 conditional novelty 7.0 of 10

    A per-prompt cross-attention spike detector plus repulsive-attractive guidance (GUARD) substantially reduces verbatim and template memorization in Stable Diffusion at inference time.

  2. Finding DoRI: Discovery of Retained Images in Diffusion Models

    cs.CV 2025-07 conditional novelty 7.0 of 10

    Adversarially optimized text embeddings re-trigger supposedly removed memorized images in pruned diffusion models, showing memorization is distributed rather than local.

Pith tools