Pith. sign in

Paper Citation Record · LEDGER

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

As of 13 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2412.05538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05538 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:16.429938Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02c3b149-e711-4e9f-b4f9-54b4e28b75d7 · outbound

This paper cites GPT-4 Technical Report.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.947995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.947995Z digest=sha256:71b703650e41f067fb52c13275feff657e96afa7d498bb753b46c849016358ed

Observation abd15361-39cc-430a-9b77-130ad3f1c618 · outbound

This paper cites Elijah: Eliminating backdoors injected in diffusion models via distribution shift.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Elijah: Eliminating backdoors injected in diffusion models via distribution shift

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.897375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:15.953987Z digest=sha256:2da4f6f1c8033230898a92e3124234b0ba1d8f27366c93ba6a356dcd42f642bd

Observation 0e15bbd9-5082-4a9d-9b0a-e90da71f9b81 · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defense-prefix for pre- venting typographic attacks on clip

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.880241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:15.964529Z digest=sha256:440cf10641e85ab4ed1354f1a632c751ebc68870c75c0ff64c7c1bd845beca8a

Observation 8652d655-23e9-48d5-a5b0-0e03519bbee9 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models In- structpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.863071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:15.977376Z digest=sha256:15b22245bd625e742f0478265129589d538b32f96e803260252e28b5e3bbbecc

Observation ec4da44d-6208-4b9a-be51-1f4b2a33fdf3 · outbound

This paper cites Controllable generation with text-to-image diffusion models: A survey.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Controllable generation with text-to-image diffusion models: A survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.986038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.986038Z digest=sha256:6e7f62848da0363ef03165d9d18d91eb2d232ecd5091a8ed77bc0d2cc5685733

Observation a36bde21-ca20-4645-84da-095867d5f8f4 · outbound

This paper cites Trojdiff: Trojan at- tacks on diffusion models with diverse targets.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Trojdiff: Trojan at- tacks on diffusion models with diverse targets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.992611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.992611Z digest=sha256:09ffe9d38bfa48a5ed9ec79ca25abbc8035a8b8f02b21de2725bc1b1d891bf69

Observation 6ec6bbc9-f0e4-4517-85ff-7d68f2b3b65f · outbound

This paper cites Rbformer: improve adversarial robustness of trans- former by robust bias.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Rbformer: improve adversarial robustness of trans- former by robust bias

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.830272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.006151Z digest=sha256:b14f401b494c8320a8ffbdc073f000243c016efc6ed7631000f47c1bb4c821cf

Observation efe0e062-367b-4ca2-be35-99b0f8466608 · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.813758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.016806Z digest=sha256:74037a2d5c66ea3d78fa9eb99daad3c49cde04f4bab67b7bf8f6267e71cb2b0b

Observation a3a5a015-2408-493d-870a-f6109ce08919 · outbound

This paper cites Villan- diffusion: A unified backdoor attack framework for diffu- sion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Villan- diffusion: A unified backdoor attack framework for diffu- sion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.795940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.023213Z digest=sha256:9a4653da7be213f5bed5a8542b7e90a3075756e284bb504f881321f8e727ed9c

Observation 16e1dc5d-bbac-4df9-99be-1d86e7b699c6 · outbound

This paper cites Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.779803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.029599Z digest=sha256:c30ede5560a9f1becb2d3b31ba3ed5f6e12c8f3106dee27bb4734678a2c8cac2

Observation 9d9acdc3-f871-4b96-85d0-ce93883dfc4d · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.035929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.035929Z digest=sha256:2507c3223f10630fa65618ae0e7e09e0c4d4cc361de3fe6e1abdea2682b98da5

Observation e05161cf-4914-4abd-8678-1f1d68a22132 · outbound

This paper cites Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.752409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.042720Z digest=sha256:fdc3c3ad9d52a51ddafe967e0d479bb9c3b4b5a25765bf4597f0a2b78306ccfd

Observation 2917006e-fb9a-409d-99d4-e0e4a1009d8c · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.736563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.047918Z digest=sha256:8913d99259d7327e82dfc97334d145fd2519f571a2c3af0a5da5c4babcf06d48

Observation 0c41b359-e9af-422b-8fca-b3ff5eae6d77 · outbound

This paper cites HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.054907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.054907Z digest=sha256:19237cd7292c2af8cfe327e0eafc43cc4144b8d907a3a18f08578fe443df450d

Observation 4be1fb8b-9bcd-4c1e-ad14-8c4871f8c365 · outbound

This paper cites Generative adversarial networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative adversarial networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.063958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.063958Z digest=sha256:5447edc87ca7e3964cf42f1cf1c4ef147470a39015f3c75be02d856d3a2374c1

Observation f7930752-df99-4fad-b256-ba8d862dca48 · outbound

This paper cites A Survey on Responsible Generative AI: What to Generate and What Not.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A Survey on Responsible Generative AI: What to Generate and What Not

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.069137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.069137Z digest=sha256:8d2483092820e528907cb5a1920e215fccfcae52de64c8a52af43d92fae831ff

Observation 3d8ba173-6825-4b65-b867-427842bf5224 · outbound

This paper cites Detoxify.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Detoxify

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.706588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.074800Z digest=sha256:8cf573d577b1208ed25daed7aa7063a2cd2b9319682ab7cb8f4253b05640affa

Observation eebffb7d-1594-474f-a07f-2d6b85a74a38 · outbound

This paper cites Defending against Backdoor Attack on Deep Neural Networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defending against Backdoor Attack on Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.079816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.079816Z digest=sha256:8494a13b4140f267169a6eed4aaaf3d3d663fb3f4eb82959c845b308d14e9020

Observation 5bd09f48-55c7-4aa7-a65b-0af1229f569b · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.690423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.086090Z digest=sha256:3e877b7043543ba5c8349384dbb7f6019cc95334d5d4aafda2c8a20283447a26

Observation 4ee3f535-3742-4566-bde3-bb6122d1ba35 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.092015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.092015Z digest=sha256:5836e55b6ef7bafbd397f30533e1c7c82f05416ef311e5b9eaa90eca4308e5e4

Observation 556bce85-0274-4106-8ef4-9007da1ab1df · outbound

This paper cites All but one: Surgical concept erasing with model preservation in text-to- image diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models All but one: Surgical concept erasing with model preservation in text-to- image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.662419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.098181Z digest=sha256:255a7d78e1fac95f29fd906a8bd8ab9c432ff6a11e3ccacfef4c0ad694d7a87b

Observation 0481bf68-d92a-4e53-84db-4cc62e665811 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.103292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.103292Z digest=sha256:e2cc510265490ad7ec6d6facb1a191dcdca3c40ea3a99bb0df69f4c90e74e7f2

Observation f5ad5d92-219f-4c07-8b82-ece074facc03 · outbound

This paper cites Progressive Growing of GANs for Improved Quality, Stability, and Variation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Progressive Growing of GANs for Improved Quality, Stability, and Variation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.109298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.109298Z digest=sha256:e6e61e33fc6d9ed906902bd8a9c72fe0e08621a6a5e5b105b182780f1b6edc13

Observation 70d9bd6f-303d-43ab-bc83-c483a9a15efe · outbound

This paper cites Auto-Encoding Variational Bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Auto-Encoding Variational Bayes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.116002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.116002Z digest=sha256:453f6a6da247f9aa235dab5f07cf4de821c84b5b6d174e6b7183e73423fced60

Observation 92f5ceb6-770f-4805-883a-6b85eeaebe98 · outbound

This paper cites Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.644039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.123141Z digest=sha256:5229dcf45d8d603c63c932619d8073a7c2167aa83da2a4e52a7ec48f25aadc8f

Observation 3dd086a4-7e79-471b-ab56-56afba6db1ab · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.137150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.137150Z digest=sha256:b4e3b8ee56f56a0cf4e9c6aeb62acdcf8690bf6282ea8d04daa3c4c59be277fc

Observation 71f0f88a-8fbf-4865-a174-bfb8436f0712 · outbound

This paper cites Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.626817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.144757Z digest=sha256:720f39e69ba5c9022a94b48c59497c6e8e6ca36c7f6141167f48332c566fab74

Observation e1d08a69-acfa-4ea1-9b94-ef572a375614 · outbound

This paper cites Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.152540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.152540Z digest=sha256:15c974da9bfe784aaa069f6ca7e0a75c748649d2563b87a400dc9e50b7bedfce

Observation 8bc0d17b-7152-4eb6-af9e-3b002e5dac2c · outbound

This paper cites Which model generated this image? a model- agnostic approach for origin attribution.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Which model generated this image? a model- agnostic approach for origin attribution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.607534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.158775Z digest=sha256:fc18716239c4cdc517ec6e0e580a259d3954ddb0be6b09f2e4c0670ceb49bce6

Observation 2ccc9a99-8687-48b0-9a78-b768756d6e54 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Improved Baselines with Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.165369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.165369Z digest=sha256:1cf2229e913e308916ca9459ee477b97ca5ba6e453c5666da32a6d7003d4b041

Observation 467148d4-0fec-4a41-84a1-ba0a941534a0 · outbound

This paper cites Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.171091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.171091Z digest=sha256:79af2063a92e7daa7033e2694a40758211e7f23222f6cec488ca802b6a66526a

Observation fcd2f6c9-6a9c-468f-8eb5-57078ab22787 · outbound

This paper cites Latent guard: a safety frame- work for text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Latent guard: a safety frame- work for text-to-image generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.590711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.182742Z digest=sha256:58f35848a2ee5bead6f93047ca2593cf553d65d7824a54c9969240398c6a28f1

Observation c05fc1f7-6df0-48dc-80aa-8acbf6ed9b58 · outbound

This paper cites Multimodal prag- matic jailbreak on text-to-image models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Multimodal prag- matic jailbreak on text-to-image models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.573013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.188138Z digest=sha256:340943edecab953ac02eb6ff4052a7a1f6c4abaf91a7312f77caa2e26a0a65b2

Observation 16392452-1f48-41ed-a583-505200cdc8f9 · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.195700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.195700Z digest=sha256:ce6d451356bdcc8527e459b2b0e1b81b06778e233a57e2f00d302a08a9103ddc

Observation 2435339a-a0c1-480a-a9b8-1a3770a8d3de · outbound

This paper cites Large-scale celebfaces attributes (celeba) dataset.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Large-scale celebfaces attributes (celeba) dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.554589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.202074Z digest=sha256:f0b8a77f845c65cbf805e300922622ccd18cf1efbcadf093199bcf6b01606c6e

Observation a06b4bea-ebaa-42a4-a776-85fd4c93f603 · outbound

This paper cites Information constraints on auto-encoding variational bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Information constraints on auto-encoding variational bayes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.537226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.210370Z digest=sha256:c261e06662ce810afc762fa21a21e9f425f3e9a777edd23b9588f9aacf9f0593

Observation a0275460-a58a-46f9-ac1a-58dd1529dece · outbound

This paper cites An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.518701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.219957Z digest=sha256:7682a070ab89a800144ed3324a88b99c1e7913502a45f2059ad3ff48af8d9028

Observation 198e468f-803d-4cee-a8a4-9e0d4bd57a69 · outbound

This paper cites Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.226826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.226826Z digest=sha256:6501a3e2f9f11131e62cf6297f0267ad5a02d99e643cb526f3cf66bab99a12f6

Observation e2edfa60-77c8-47ce-934a-0d9a200ea5dd · outbound

This paper cites A holistic approach to undesired content detection in the real world.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A holistic approach to undesired content detection in the real world

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.500221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.233197Z digest=sha256:2c21204bd3755df637fe84b325184fb4b16956eb294c2b017c0a8f7f5826e958

Observation 7626b988-ae74-4559-bbe6-3495c130d857 · outbound

This paper cites Dreamguider: Improved Training free Diffusion-based Conditional Generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Dreamguider: Improved Training free Diffusion-based Conditional Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.247131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.247131Z digest=sha256:979dda4f0f08dc26b5cb5a8b05c64d798b3aaca0cadbec754577f2133f1a3f02

Observation 1dc9f6e0-757f-4237-bf2d-85d1d5b31f90 · outbound

This paper cites At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.484104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.262943Z digest=sha256:c19d4d1d3fb1a97dcfb9ba7f7ee3f8f08efe05d4484a542bf7c55787a7a0e00b

Observation d700a64a-157c-41d8-aa2a-db3adfbc0425 · outbound

This paper cites Contrastive denoising score for text-guided latent diffusion image editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Contrastive denoising score for text-guided latent diffusion image editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.467275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.272875Z digest=sha256:a714b692d811f01d224d7e2f8b2976c1c056eea5120ef0d0d13b715dae7c66fc

Observation 11ea48a6-e231-4b9b-a4a2-dc2a461d0fb1 · outbound

This paper cites White-box Membership Inference Attacks against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models White-box Membership Inference Attacks against Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.279099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.279099Z digest=sha256:ebbb1992656287503fdf9d690b5bcc877804c1d9e14d7c6920804b170e022a96

Observation 580454fc-c7ef-4465-8431-183ef9cace1e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.285778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.285778Z digest=sha256:8a649bed8e871eb85b91074743790edba204fd58e82ca0f617a528f91d947358

Observation d24ded01-08f6-4477-91f8-1bef07cad98f · outbound

This paper cites Safe-clip: Removing nsfw concepts from vision-and-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Safe-clip: Removing nsfw concepts from vision-and-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.450071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.294864Z digest=sha256:51aad7f945eb1f9c9500402038bcb643a267c43216fc68fbffc14008a8d35f1f

Observation f3fdfe8d-e283-48ab-8b0c-c59edf536f87 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.431641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.302095Z digest=sha256:f037c552ac7003dac182e4f46c4c3e5d9dd9bf43bc1af632558de765205a4b1f

Observation 79bb20c5-c338-4f24-9e9c-1e481733286c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervi- sion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.311262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.311262Z digest=sha256:55f82390aa36a218ca061de121823b85a9d96b49d7f46dfa09e1a2bc0ad6e772

Observation e85317c9-32b4-4461-8a2b-435af4216843 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.317220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.317220Z digest=sha256:17b2923865468a7ff751eb3d0b27d11adfa9f1445d8d1dd25ac29df390bb59d7

Observation e1901301-6bd3-469d-9c37-12a2c11fd13e · outbound

This paper cites Red-Teaming the Stable Diffusion Safety Filter.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Red-Teaming the Stable Diffusion Safety Filter

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.323922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.323922Z digest=sha256:66810a6bbfc3a1cbc12d95089f8f443347eea72022e7dbcdaff2d2780221df36

Observation c466172a-b985-4d33-9e9a-f83673b770fc · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gener- ating diverse high-fidelity images with vq-vae-2

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.330422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.330422Z digest=sha256:67e7a22037a2f277b574d6d2bfb694fc2ca13fdb1e5d84375536cb500dda2348

Observation 8750fd41-ada9-4845-9cb1-f114b81e8ce3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.335821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.335821Z digest=sha256:e3b60ab000a87ea8673a17d67fc276006d32be1a398441da5328d27383148556

Observation d081260a-f4b2-496e-baf4-607ff318d5d7 · outbound

This paper cites Raising the Cost of Malicious AI-Powered Image Editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Raising the Cost of Malicious AI-Powered Image Editing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.342958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.342958Z digest=sha256:9d16744438d0d485e9470601da6c5322a30ab5b286cbe8a32fbb307c8bc359b2

Observation 4c45e4a7-1f94-466f-ab87-acc99dd4d8b5 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Deep unsupervised learning using nonequilibrium thermodynamics

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.349013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.349013Z digest=sha256:f3a02fbd5ba4432238c59def0f35d2c3f384e3b4b4a83759601695e80f8341f9

Observation f961a41a-8ff6-474c-966e-e71c5bb3d1e8 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Emu: Generative pretraining in multimodality

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.372299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.354839Z digest=sha256:ab21814d9228361aaa570401129828b50cee8e6083e3c02f5ae0289c00192b1c

Observation e802768e-17cc-4a74-9438-ca6a84111fd6 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative multimodal mod- els are in-context learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.344759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.361162Z digest=sha256:c7e75e9516a054bfc820662489294df37fbdf9950896bc8e64f716a338695003

Observation b939e121-7ed5-4950-9f0e-5c6f072d5143 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.368896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.368896Z digest=sha256:870c7973d139451409dd34ec72c7621405f4c62753e9528674f571a34dc75ad7

Observation 8ba5fa7b-4918-401d-8eb1-8e8bb97917b7 · outbound

This paper cites Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.325116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.379570Z digest=sha256:b96c6bc45a0fbcb731c62f84376e98a266ed64c82fa73eba193c22fbc97d230e

Observation 31f190b1-04b0-4372-8834-c35e970dbfb7 · outbound

This paper cites Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.386103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.386103Z digest=sha256:5b966156cafcf81b9f4f645458e30e6bfc8d1eccd4b6e36fab73f1d6c7e8c3c4

Observation e8f7d067-0e3b-4097-8732-9a4e1e7fa217 · outbound

This paper cites Sneakyprompt: Jailbreaking text-to-image generative models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Sneakyprompt: Jailbreaking text-to-image generative models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.307571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.396910Z digest=sha256:46d37a99872af72592a04b0da2a050d6c73baef2597b98edc351f85e5414a654

Observation cc4ba249-05ba-46ef-9cf8-ef0e9ef2fa5b · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.402465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.402465Z digest=sha256:5fd3e075eb4749cdaaea7cd0de210056251d7e8efc0ed13caa9997bf0970e367

Observation 8c07f2aa-0b95-473c-8239-696ad4c3029f · outbound

This paper cites Inversion-based style transfer with diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Inversion-based style transfer with diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.291166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.408500Z digest=sha256:5f100c115a9f7e9be47831beb399658fd77e09e8c9c5c7f3b3be626ae561d46e

Observation d2ec859b-e882-477c-a0fb-e1c0ca8066d8 · outbound

This paper cites Defensive unlearning with adversarial training for robust concept erasure in diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defensive unlearning with adversarial training for robust concept erasure in diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.274376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:41:16.422936Z digest=sha256:b23b7dff7e8ad46be4fbd6afc6aeb62a3ec9d5cc9c5eb4c08db1062871888441

Observation b5acb7ce-93af-49d7-8d8b-8ddd9a55b355 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:41:16.429938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.429938Z digest=sha256:7ac20c4420eff6695ed2c9210ca16ff1236de05a6b3dcaefa10f988f6cde599c

Pith citing papers

No inbound Pith citation observations are available.