Pith. sign in

Paper Citation Record · LEDGER

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

As of 13 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2412.05538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05538 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:16.429938Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02c3b149-e711-4e9f-b4f9-54b4e28b75d7 · outbound

This paper cites GPT-4 Technical Report.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.947995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.947995Z digest=sha256:71b703650e41f067fb52c13275feff657e96afa7d498bb753b46c849016358ed

Observation abd15361-39cc-430a-9b77-130ad3f1c618 · outbound

This paper cites Elijah: Eliminating backdoors injected in diffusion models via distribution shift.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Elijah: Eliminating backdoors injected in diffusion models via distribution shift

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.897375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:15.953987Z digest=sha256:2b69c38eabe05958c4b98078e80665932402da52d0997dab2513fb52472d521d

Observation 0e15bbd9-5082-4a9d-9b0a-e90da71f9b81 · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defense-prefix for pre- venting typographic attacks on clip

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.880241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:15.964529Z digest=sha256:cca3bb9ea0226b000e9e3b74615e3df57a80af39fdbae866f1c7dd4508fd69a5

Observation 8652d655-23e9-48d5-a5b0-0e03519bbee9 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models In- structpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.863071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:15.977376Z digest=sha256:d384869d526e9bcbc77176ec10837b6255df00ad0aa8e1e4c5859ce2954390af

Observation ec4da44d-6208-4b9a-be51-1f4b2a33fdf3 · outbound

This paper cites Controllable generation with text-to-image diffusion models: A survey.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Controllable generation with text-to-image diffusion models: A survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.986038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.986038Z digest=sha256:6e7f62848da0363ef03165d9d18d91eb2d232ecd5091a8ed77bc0d2cc5685733

Observation a36bde21-ca20-4645-84da-095867d5f8f4 · outbound

This paper cites Trojdiff: Trojan at- tacks on diffusion models with diverse targets.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Trojdiff: Trojan at- tacks on diffusion models with diverse targets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.992611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.992611Z digest=sha256:09ffe9d38bfa48a5ed9ec79ca25abbc8035a8b8f02b21de2725bc1b1d891bf69

Observation 6ec6bbc9-f0e4-4517-85ff-7d68f2b3b65f · outbound

This paper cites Rbformer: improve adversarial robustness of trans- former by robust bias.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Rbformer: improve adversarial robustness of trans- former by robust bias

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.830272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.006151Z digest=sha256:e80160c1a051148e5f3a1f3f53736d5c9e172b6c13ec1904637d96d8db1fc8a7

Observation efe0e062-367b-4ca2-be35-99b0f8466608 · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.813758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.016806Z digest=sha256:225d7dd919efc7ee81fa40025b58b9efb8f1151ee847ecac30e7f7f1ce51459e

Observation a3a5a015-2408-493d-870a-f6109ce08919 · outbound

This paper cites Villan- diffusion: A unified backdoor attack framework for diffu- sion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Villan- diffusion: A unified backdoor attack framework for diffu- sion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.795940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.023213Z digest=sha256:4d366e8ce678ed5d2581ba6d512e5e6ce0aca400c61558f3779d1b16e948da94

Observation 16e1dc5d-bbac-4df9-99be-1d86e7b699c6 · outbound

This paper cites Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.779803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.029599Z digest=sha256:4bed8da2bfeb112c5b47ff7cca3101d015972dea11a2f5098509b2b097c98171

Observation 9d9acdc3-f871-4b96-85d0-ce93883dfc4d · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.035929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.035929Z digest=sha256:2507c3223f10630fa65618ae0e7e09e0c4d4cc361de3fe6e1abdea2682b98da5

Observation e05161cf-4914-4abd-8678-1f1d68a22132 · outbound

This paper cites Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.752409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.042720Z digest=sha256:1307cdd91f27752375f0ac0e1c2603e341624db841facb6d039155e5801ba53a

Observation 2917006e-fb9a-409d-99d4-e0e4a1009d8c · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.736563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.047918Z digest=sha256:277bd5c15122e2eb607a55824d707437d3ef68455a8ad104f1ed0e8a8b8296b3

Observation 0c41b359-e9af-422b-8fca-b3ff5eae6d77 · outbound

This paper cites HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.054907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.054907Z digest=sha256:19237cd7292c2af8cfe327e0eafc43cc4144b8d907a3a18f08578fe443df450d

Observation 4be1fb8b-9bcd-4c1e-ad14-8c4871f8c365 · outbound

This paper cites Generative adversarial networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative adversarial networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.063958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.063958Z digest=sha256:5447edc87ca7e3964cf42f1cf1c4ef147470a39015f3c75be02d856d3a2374c1

Observation f7930752-df99-4fad-b256-ba8d862dca48 · outbound

This paper cites A Survey on Responsible Generative AI: What to Generate and What Not.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A Survey on Responsible Generative AI: What to Generate and What Not

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.069137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.069137Z digest=sha256:8d2483092820e528907cb5a1920e215fccfcae52de64c8a52af43d92fae831ff

Observation 3d8ba173-6825-4b65-b867-427842bf5224 · outbound

This paper cites Detoxify.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Detoxify

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.706588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.074800Z digest=sha256:560d0ac0903bf4335c93cc27625310f50008d782286b1784e46998079509edfc

Observation eebffb7d-1594-474f-a07f-2d6b85a74a38 · outbound

This paper cites Defending against Backdoor Attack on Deep Neural Networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defending against Backdoor Attack on Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.079816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.079816Z digest=sha256:8494a13b4140f267169a6eed4aaaf3d3d663fb3f4eb82959c845b308d14e9020

Observation 5bd09f48-55c7-4aa7-a65b-0af1229f569b · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.690423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.086090Z digest=sha256:3cad92016195e9fff517ad582857361695a415d408bc6db886293b42f4ee9d4a

Observation 4ee3f535-3742-4566-bde3-bb6122d1ba35 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.092015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.092015Z digest=sha256:5836e55b6ef7bafbd397f30533e1c7c82f05416ef311e5b9eaa90eca4308e5e4

Observation 556bce85-0274-4106-8ef4-9007da1ab1df · outbound

This paper cites All but one: Surgical concept erasing with model preservation in text-to- image diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models All but one: Surgical concept erasing with model preservation in text-to- image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.662419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.098181Z digest=sha256:e66b0f2baf23460d8c109d4f6076bc8d89b3ae7f49ed1f460167f5432e1ba1d8

Observation 0481bf68-d92a-4e53-84db-4cc62e665811 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.103292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.103292Z digest=sha256:e2cc510265490ad7ec6d6facb1a191dcdca3c40ea3a99bb0df69f4c90e74e7f2

Observation f5ad5d92-219f-4c07-8b82-ece074facc03 · outbound

This paper cites Progressive Growing of GANs for Improved Quality, Stability, and Variation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Progressive Growing of GANs for Improved Quality, Stability, and Variation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.109298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.109298Z digest=sha256:e6e61e33fc6d9ed906902bd8a9c72fe0e08621a6a5e5b105b182780f1b6edc13

Observation 70d9bd6f-303d-43ab-bc83-c483a9a15efe · outbound

This paper cites Auto-Encoding Variational Bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Auto-Encoding Variational Bayes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.116002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.116002Z digest=sha256:453f6a6da247f9aa235dab5f07cf4de821c84b5b6d174e6b7183e73423fced60

Observation 92f5ceb6-770f-4805-883a-6b85eeaebe98 · outbound

This paper cites Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.644039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.123141Z digest=sha256:a8737df74c86bbdef0cd7bc3c83e4e79f4e7672fe904ecb204f128aa9e61d1de

Observation 3dd086a4-7e79-471b-ab56-56afba6db1ab · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.137150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.137150Z digest=sha256:b4e3b8ee56f56a0cf4e9c6aeb62acdcf8690bf6282ea8d04daa3c4c59be277fc

Observation 71f0f88a-8fbf-4865-a174-bfb8436f0712 · outbound

This paper cites Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.626817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.144757Z digest=sha256:0eb950b9d06a00ab184b7222598aa101faae37a3b2af2de54a4bed5d77daf929

Observation e1d08a69-acfa-4ea1-9b94-ef572a375614 · outbound

This paper cites Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.152540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.152540Z digest=sha256:15c974da9bfe784aaa069f6ca7e0a75c748649d2563b87a400dc9e50b7bedfce

Observation 8bc0d17b-7152-4eb6-af9e-3b002e5dac2c · outbound

This paper cites Which model generated this image? a model- agnostic approach for origin attribution.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Which model generated this image? a model- agnostic approach for origin attribution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.607534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.158775Z digest=sha256:28605734a4a27b14f8221b08e0cdbc927b4dc69b7bc979b56ead062b33cc95af

Observation 2ccc9a99-8687-48b0-9a78-b768756d6e54 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Improved Baselines with Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.165369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.165369Z digest=sha256:1cf2229e913e308916ca9459ee477b97ca5ba6e453c5666da32a6d7003d4b041

Observation 467148d4-0fec-4a41-84a1-ba0a941534a0 · outbound

This paper cites Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.171091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.171091Z digest=sha256:79af2063a92e7daa7033e2694a40758211e7f23222f6cec488ca802b6a66526a

Observation fcd2f6c9-6a9c-468f-8eb5-57078ab22787 · outbound

This paper cites Latent guard: a safety frame- work for text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Latent guard: a safety frame- work for text-to-image generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.590711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.182742Z digest=sha256:44c4a6acd622c3039a7de71d095bb97c5993e9643974e7293996fd8bcf0a578f

Observation c05fc1f7-6df0-48dc-80aa-8acbf6ed9b58 · outbound

This paper cites Multimodal prag- matic jailbreak on text-to-image models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Multimodal prag- matic jailbreak on text-to-image models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.573013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.188138Z digest=sha256:5c099f33e79456ef86c38766262528ad2550a8dfb21cf6ec561cd8442ac9ceeb

Observation 16392452-1f48-41ed-a583-505200cdc8f9 · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.195700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.195700Z digest=sha256:ce6d451356bdcc8527e459b2b0e1b81b06778e233a57e2f00d302a08a9103ddc

Observation 2435339a-a0c1-480a-a9b8-1a3770a8d3de · outbound

This paper cites Large-scale celebfaces attributes (celeba) dataset.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Large-scale celebfaces attributes (celeba) dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.554589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.202074Z digest=sha256:29ee930a7ece65c728e103cae80115cefab09b2b02144387d160ea4e39a25e89

Observation a06b4bea-ebaa-42a4-a776-85fd4c93f603 · outbound

This paper cites Information constraints on auto-encoding variational bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Information constraints on auto-encoding variational bayes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.537226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.210370Z digest=sha256:194a8caf190d865ed4cb1714a2af43cec7a669711042094f3dac5e0296e44843

Observation a0275460-a58a-46f9-ac1a-58dd1529dece · outbound

This paper cites An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.518701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.219957Z digest=sha256:7877d671904119fc0e823febd0abda09a2f598d676b362a5343c6e3baf6ecb76

Observation 198e468f-803d-4cee-a8a4-9e0d4bd57a69 · outbound

This paper cites Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.226826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.226826Z digest=sha256:6501a3e2f9f11131e62cf6297f0267ad5a02d99e643cb526f3cf66bab99a12f6

Observation e2edfa60-77c8-47ce-934a-0d9a200ea5dd · outbound

This paper cites A holistic approach to undesired content detection in the real world.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A holistic approach to undesired content detection in the real world

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.500221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.233197Z digest=sha256:1a4f22ab9cd0cf83fe30009ae08351a54d46f997115bedb8c2614c10d2b11e29

Observation 7626b988-ae74-4559-bbe6-3495c130d857 · outbound

This paper cites Dreamguider: Improved Training free Diffusion-based Conditional Generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Dreamguider: Improved Training free Diffusion-based Conditional Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.247131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.247131Z digest=sha256:979dda4f0f08dc26b5cb5a8b05c64d798b3aaca0cadbec754577f2133f1a3f02

Observation 1dc9f6e0-757f-4237-bf2d-85d1d5b31f90 · outbound

This paper cites At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.484104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.262943Z digest=sha256:72c5d043acc7e6719d463ec42e03926a44f631795d26f4fc8dc3f4e76a7ff92e

Observation d700a64a-157c-41d8-aa2a-db3adfbc0425 · outbound

This paper cites Contrastive denoising score for text-guided latent diffusion image editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Contrastive denoising score for text-guided latent diffusion image editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.467275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.272875Z digest=sha256:61d561845f4dfbd2840141382f5f3084522a8dfc56570ccdfbbb28295ff83578

Observation 11ea48a6-e231-4b9b-a4a2-dc2a461d0fb1 · outbound

This paper cites White-box Membership Inference Attacks against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models White-box Membership Inference Attacks against Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.279099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.279099Z digest=sha256:ebbb1992656287503fdf9d690b5bcc877804c1d9e14d7c6920804b170e022a96

Observation 580454fc-c7ef-4465-8431-183ef9cace1e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.285778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.285778Z digest=sha256:8a649bed8e871eb85b91074743790edba204fd58e82ca0f617a528f91d947358

Observation d24ded01-08f6-4477-91f8-1bef07cad98f · outbound

This paper cites Safe-clip: Removing nsfw concepts from vision-and-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Safe-clip: Removing nsfw concepts from vision-and-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.450071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.294864Z digest=sha256:9371d558bcc76452a660c3c8322c9d55473e5e61fa69b37d46fce4d470257675

Observation f3fdfe8d-e283-48ab-8b0c-c59edf536f87 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.431641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.302095Z digest=sha256:9fe7da36da36f8bc591842f235c7916577a49c9c177923f15a56c6dab3f208ec

Observation 79bb20c5-c338-4f24-9e9c-1e481733286c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervi- sion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.311262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.311262Z digest=sha256:55f82390aa36a218ca061de121823b85a9d96b49d7f46dfa09e1a2bc0ad6e772

Observation e85317c9-32b4-4461-8a2b-435af4216843 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.317220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.317220Z digest=sha256:17b2923865468a7ff751eb3d0b27d11adfa9f1445d8d1dd25ac29df390bb59d7

Observation e1901301-6bd3-469d-9c37-12a2c11fd13e · outbound

This paper cites Red-Teaming the Stable Diffusion Safety Filter.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Red-Teaming the Stable Diffusion Safety Filter

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.323922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.323922Z digest=sha256:66810a6bbfc3a1cbc12d95089f8f443347eea72022e7dbcdaff2d2780221df36

Observation c466172a-b985-4d33-9e9a-f83673b770fc · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gener- ating diverse high-fidelity images with vq-vae-2

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.330422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.330422Z digest=sha256:67e7a22037a2f277b574d6d2bfb694fc2ca13fdb1e5d84375536cb500dda2348

Observation 8750fd41-ada9-4845-9cb1-f114b81e8ce3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.335821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.335821Z digest=sha256:e3b60ab000a87ea8673a17d67fc276006d32be1a398441da5328d27383148556

Observation d081260a-f4b2-496e-baf4-607ff318d5d7 · outbound

This paper cites Raising the Cost of Malicious AI-Powered Image Editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Raising the Cost of Malicious AI-Powered Image Editing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.342958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.342958Z digest=sha256:9d16744438d0d485e9470601da6c5322a30ab5b286cbe8a32fbb307c8bc359b2

Observation 4c45e4a7-1f94-466f-ab87-acc99dd4d8b5 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Deep unsupervised learning using nonequilibrium thermodynamics

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.349013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.349013Z digest=sha256:f3a02fbd5ba4432238c59def0f35d2c3f384e3b4b4a83759601695e80f8341f9

Observation f961a41a-8ff6-474c-966e-e71c5bb3d1e8 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Emu: Generative pretraining in multimodality

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.372299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.354839Z digest=sha256:87f40e39ab5157199240c594340be80f512a720ba89b8a511edb1ee73eb02126

Observation e802768e-17cc-4a74-9438-ca6a84111fd6 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative multimodal mod- els are in-context learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.344759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.361162Z digest=sha256:810d175f58cad4d9e23b58fd273ee60fbaca2d343baafd39b8416c65e5ea4257

Observation b939e121-7ed5-4950-9f0e-5c6f072d5143 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.368896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.368896Z digest=sha256:870c7973d139451409dd34ec72c7621405f4c62753e9528674f571a34dc75ad7

Observation 8ba5fa7b-4918-401d-8eb1-8e8bb97917b7 · outbound

This paper cites Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.325116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.379570Z digest=sha256:37a7b029e630ccc84933d51793dd8e61d3a8ebd8950f8b7d8a9aa4830c246d3e

Observation 31f190b1-04b0-4372-8834-c35e970dbfb7 · outbound

This paper cites Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.386103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.386103Z digest=sha256:5b966156cafcf81b9f4f645458e30e6bfc8d1eccd4b6e36fab73f1d6c7e8c3c4

Observation e8f7d067-0e3b-4097-8732-9a4e1e7fa217 · outbound

This paper cites Sneakyprompt: Jailbreaking text-to-image generative models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Sneakyprompt: Jailbreaking text-to-image generative models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.307571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.396910Z digest=sha256:e30bd750aed0f761e065b8a38125e5ce99bd61bf2233f521916bd835ee78d9ca

Observation cc4ba249-05ba-46ef-9cf8-ef0e9ef2fa5b · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.402465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.402465Z digest=sha256:5fd3e075eb4749cdaaea7cd0de210056251d7e8efc0ed13caa9997bf0970e367

Observation 8c07f2aa-0b95-473c-8239-696ad4c3029f · outbound

This paper cites Inversion-based style transfer with diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Inversion-based style transfer with diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.291166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.408500Z digest=sha256:558e7d0306b6ccd9e3a3872172027f16b8af414062e7a6fa9856fe20145dc0ae

Observation d2ec859b-e882-477c-a0fb-e1c0ca8066d8 · outbound

This paper cites Defensive unlearning with adversarial training for robust concept erasure in diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defensive unlearning with adversarial training for robust concept erasure in diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.274376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:41:16.422936Z digest=sha256:5a25b1488ccfc0017a941e63b4b446e235606a61975beb37548281da420303cb

Observation b5acb7ce-93af-49d7-8d8b-8ddd9a55b355 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:41:16.429938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.429938Z digest=sha256:7ac20c4420eff6695ed2c9210ca16ff1236de05a6b3dcaefa10f988f6cde599c

Pith citing papers

No inbound Pith citation observations are available.