Pith. sign in

Paper Citation Record · LEDGER

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

As of 22 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.00114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00114 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:46.283268Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:24:49.439929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:28:10.399987Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959f4d25-889c-4b7e-8222-165a1f45a24e · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.512263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:44.876264Z digest=sha256:ee72911a3a70b4eaf3296a4976bedb58b712384d33e4bf987120262fcd9018bb

Observation 42e7d51d-2a0b-47b5-a97f-57815e83a9dc · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.941580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.941580Z digest=sha256:fc26a6083fe9283f854764ef0351dfbc290f69328a08bd0cc703660298041a18

Observation 93cd3123-85b5-4ec7-aff2-0fd7fc026c59 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Learning transferable visual models from natural language supervi- sion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.946224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.946224Z digest=sha256:faceb95594d3a41817a1c3e836d2eef91ac946f12f288360a25a02f477002b5d

Observation 1c5de27f-7e44-42aa-8475-5cb864f2e4c0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.950726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.950726Z digest=sha256:e567e4a439019dd73f7f0b16681c562fc82fe93954ae26c6ac768a1d423a4925

Observation dd7e1a67-32cb-4e9e-8dac-39c719987036 · outbound

This paper cites Visual instruction tuning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.954832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.954832Z digest=sha256:41909c79ff9a650a091614eafb0d218ceb9758ddd410ad6b5203e5370059194c

Observation 09c59ec9-9fc4-4daf-b530-db04b9f84557 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.958598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.958598Z digest=sha256:9c3ca198ffe85cc1de31b309071942bb8b994217043e690e4d2ae85e543b3f22

Observation ce35016b-867c-45e8-851c-c68f624124e4 · outbound

This paper cites Irad: implicit representation-driven image resampling against adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Irad: implicit representation-driven image resampling against adversarial attacks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.481016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:44.962914Z digest=sha256:5e3cba1beec6f7ef47f1e1fc49294b73907af07374f19e667bbc894c97d83979

Observation 2d383b11-c351-4d92-ba72-cd6e3a1d9365 · outbound

This paper cites Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.468554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:44.966374Z digest=sha256:b29632a90f6b100aaac24f43b1765dec642681c1cc2199f1a2a37618989d7f08

Observation b78f2d46-0d51-47ef-9122-d9ed6bb4c198 · outbound

This paper cites On the Robustness of Segment Anything.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the Robustness of Segment Anything

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.970401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.970401Z digest=sha256:ff8d42a28ae5e07632abc45647102b5db2cac851e5fe6d26764d0e0a9d88f8ae

Observation 3a4e8087-764a-4fde-ab6b-7e2d6d3dade8 · outbound

This paper cites Adversarial relighting against face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial relighting against face recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.391424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.013666Z digest=sha256:de1ad7fdb219667e3e0a0cd302bdb56bf6372c3978c78ab7d94b8ee6166dc139

Observation 145a7d42-1fff-401a-a238-e83ccaf636cf · outbound

This paper cites MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:46.475722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.061056Z digest=sha256:f995dc61b97d5a486d809201f63188867f3038e404c907bbca9473dd0a56d442

Observation 3216e681-342c-400c-96e0-d4848e02de32 · outbound

This paper cites ALA: Naturalness-aware Adversarial Lightness Attack.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments ALA: Naturalness-aware Adversarial Lightness Attack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.066082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.066082Z digest=sha256:c1d730f36c2fea3ce684a669ad3e045e40824aed960ad707edb6bcbf8025b54f

Observation ac7130ca-b754-4af9-a1ad-cb6de616d5e7 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On evaluating adversarial robustness of large vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.260042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.070592Z digest=sha256:4578f00c59a3854057a07ed3ed3256974776ac4af2ca25975a09cf563c550a4b

Observation f5256f5e-fe27-4a98-bd75-772143bdbc7a · outbound

This paper cites InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.074404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.074404Z digest=sha256:0a69b60ef840f0daab7eba4aafbc9c3505d5574d9eb210b2032236f58baadfea

Observation cd03d5ca-c492-46a3-89e3-2166e0b99744 · outbound

This paper cites Transferable multimodal attack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Transferable multimodal attack on vision-language pre-training models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.248497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.081051Z digest=sha256:93349512486f54f2c041f9b0e3dde1d9eb793877d1d035e2c3047baa775d7e83

Observation c0cdd27f-7577-49e8-84f5-1c4ae4132fa9 · outbound

This paper cites Towards adversarial at- tack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards adversarial at- tack on vision-language pre-training models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.237153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.130705Z digest=sha256:fc208ab7e632eaffc03a348f9c807b5911d373319bfe82a95f9d9c9f1ddabbb5

Observation 11faa285-7bac-4ba0-873f-ab61e1f4b131 · outbound

This paper cites Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.101647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.206041Z digest=sha256:5864962983b52b5c309159a59f6e20f6dd465e6dbb335ccfc6bc35d4543204bf

Observation fb43b57d-0e3f-43d5-a6a7-51b8d61b2136 · outbound

This paper cites Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.059640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.283076Z digest=sha256:4f054865dd73c5bdc7bedbba0b125ab481146ebe3ebd710698bb789e0c4621ff

Observation ffa708c9-3947-4b29-a55c-9cf780b3067a · outbound

This paper cites Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.287680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.287680Z digest=sha256:6fb3aa6c0c33f820ffd56593b17864e2d5033064fd862df7cada99d4761ad504

Observation 474571cb-8333-43b4-815a-9125db1e13e4 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.997162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.292824Z digest=sha256:eb0af724d800c7f6f3f38d121705e3af35e30f299f78a59392b55854c02bb392

Observation 09c6dfb6-7f04-4e80-9e47-3c58018adf70 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.296501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.296501Z digest=sha256:f13a17fe802d4f656544dbd7088dd8c3391d5495abae431e3872d3604b3345ce

Observation 3c31f85b-ca59-4af3-b03c-cc86b6626988 · outbound

This paper cites On the robustness of large multimodal mod- els against image adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the robustness of large multimodal mod- els against image adversarial attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.950273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.300764Z digest=sha256:8e12d2e5c7a1be631e05becb147827ac11e5620758b12856f61ae053fb069293

Observation e49005a2-0855-4804-b1a8-fc25f761621d · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.365436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.365436Z digest=sha256:e00e635ab8377dbb1fba1c0bb00b1775f1f7da872592cfd39ddf673b6307df2c

Observation 7977e5ac-6b8b-49aa-b551-2e6a9470ce12 · outbound

This paper cites Multimodal neurons in artificial neural networks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Multimodal neurons in artificial neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.385682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.385682Z digest=sha256:f24273be02187701a0cd1ea9760218f1b50932d481f9aa8ec523e998bfd61688

Observation f87c4e92-7e62-42c5-b29a-d313afa8ad4a · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Blended diffusion for text-driven editing of natural images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.389653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.389653Z digest=sha256:df81d79b664e84fe2dc3572bb14246e33cf383512c9e76456258bb55abb95ec8

Observation f7952bc4-54b3-4cea-802d-1ac0f53be564 · outbound

This paper cites Dis- entangling visual and written concepts in clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Dis- entangling visual and written concepts in clip

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.838749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.393835Z digest=sha256:26ab47fb7f7fe586b56dd62dde66e48d97833e2d8a849defb7d76d217a4c51c8

Observation 6695e2fb-87b9-4fa1-98bf-2e5cb403c4c6 · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Patching open-vocabulary models by interpolating weights

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.397035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.397035Z digest=sha256:c6735e641be14fb9929727eccfc006154315554c1ce5d2806d8a9f14c61dc902

Observation 60625371-7448-446f-be21-67bd553096ee · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defense-prefix for pre- venting typographic attacks on clip

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.788280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.400822Z digest=sha256:355a14bf4604c047a327ccc9339eb8999844b856946e5be2153f621fb1028acb

Observation 070e2b28-f9a3-4f9a-a4c0-fefd294b99b6 · outbound

This paper cites Defending lvlms against vision attacks through partial-perception supervision, 2024.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defending lvlms against vision attacks through partial-perception supervision, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.681126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.404438Z digest=sha256:a0030a566ac6aed6535c778d9b37d9236cc7cf49ade4ef49b4a28bb30d92f056

Observation cbe3766e-6983-4894-a68a-269518e540d8 · outbound

This paper cites Adversarial Machine Learning at Scale.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial Machine Learning at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.408168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.408168Z digest=sha256:76194012682d4a24c1eb610d0da81b5c46884bebb81b3484de1f82793cf08d12

Observation 4e34ba8c-1c18-42cf-af62-868a7192a2f2 · outbound

This paper cites Adver- sarial examples in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adver- sarial examples in the physical world

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.549168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.549168Z digest=sha256:71b873eaa4a3cb7bf67c204388b78ed6d9075e3816645c454d308c5e5bd8f5a5

Observation 8c0960cc-a898-421b-ba7d-434aa136a5dd · outbound

This paper cites Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.601997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.580877Z digest=sha256:bba417b7b3895b955918187e921b1df6e9dd2d3f0011a76025713d2376eb24da

Observation acbefc5f-ff81-44d3-b807-6ad022e671fb · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Robust physical-world attacks on deep learning visual classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.588605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.584533Z digest=sha256:ac93610896d99bb02df1894e665532b1777b0463b0bec0f60cf1facd6125fea3

Observation 2cdecedd-0da2-4719-b7f6-7a826f997784 · outbound

This paper cites Towards transferable targeted 3d adversarial attack in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards transferable targeted 3d adversarial attack in the physical world

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.398716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.588548Z digest=sha256:95b26c0039650773b9e7b15e167e5bc1e1d871c3d3f3da4a96d88653d70cd45e

Observation b3b4cd52-3786-4608-ae85-fbe44f906476 · outbound

This paper cites Adversarial t-shirt! evading person detectors in a physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial t-shirt! evading person detectors in a physical world

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.353693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.592280Z digest=sha256:3a576e84785d152096a59dec5cb0354b8dd71bb5f567ea0bbf11c83f13634ae1

Observation f3ceabdf-d21f-4450-b350-97bd03b9b2e9 · outbound

This paper cites Fooling thermal infrared pedestrian detectors in real world using small bulbs.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Fooling thermal infrared pedestrian detectors in real world using small bulbs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.343421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.597309Z digest=sha256:49c1c09cff7422c4d86dce29736e64fb6ab869a2e63f9d410e05d08ef2648598

Observation 2b2d1cd6-6f27-4c9a-9029-48c52e810cef · outbound

This paper cites Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.330137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.636881Z digest=sha256:519c43cb36ec183e506760e2fdceb4a42aa8f907cfdd395796a4c6a27c568cb3

Observation af9c87b8-c84c-4b71-a13a-e3dc5b7207f4 · outbound

This paper cites Hotcold block: Fooling thermal infrared detectors with a novel wearable design.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Hotcold block: Fooling thermal infrared detectors with a novel wearable design

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.280549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.682499Z digest=sha256:0da6e4bf2c61ae6f692ee8f177e4afe5f0cc4082dab3e774e6ff5c67ed762c22

Observation 69b89cfe-1f66-4efa-affc-6774cc2a4bf3 · outbound

This paper cites Adversarial camouflage: Hiding physical- world attacks with natural styles.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial camouflage: Hiding physical- world attacks with natural styles

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.213692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.700581Z digest=sha256:12cc7b920bbc73c1d82ea0033a498d4dbaafbe335ba38aaa6d9cfbe8ba15cf87

Observation dc9c1e9a-f670-481d-9cef-c6846de4cbec · outbound

This paper cites Uni- fied adversarial patch for cross-modal attacks in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Uni- fied adversarial patch for cross-modal attacks in the physical world

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.068983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.704277Z digest=sha256:c28ab0e225db8a44516f43a913f2e29cfe5c36bc27bb1804e166e3a74dec1dad

Observation 6b928e87-84ca-4978-ab13-29b03ccb9fdf · outbound

This paper cites Visual instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.057281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.708071Z digest=sha256:f2eb8f7f0187538a5f505db3856e269029086485992981156b91939aaa3fee61

Observation 1db9c836-f655-43c3-be4d-8e930646beb8 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.044858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.711818Z digest=sha256:d7cf9fb7b9c4f2775a85871dba4bcfccc2041de8037e0face06235dda46665f6

Observation 676dac00-ae12-478b-9c19-07031a83cdb6 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.715928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.715928Z digest=sha256:971fce99b7828cff18f4239616948f4ae543cd8fc2266bb4ce3d9a5d17fae8d6

Observation d3e52cf9-54d1-4c00-9fb4-86d53d62740c · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser: Diffusion models as text painters

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.030989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:45.720628Z digest=sha256:347493d06dbbfbe10faf277507a7941df41c79f7a072edbfa9cdb5549378cbba

Observation 1e92ab44-e181-4564-93c5-4b565c4bc00c · outbound

This paper cites LingoQA: Visual Question Answering for Autonomous Driving.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments LingoQA: Visual Question Answering for Autonomous Driving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.724257Z digest=sha256:00c984f4b54fecabc4261b3dd4990cce5deab2205eb2e3137cddb7b87ae817bf

Observation 8b04b758-3b03-4cd7-a2ab-a7f3b704e047 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.728189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.728189Z digest=sha256:cb705e2d11d8e6d9bbbba0cab30f03fe5c4a4a26bb62322a76b69b014d9a5a8c

Observation 3541e0bc-722a-44f4-9506-9c0f60ac74cd · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.868817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.868817Z digest=sha256:8d95d781b1599587a592f2cb8ee6d2753ed78750131d9a6544cdcd9da1247782

Observation 833ecf8b-6c0a-45c1-b5a0-0b8c86684a08 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.950480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.001670Z digest=sha256:8b177e515e70f1e369fe936919c6f50523a77566441b3fd89388e56882ee1dc0

Observation a796c973-ab8b-4611-b8f6-56d235886a28 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.937711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.006249Z digest=sha256:fc6faafb2b8f0eb60a55e4f737c79ae0dff2ca097920ae6741eb2df62e1a31f2

Observation 21916217-b062-4b09-aa23-d164e0d419e9 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.879841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.010987Z digest=sha256:341b144c17d361fbc852982f3ac59473e1a653f17e7dada7a5df617a2577f527

Observation 3750fd6c-f257-4ab3-917a-4af553051763 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.788254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.015327Z digest=sha256:6c74b96ad75733d6aa23be9033bfcd53a63703a3ffb8c1b3b1b2717a4a0f8f4b

Observation 52dd5a22-b884-479b-be75-0153a69ef098 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.776382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.119310Z digest=sha256:0856135f90576b75f5c0def2471511a67629f446bab44bf7c7775a37aae1b0a5

Observation f1d58c31-a94e-4c01-adec-81ef33110274 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.720866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.176161Z digest=sha256:8512cdc0df2fd45984de139e0add4eef662e2be8d804366c4d6ec47671a8f30c

Observation 44c9d7d8-c66f-45b5-a7b4-f06cc5e64c39 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.653487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.271794Z digest=sha256:e527969ec6ed6bd2846ba1cc5ccd924d5244ea6942b67789a8a9eb3a7f6d5bcd

Observation 844466fc-bf8e-45b0-b5fd-b3452e64edbb · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.642018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.275743Z digest=sha256:aaa6425f789b62e3dfb627194c372c24412e44820a58303f5de6bd49dfd3f54b

Observation f80e2030-bbc6-4cf3-81fe-b1914748ed19 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.631034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.279441Z digest=sha256:5dd3853bd3ab78d508de96f31107655a9362f4073f528426f29c4b9b84edad10

Observation 0271bf1b-1787-46e6-8cab-e11bbb31e059 · outbound

This paper cites colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:46.561343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T10:48:46.283268Z digest=sha256:33e4bcf1cc8e53bd15bfeaeeba250395f0c95fdf378224b476dbf1088d65265b

Pith citing papers

Observation 90f37915-c44a-4850-a22b-07dd97cc1639 · inbound

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents cites this paper.

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:49.439929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:24:49.439929Z digest=sha256:108f1ed91c9a1c75be164356d1a2126858408aabaccc32d7902668ecd2bd1f46

Observation f6e69485-f1b8-4bfb-a5b2-c017f58bba82 · inbound

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision cites this paper.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.423820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.423820Z digest=sha256:6abd22298b41071bdb8c17f1a32b0e0479f4ee920de96e28a7f6929a31625268

Observation 906c3d0e-f8b3-4f13-a6de-6af3f225549f · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:8b78e48398055a84d3fc0feac83d40c8af4b78f43cf6dba58e82ad923ec965da

Observation 82290c1d-6857-4760-b304-ac8dc1d4b0fc · inbound

Not What You Asked For: Typographic Attacks in Household Robot Manipulation cites this paper.

Not What You Asked For: Typographic Attacks in Household Robot Manipulation SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:28:10.401596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T09:26:39.444806Z digest=sha256:9ac9bef2b9e132ffba675795993bb35e641dae6d4104becd6dbf9658016dd340