Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:46.283268Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.00114.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:46.283268Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:24:49.439929Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T09:28:10.399987Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 959f4d25-889c-4b7e-8222-165a1f45a24e · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42e7d51d-2a0b-47b5-a97f-57815e83a9dc · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93cd3123-85b5-4ec7-aff2-0fd7fc026c59 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Learning transferable visual models from natural language supervi- sion
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5de27f-7e44-42aa-8475-5cb864f2e4c0 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7e1a67-32cb-4e9e-8dac-39c719987036 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c59ec9-9fc4-4daf-b530-db04b9f84557 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce35016b-867c-45e8-851c-c68f624124e4 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Irad: implicit representation-driven image resampling against adversarial attacks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d383b11-c351-4d92-ba72-cd6e3a1d9365 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b78f2d46-0d51-47ef-9122-d9ed6bb4c198 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the Robustness of Segment Anything
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4e8087-764a-4fde-ab6b-7e2d6d3dade8 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial relighting against face recognition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 145a7d42-1fff-401a-a238-e83ccaf636cf · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3216e681-342c-400c-96e0-d4848e02de32 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments ALA: Naturalness-aware Adversarial Lightness Attack
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7130ca-b754-4af9-a1ad-cb6de616d5e7 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On evaluating adversarial robustness of large vision-language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f5256f5e-fe27-4a98-bd75-772143bdbc7a · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd03d5ca-c492-46a3-89e3-2166e0b99744 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Transferable multimodal attack on vision-language pre-training models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0cdd27f-7577-49e8-84f5-1c4ae4132fa9 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards adversarial at- tack on vision-language pre-training models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 11faa285-7bac-4ba0-873f-ab61e1f4b131 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb43b57d-0e3f-43d5-a6a7-51b8d61b2136 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ffa708c9-3947-4b29-a55c-9cf780b3067a · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474571cb-8333-43b4-815a-9125db1e13e4 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser-2: Unleashing the power of language models for text rendering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09c6dfb6-7f04-4e80-9e47-3c58018adf70 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c31f85b-ca59-4af3-b03c-cc86b6626988 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the robustness of large multimodal mod- els against image adversarial attacks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e49005a2-0855-4804-b1a8-fc25f761621d · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7977e5ac-6b8b-49aa-b551-2e6a9470ce12 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Multimodal neurons in artificial neural networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87c4e92-7e62-42c5-b29a-d313afa8ad4a · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Blended diffusion for text-driven editing of natural images
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7952bc4-54b3-4cea-802d-1ac0f53be564 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Dis- entangling visual and written concepts in clip
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6695e2fb-87b9-4fa1-98bf-2e5cb403c4c6 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Patching open-vocabulary models by interpolating weights
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60625371-7448-446f-be21-67bd553096ee · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defense-prefix for pre- venting typographic attacks on clip
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 070e2b28-f9a3-4f9a-a4c0-fefd294b99b6 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defending lvlms against vision attacks through partial-perception supervision, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cbe3766e-6983-4894-a68a-269518e540d8 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial Machine Learning at Scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e34ba8c-1c18-42cf-af62-868a7192a2f2 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adver- sarial examples in the physical world
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0960cc-a898-421b-ba7d-434aa136a5dd · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation acbefc5f-ff81-44d3-b807-6ad022e671fb · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Robust physical-world attacks on deep learning visual classification
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2cdecedd-0da2-4719-b7f6-7a826f997784 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards transferable targeted 3d adversarial attack in the physical world
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3b4cd52-3786-4608-ae85-fbe44f906476 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial t-shirt! evading person detectors in a physical world
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f3ceabdf-d21f-4450-b350-97bd03b9b2e9 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Fooling thermal infrared pedestrian detectors in real world using small bulbs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2b2d1cd6-6f27-4c9a-9029-48c52e810cef · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af9c87b8-c84c-4b71-a13a-e3dc5b7207f4 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Hotcold block: Fooling thermal infrared detectors with a novel wearable design
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69b89cfe-1f66-4efa-affc-6774cc2a4bf3 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial camouflage: Hiding physical- world attacks with natural styles
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc9c1e9a-f670-481d-9cef-c6846de4cbec · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Uni- fied adversarial patch for cross-modal attacks in the physical world
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b928e87-84ca-4978-ab13-29b03ccb9fdf · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning, 2023
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1db9c836-f655-43c3-be4d-8e930646beb8 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 676dac00-ae12-478b-9c19-07031a83cdb6 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e52cf9-54d1-4c00-9fb4-86d53d62740c · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser: Diffusion models as text painters
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e92ab44-e181-4564-93c5-4b565c4bc00c · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments LingoQA: Visual Question Answering for Autonomous Driving
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b04b758-3b03-4cd7-a2ab-a7f3b704e047 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3541e0bc-722a-44f4-9506-9c0f60ac74cd · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 833ecf8b-6c0a-45c1-b5a0-0b8c86684a08 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a796c973-ab8b-4611-b8f6-56d235886a28 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 21916217-b062-4b09-aa23-d164e0d419e9 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3750fd6c-f257-4ab3-917a-4af553051763 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52dd5a22-b884-479b-be75-0153a69ef098 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1d58c31-a94e-4c01-adec-81ef33110274 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 44c9d7d8-c66f-45b5-a7b4-f06cc5e64c39 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 844466fc-bf8e-45b0-b5fd-b3452e64edbb · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f80e2030-bbc6-4cf3-81fe-b1914748ed19 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0271bf1b-1787-46e6-8cab-e11bbb31e059 · outbound
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90f37915-c44a-4850-a22b-07dd97cc1639 · inbound
MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e69485-f1b8-4bfb-a5b2-c017f58bba82 · inbound
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906c3d0e-f8b3-4f13-a6de-6af3f225549f · inbound
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 82290c1d-6857-4760-b304-ac8dc1d4b0fc · inbound
Not What You Asked For: Typographic Attacks in Household Robot Manipulation SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.