Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:32:16.575532Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2512.11899.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:32:16.575532Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:13:30.825592Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation b62e09d8-2198-4298-a28d-0503e3fc879f · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650861ec-712b-40e4-a232-f85000d8a688 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Defense-prefix for pre- venting typographic attacks on clip
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131f41b8-e007-4814-97dc-fba63b42e6ef · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138f1e21-53c7-4995-93de-e40075da3863 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vizwiz: nearly real-time answers to visual questions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c7affb-f98b-421d-ad81-b510ef3fb9b7 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scene text visual question answering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d4de9c-bfb2-4554-9f91-cb98fd1187ee · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Adversarial Patch
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0973c8b4-144f-46c8-ad17-d5e8249779e0 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scenetap: Scene- coherent typographic adversarial planner against vision- language models in real-world environments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0b8aa7-d88f-4a30-82ac-00aab4f201a5 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2968110c-c787-4671-a785-fc8a5fa9e2cb · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b386fd-4eb8-40e8-83ab-911388c9cd63 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Sari sandbox: A virtual retail store environment for embodied ai agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b9e79c-f634-44c5-b915-a523806e4b02 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80961c2b-d8f3-4c6f-a3b3-f968d5e168ce · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Multimodal neurons in artificial neural networks.Dis- till, 6(3):e30, 2021
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e65586c-d3e4-4fa0-b967-4987a0fe5456 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Explaining and harnessing adversarial examples
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3ace6b-f35b-44bd-9079-c44208936c5b · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c0390c-2e70-4eb3-9100-82770a31394c · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce04df3-e2d0-4fe9-9530-2c399b60e5f0 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards mech- anistic defenses against typographic attacks in clip.arXiv preprint arXiv:2508.20570, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b47a868-7136-425e-9eb9-6fb24fb95377 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1192ff8-7952-446d-8337-ed1ac55be160 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models A diagram is worth a dozen images
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d9b2d6-f613-4a79-8c36-a48e44a29391 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ocr-free document understanding transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04dd6a8-432c-4a20-aeb2-84c187b08f7d · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9afa5a5-68a3-44c6-b85e-50b7316d86ad · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf4c22d-5771-4006-911d-cdf510307cdc · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f520a9a5-0818-40f4-ae91-33b9da07f6d8 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06937145-2ced-4d36-ac1f-a495c0bcef7f · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727b6360-920e-4756-a00a-d1e014e06490 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards deep learning models resistant to adversarial attacks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3519e281-73ed-40fa-8c2f-f5447c284c62 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models SmolVLM: Redefining small and efficient multimodal models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cd3e02-d5c3-4fbc-932f-2a8abf1f5f63 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cff08e-5b1f-47a5-8fd5-5a3302427dfb · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d02b1dd-17a3-44fe-96c5-729d8eed5086 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Dis- entangling visual and written concepts in clip
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5bfdb1b-d947-482d-b912-028e20a2ad05 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Docvqa: A dataset for vqa on document images
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb294161-0c90-45ce-94e8-03e4905e4f15 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Infographicvqa
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae15cad-b497-4d7f-8413-2808442bd699 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77e97fb-e423-4b34-b146-e0690934aa75 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learning transferable visual models from natural language supervi- sion
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a456c5-9456-4519-87d3-5d3b50344b24 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Roadtext- 1k: Text detection & recognition dataset for driving videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c77226-9f1d-4caa-a05c-7ea2a8226778 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards vqa models that can read
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6111572f-1e48-45e5-aa45-2f7aa93ad864 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4167bf86-f16d-4b8e-811c-15ec23b1db84 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Mtvqa: Benchmarking multilingual text-centric visual question answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba91600-0ff3-4896-a8aa-28286e15e3d7 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Reading between the lanes: Text videoqa on the road
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6223ad18-9750-4a15-9e0c-5d507c4f92a5 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Clip in mirror: Disentangling text from visual images through re- flection.Advances in Neural Information Processing Sys- tems, 37:24523–24546, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1f5b2d-c3ea-49d6-940f-8a392d447906 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25f49f4-1556-4a45-aaf9-ee80c7fb7b5d · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models What word is written on the sign?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d17f6bf-6570-49e6-aca2-ecf4e1f569f4 · outbound
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c11ea9-d6e0-4f52-8736-23b2fab62fae · inbound
Token-Efficient Multimodal Reasoning via Image Prompt Packaging Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 351d7fa3-3160-45e9-b9a4-3a94618d81d5 · inbound
Towards Robustness against Typographic Attack with Training-free Concept Localization Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68e85a4b-f1a1-4ab9-8edb-9d8d4677b56d · inbound
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.