Pith. sign in

Paper Citation Record · LEDGER

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2512.11899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.11899 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:32:16.575532Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:13:30.825592Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b62e09d8-2198-4298-a28d-0503e3fc879f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.163631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.163631Z digest=sha256:2aa8bcd92d6ddfb556f6784bdd6c6f06e52cd6dbb02116cd96a194176c06bf6f

Observation 650861ec-712b-40e4-a232-f85000d8a688 · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Defense-prefix for pre- venting typographic attacks on clip

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.213168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.213168Z digest=sha256:1cd7d7f83b73bd5497164d7a0b73c2d58f0105ee0c2e1a63441fb85bd72f02ad

Observation 131f41b8-e007-4814-97dc-fba63b42e6ef · outbound

This paper cites Qwen Technical Report.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.292022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.292022Z digest=sha256:8a4daa1b77dda8f25356bdf48d09dea41bc64d0e5e6b5fc7b41ef4af59ad2c89

Observation 138f1e21-53c7-4995-93de-e40075da3863 · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vizwiz: nearly real-time answers to visual questions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.427960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.427960Z digest=sha256:7644be99dc90ae31ae34e4f4288a62d6380bfcf22afd5293109f50edb28e3177

Observation 51c7affb-f98b-421d-ad81-b510ef3fb9b7 · outbound

This paper cites Scene text visual question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scene text visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.572998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.572998Z digest=sha256:6ec883c8a2e7370cb598ca6c004757e3bdd27d563e07c49c1f64b3b05352e849

Observation 21d4de9c-bfb2-4554-9f91-cb98fd1187ee · outbound

This paper cites Adversarial Patch.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Adversarial Patch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.815967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.815967Z digest=sha256:3db67cbf276ee67a2b0ddfeeb2fb9119a30825685813c881386c8be0ff6027bf

Observation 0973c8b4-144f-46c8-ad17-d5e8249779e0 · outbound

This paper cites Scenetap: Scene- coherent typographic adversarial planner against vision- language models in real-world environments.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scenetap: Scene- coherent typographic adversarial planner against vision- language models in real-world environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.977043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.977043Z digest=sha256:e63d287bf5eaac205baa438471b977f20ad1f387c11d38cfc54ddfd657a9fdb2

Observation 2d0b8aa7-d88f-4a30-82ac-00aab4f201a5 · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.179174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.179174Z digest=sha256:3bd0f4fea8e47758ddcd6fae78efeb7d0e4af4721c51dbe0cb3cbb0d1c0ba92c

Observation 2968110c-c787-4671-a785-fc8a5fa9e2cb · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.529564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.529564Z digest=sha256:966e466ad19241258f8a35f936aa3e03985182dff7d27b8173282cbc8fcbcfe7

Observation 80b386fd-4eb8-40e8-83ab-911388c9cd63 · outbound

This paper cites Sari sandbox: A virtual retail store environment for embodied ai agents.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Sari sandbox: A virtual retail store environment for embodied ai agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.735538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.735538Z digest=sha256:bf67bd5baec1d7429983c39142812bf2d1dbb0897e6b2d2ef825c5fefa9cad47

Observation b7b9e79c-f634-44c5-b915-a523806e4b02 · outbound

This paper cites Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.899059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.899059Z digest=sha256:dbdb073c00760566533165d8e3faa77eb6c6269f9320537d29fd5dd32f8e7b95

Observation 80961c2b-d8f3-4c6f-a3b3-f968d5e168ce · outbound

This paper cites Multimodal neurons in artificial neural networks.Dis- till, 6(3):e30, 2021.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Multimodal neurons in artificial neural networks.Dis- till, 6(3):e30, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.046639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.046639Z digest=sha256:e0cae41b05a70de2dca0c7e7c2d63a980ed7ee9702aa695376540d0a6a620e1a

Observation 3e65586c-d3e4-4fa0-b967-4987a0fe5456 · outbound

This paper cites Explaining and harnessing adversarial examples.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Explaining and harnessing adversarial examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.141883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.141883Z digest=sha256:830f322311c33cd4e8d40d6e09f4cbe3884086eaa2c7822abc2ba91c0c9153c5

Observation 7e3ace6b-f35b-44bd-9079-c44208936c5b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.202228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.202228Z digest=sha256:df7f4b724ab43bf94bb76a1fd59956eefc9335ba1821b8cab8a3ea003b9c800c

Observation 73c0390c-2e70-4eb3-9100-82770a31394c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.274366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.274366Z digest=sha256:fbdcfb0fb6f3b1ca2974e56c86b62e92b793bdaaad160d50dd8e2e06758469fc

Observation 7ce04df3-e2d0-4fe9-9530-2c399b60e5f0 · outbound

This paper cites Towards mech- anistic defenses against typographic attacks in clip.arXiv preprint arXiv:2508.20570, 2025.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards mech- anistic defenses against typographic attacks in clip.arXiv preprint arXiv:2508.20570, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.328818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.328818Z digest=sha256:de28550fac05600004e1030b4aeb8a5bc3ef7fa45b1970ae14e7fcbf1783b3e2

Observation 6b47a868-7136-425e-9eb9-6fb24fb95377 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.406664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.406664Z digest=sha256:907c58eb5d7959607e4afd98594ac6ee19352ae4360d43b7d09792868e69daf2

Observation b1192ff8-7952-446d-8337-ed1ac55be160 · outbound

This paper cites A diagram is worth a dozen images.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models A diagram is worth a dozen images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.535410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.535410Z digest=sha256:30b6e98670e7514cae6d1084ed3bc0300d956a58d5490b2dfa0b41e0e6860b17

Observation f2d9b2d6-f613-4a79-8c36-a48e44a29391 · outbound

This paper cites Ocr-free document understanding transformer.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ocr-free document understanding transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.609358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.609358Z digest=sha256:56d87cd3aefd1281df879161bde65b412027be10e2ec8a57d6d012e30674c227

Observation d04dd6a8-432c-4a20-aeb2-84c187b08f7d · outbound

This paper cites an unresolved cited work.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.658093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.658093Z digest=sha256:0f9275af56b69113e07c7eaa7384a641b367bf59769741e191539c64d7d342e0

Observation e9afa5a5-68a3-44c6-b85e-50b7316d86ad · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.738428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.738428Z digest=sha256:bc714c870e85bf03c9714d2f4b96330d5dbbe3e9cf3dbf36424d9c20f77f5093

Observation ccf4c22d-5771-4006-911d-cdf510307cdc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.868129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.868129Z digest=sha256:10fd5e415d02e2de0045cc4ee3569599e29ae204165e145515525bff10ff36b2

Observation f520a9a5-0818-40f4-ae91-33b9da07f6d8 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.013156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.013156Z digest=sha256:6f40a951952ae7926b09c6642bee6cedd9b80d7f3f40f8bc51432955f97c2abb

Observation 06937145-2ced-4d36-ac1f-a495c0bcef7f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.096645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.096645Z digest=sha256:75f92937cee036176a91a146b478cd1f95c2ecae93ecdc1608d8f9014d571239

Observation 727b6360-920e-4756-a00a-d1e014e06490 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards deep learning models resistant to adversarial attacks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.180825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.180825Z digest=sha256:fa2d74e7ad8efc580f5e7e5876a8d405e0ece15ca80b504e4e0806845876095b

Observation 3519e281-73ed-40fa-8c2f-f5447c284c62 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.264270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.264270Z digest=sha256:7d92383946136222d6d4c701bf65b618e6bd05d7f9477d98a5997cb2f3ad475b

Observation b1cd3e02-d5c3-4fbc-932f-2a8abf1f5f63 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.357052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.357052Z digest=sha256:04fe530179739680b739029056d74283974ba5ae1f994da72e9ff34121a7c496

Observation 45cff08e-5b1f-47a5-8fd5-5a3302427dfb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.434754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.434754Z digest=sha256:4bfffcaeaa17f64e531728540159902ede206b342eeee4c022a76057b0ce9be2

Observation 4d02b1dd-17a3-44fe-96c5-729d8eed5086 · outbound

This paper cites Dis- entangling visual and written concepts in clip.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Dis- entangling visual and written concepts in clip

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.520803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.520803Z digest=sha256:da34ffdad31fe62a55cc255c5aba7d7a67f80cc97ad954ab6d863ad14b0214bc

Observation b5bfdb1b-d947-482d-b912-028e20a2ad05 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.588713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.588713Z digest=sha256:75b127ef8c2a5f4d5490aafefa39461d634168d3b9d671d4268e72cb7004b63b

Observation eb294161-0c90-45ce-94e8-03e4905e4f15 · outbound

This paper cites Infographicvqa.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Infographicvqa

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.635975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.635975Z digest=sha256:21e8147a80db49fef2c48b93a9f7ce725a9a78992150a882ff4d8afc3ab64a62

Observation bae15cad-b497-4d7f-8413-2808442bd699 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.720550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.720550Z digest=sha256:f64ae0ad0470a4e5921f0a02b0db5384b0a7892acffa730d901da1c4a745d6a2

Observation d77e97fb-e423-4b34-b146-e0690934aa75 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.788848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.788848Z digest=sha256:bc4607b539edbaab6d682aa44bcb2c2aab7bd43a8360b961b5bab5717d9105f8

Observation 07a456c5-9456-4519-87d3-5d3b50344b24 · outbound

This paper cites Roadtext- 1k: Text detection & recognition dataset for driving videos.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Roadtext- 1k: Text detection & recognition dataset for driving videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.820636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.820636Z digest=sha256:d0a21f4a5fd1a365c74ec516d1db40280f72e121f6602a67f6c1a39e766b8a01

Observation 86c77226-9f1d-4caa-a05c-7ea2a8226778 · outbound

This paper cites Towards vqa models that can read.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards vqa models that can read

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.889019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.889019Z digest=sha256:a4b33e86f8f26441c1236293f235533691e673a6ab6c60156e52e047d7b7429a

Observation 6111572f-1e48-45e5-aa45-2f7aa93ad864 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.997144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.997144Z digest=sha256:7ac9d0794ccacd48e4261c50ee705e7d6be5bc2e2cfb563481ae46e444827cc4

Observation 4167bf86-f16d-4b8e-811c-15ec23b1db84 · outbound

This paper cites Mtvqa: Benchmarking multilingual text-centric visual question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Mtvqa: Benchmarking multilingual text-centric visual question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.115876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.115876Z digest=sha256:0618fd4da2227e75fe4505d630443fa8d91c5339705e03bebef0df05e8dfcbdc

Observation 1ba91600-0ff3-4896-a8aa-28286e15e3d7 · outbound

This paper cites Reading between the lanes: Text videoqa on the road.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Reading between the lanes: Text videoqa on the road

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.226671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.226671Z digest=sha256:41a2d2aaa2e4b4e78a72f436d4af6dce2e5a57908037ee6882840dc379b854eb

Observation 6223ad18-9750-4a15-9e0c-5d507c4f92a5 · outbound

This paper cites Clip in mirror: Disentangling text from visual images through re- flection.Advances in Neural Information Processing Sys- tems, 37:24523–24546, 2024.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Clip in mirror: Disentangling text from visual images through re- flection.Advances in Neural Information Processing Sys- tems, 37:24523–24546, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.318524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.318524Z digest=sha256:43529708ba3213f35870d2a4d9c9ed56924b803278821f7f00d73686137bd19f

Observation 6a1f5b2d-c3ea-49d6-940f-8a392d447906 · outbound

This paper cites Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.451887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.451887Z digest=sha256:81a7c4066ffbf60cc32661aa00d3ddd14edd0f556e5a554c1c64c7c8eb46ff6c

Observation a25f49f4-1556-4a45-aaf9-ee80c7fb7b5d · outbound

This paper cites What word is written on the sign?.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models What word is written on the sign?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.575532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.575532Z digest=sha256:08a72bf08894e1ed45a42fb103cbcd3064723729e9835afe05985fc8b57ab522

Observation 7d17f6bf-6570-49e6-aca2-ecf4e1f569f4 · outbound

This paper cites an unresolved cited work.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-03T17:32:13.313104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.313104Z digest=sha256:5f03f90c52d87e4b7bc4982f0ddd4c1f61b17a6c73cd02913469040f11141e3a

Pith citing papers

Observation 49c11ea9-d6e0-4f52-8736-23b2fab62fae · inbound

Token-Efficient Multimodal Reasoning via Image Prompt Packaging cites this paper.

Token-Efficient Multimodal Reasoning via Image Prompt Packaging Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:23:07.125862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:52:25.713970Z digest=sha256:8b094165cd592938202b94feae09d8470e6c16fd5182fa34e4843868fbeed103

Observation 351d7fa3-3160-45e9-b9a4-3a94618d81d5 · inbound

Towards Robustness against Typographic Attack with Training-free Concept Localization cites this paper.

Towards Robustness against Typographic Attack with Training-free Concept Localization Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:23:07.125862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:41:21.582578Z digest=sha256:11d222e5163906b4bda183170acd211e585b049c3494210a2f1cda5be955bc31

Observation 68e85a4b-f1a1-4ab9-8edb-9d8d4677b56d · inbound

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models cites this paper.

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T00:13:30.825592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:13:30.825592Z digest=sha256:e25deb5c818f0b48fc8c1001c9acc15bd02771a468e80503fbd1263f95d88da8