Pith. sign in

Paper Citation Record · LEDGER

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2512.11899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.11899 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:32:16.575532Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:13:30.825592Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b62e09d8-2198-4298-a28d-0503e3fc879f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.163631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.163631Z digest=sha256:01308e3637d4413c99af6e613c61c51c2e6cde539b7ff4f3657223adc19b045d

Observation 650861ec-712b-40e4-a232-f85000d8a688 · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Defense-prefix for pre- venting typographic attacks on clip

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.213168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.213168Z digest=sha256:9ea3135ac51c5ec08523bca600ae9be51eda7d506943db6efb3ccd64514f16ab

Observation 131f41b8-e007-4814-97dc-fba63b42e6ef · outbound

This paper cites Qwen Technical Report.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.292022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.292022Z digest=sha256:5bb0348010fb7017ebca207edf8fc20e19d4887a37a70f2760baa307838cd47d

Observation 138f1e21-53c7-4995-93de-e40075da3863 · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vizwiz: nearly real-time answers to visual questions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.427960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.427960Z digest=sha256:745309420e63e48986be1ba47109263cea8c9f67fa7397892b36105984514c34

Observation 51c7affb-f98b-421d-ad81-b510ef3fb9b7 · outbound

This paper cites Scene text visual question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scene text visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.572998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.572998Z digest=sha256:d0a91083f34fa22d218f430c7be43f9d568592b1541e90e5168092eaa22d43f7

Observation 21d4de9c-bfb2-4554-9f91-cb98fd1187ee · outbound

This paper cites Adversarial Patch.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Adversarial Patch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.815967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.815967Z digest=sha256:119861433e7c7d3e18886c5c215ab462aec97531a3e9e7fdf13dccf4d8106011

Observation 0973c8b4-144f-46c8-ad17-d5e8249779e0 · outbound

This paper cites Scenetap: Scene- coherent typographic adversarial planner against vision- language models in real-world environments.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Scenetap: Scene- coherent typographic adversarial planner against vision- language models in real-world environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:12.977043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:12.977043Z digest=sha256:1acc1eca3f33ab3c1f4555f369b55eea373e865245bc6811f7e89108e8fb1028

Observation 2d0b8aa7-d88f-4a30-82ac-00aab4f201a5 · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.179174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.179174Z digest=sha256:6e26e9bfbec83caf6947d2a248c56b08841284e21056cd1456ed7e4acdb191d5

Observation 2968110c-c787-4671-a785-fc8a5fa9e2cb · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.529564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.529564Z digest=sha256:0ea18667ab037e41dbd3962297d5921bf4ab619393fee8230cf69d37f8b801a8

Observation 80b386fd-4eb8-40e8-83ab-911388c9cd63 · outbound

This paper cites Sari sandbox: A virtual retail store environment for embodied ai agents.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Sari sandbox: A virtual retail store environment for embodied ai agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.735538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.735538Z digest=sha256:9a7b50528c9962b045d04d36ca1576931a701c86d5b95993275246d359d03535

Observation b7b9e79c-f634-44c5-b915-a523806e4b02 · outbound

This paper cites Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.899059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.899059Z digest=sha256:e888073976e7e2ed872c8a579a29f70290273a85c9cbddef277bf01e7fcc1aaa

Observation 80961c2b-d8f3-4c6f-a3b3-f968d5e168ce · outbound

This paper cites Multimodal neurons in artificial neural networks.Dis- till, 6(3):e30, 2021.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Multimodal neurons in artificial neural networks.Dis- till, 6(3):e30, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.046639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.046639Z digest=sha256:eecad814d6303d3804087b48b2f8f559e98211de471d8420029cc29321505deb

Observation 3e65586c-d3e4-4fa0-b967-4987a0fe5456 · outbound

This paper cites Explaining and harnessing adversarial examples.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Explaining and harnessing adversarial examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.141883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.141883Z digest=sha256:f678196cf63ac6b575661e684075c05911a74e70a9688b923ad93fac6d8eefc2

Observation 7e3ace6b-f35b-44bd-9079-c44208936c5b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.202228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.202228Z digest=sha256:0bf83b875d71027de5be2ada4542a2f1095c6949b14e8e087a3ed4b2d653171b

Observation 73c0390c-2e70-4eb3-9100-82770a31394c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.274366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.274366Z digest=sha256:c712bdbc0a1694b8d4c5bff66432ca16c0c87b4b14c95f6e79f98caf9eb8bb7e

Observation 7ce04df3-e2d0-4fe9-9530-2c399b60e5f0 · outbound

This paper cites Towards mech- anistic defenses against typographic attacks in clip.arXiv preprint arXiv:2508.20570, 2025.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards mech- anistic defenses against typographic attacks in clip.arXiv preprint arXiv:2508.20570, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.328818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.328818Z digest=sha256:29eb783b128282721ac0fbca21c6fcb11f3f2538404d5a3e502c9a7ffb5a6215

Observation 6b47a868-7136-425e-9eb9-6fb24fb95377 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.406664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.406664Z digest=sha256:66a3fd62503da645f86b6066bfed7efcfb15c87bd37bd8d71a25aba08f8ee051

Observation b1192ff8-7952-446d-8337-ed1ac55be160 · outbound

This paper cites A diagram is worth a dozen images.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models A diagram is worth a dozen images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.535410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.535410Z digest=sha256:71969b1272e779f50e06540ab11f1a3e37fbcb1f678e387c8b404ddeddd90522

Observation f2d9b2d6-f613-4a79-8c36-a48e44a29391 · outbound

This paper cites Ocr-free document understanding transformer.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ocr-free document understanding transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.609358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.609358Z digest=sha256:cc6753fa0dcec0bcf01b115fc609a6b7cba2c574fab0f13894e4ae4306df8051

Observation d04dd6a8-432c-4a20-aeb2-84c187b08f7d · outbound

This paper cites an unresolved cited work.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.658093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.658093Z digest=sha256:f12f2827c306034998a6e2b14368c52fe04d251ac30895182f190c214872a64d

Observation e9afa5a5-68a3-44c6-b85e-50b7316d86ad · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.738428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.738428Z digest=sha256:6d712a27d541873283f609e183811d1967f16e7dd97a1125265ddaec58e4413d

Observation ccf4c22d-5771-4006-911d-cdf510307cdc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.868129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.868129Z digest=sha256:6435e6ae4519b067f01358da3f5ab31470b7064bb0897f73d5f6c9d1e0f8b61d

Observation f520a9a5-0818-40f4-ae91-33b9da07f6d8 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.013156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.013156Z digest=sha256:39cb3b938acddfe14f08acc1bc5ced0c39e4440a3f35b9c4e6db19cc525196f1

Observation 06937145-2ced-4d36-ac1f-a495c0bcef7f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.096645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.096645Z digest=sha256:63bc71cdf02a3d0789d1783ac2de7bfad2a332c73295e4936c5af8fb7fbc7a12

Observation 727b6360-920e-4756-a00a-d1e014e06490 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards deep learning models resistant to adversarial attacks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.180825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.180825Z digest=sha256:48f5ba717d22a725239228f707c161b444edc3ceacb266a7f6e926ecc65ed220

Observation 3519e281-73ed-40fa-8c2f-f5447c284c62 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.264270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.264270Z digest=sha256:a3de73e258bb59be0925d340d199dc87a625430148b02660d47d56b6744a1f77

Observation b1cd3e02-d5c3-4fbc-932f-2a8abf1f5f63 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.357052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.357052Z digest=sha256:f0c538df9f198be0d1eed40656308bcf1ab8300a6b85bebc3b1b2ce856eb8324

Observation 45cff08e-5b1f-47a5-8fd5-5a3302427dfb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.434754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.434754Z digest=sha256:fcd16c5ad8e243f5d773609ddaafb8e987e81beb9b3520853ca550c91ff3032a

Observation 4d02b1dd-17a3-44fe-96c5-729d8eed5086 · outbound

This paper cites Dis- entangling visual and written concepts in clip.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Dis- entangling visual and written concepts in clip

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.520803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.520803Z digest=sha256:fc43f63ef50fdd0dc8a79da3fb7252430a716052ed17c035c53920ab1c030c90

Observation b5bfdb1b-d947-482d-b912-028e20a2ad05 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.588713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.588713Z digest=sha256:6de2f7f56299f5b3a1e7886bd4a070d9012e1c1ed8fcaf1d1b8906e55045d619

Observation eb294161-0c90-45ce-94e8-03e4905e4f15 · outbound

This paper cites Infographicvqa.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Infographicvqa

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.635975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.635975Z digest=sha256:6fe39dbc05bf68dbcfc22a339dc65932c3628ebeaa47072a06cca111aae06009

Observation bae15cad-b497-4d7f-8413-2808442bd699 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.720550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.720550Z digest=sha256:78c9211d52d27d3f74d86d5d899d5b62a99bd0420ee637f56cea7154689b098f

Observation d77e97fb-e423-4b34-b146-e0690934aa75 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.788848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.788848Z digest=sha256:54e7dba5bf21d64e0629e1df11824c91fed7b813106e56c845953914bdc26230

Observation 07a456c5-9456-4519-87d3-5d3b50344b24 · outbound

This paper cites Roadtext- 1k: Text detection & recognition dataset for driving videos.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Roadtext- 1k: Text detection & recognition dataset for driving videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.820636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.820636Z digest=sha256:a8f7d119cbd5b6fa92706a95cee9b8db448beea5fc5ee5f1beb09f942771bab4

Observation 86c77226-9f1d-4caa-a05c-7ea2a8226778 · outbound

This paper cites Towards vqa models that can read.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Towards vqa models that can read

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.889019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.889019Z digest=sha256:0173808961e18a96d95163ffa8cddc8cfacb65e7428bd04344c7eba3228983f3

Observation 6111572f-1e48-45e5-aa45-2f7aa93ad864 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.997144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.997144Z digest=sha256:792440ef5f968d834bf7ec06282a9c4a1da953cc890627a1d1b054943259e5e8

Observation 4167bf86-f16d-4b8e-811c-15ec23b1db84 · outbound

This paper cites Mtvqa: Benchmarking multilingual text-centric visual question answering.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Mtvqa: Benchmarking multilingual text-centric visual question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.115876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.115876Z digest=sha256:b54070bc9d51b5ec5df066a758747036944d851b9c38716e1c1abd943ac207dd

Observation 1ba91600-0ff3-4896-a8aa-28286e15e3d7 · outbound

This paper cites Reading between the lanes: Text videoqa on the road.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Reading between the lanes: Text videoqa on the road

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.226671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.226671Z digest=sha256:1e39cbc15129559fb310f83f25147fa59d02279ad510e21962a93d8f773a633f

Observation 6223ad18-9750-4a15-9e0c-5d507c4f92a5 · outbound

This paper cites Clip in mirror: Disentangling text from visual images through re- flection.Advances in Neural Information Processing Sys- tems, 37:24523–24546, 2024.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Clip in mirror: Disentangling text from visual images through re- flection.Advances in Neural Information Processing Sys- tems, 37:24523–24546, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.318524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.318524Z digest=sha256:3921aadb6cfcc6697e8a4d1820d3a40833466f2fcdccea6954a1981e8da4761a

Observation 6a1f5b2d-c3ea-49d6-940f-8a392d447906 · outbound

This paper cites Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.451887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.451887Z digest=sha256:9840adb6c7485d877bd813b4282ec39de2e7b22e95e4f7e96708824867adf620

Observation a25f49f4-1556-4a45-aaf9-ee80c7fb7b5d · outbound

This paper cites What word is written on the sign?.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models What word is written on the sign?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:16.575532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:16.575532Z digest=sha256:aff8fe2f3ba7f7d568b3080f671bef175987b16dceaf68fa23ed3b0e2b63da5d

Observation 7d17f6bf-6570-49e6-aca2-ecf4e1f569f4 · outbound

This paper cites an unresolved cited work.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-03T17:32:13.313104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.313104Z digest=sha256:d7fae043b3a872966dd74545c1eeb57ee35c05218ee98e387b061b088772bbe1

Pith citing papers

Observation 49c11ea9-d6e0-4f52-8736-23b2fab62fae · inbound

Token-Efficient Multimodal Reasoning via Image Prompt Packaging cites this paper.

Token-Efficient Multimodal Reasoning via Image Prompt Packaging Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:23:07.125862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T21:52:25.713970Z digest=sha256:f7bcc3b5a4185ad0f36c4ee3415a406e413eee8b689b5cafb6fef7aad5e5d51e

Observation 351d7fa3-3160-45e9-b9a4-3a94618d81d5 · inbound

Towards Robustness against Typographic Attack with Training-free Concept Localization cites this paper.

Towards Robustness against Typographic Attack with Training-free Concept Localization Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:23:07.125862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T14:41:21.582578Z digest=sha256:04ced17e385de3fb9c05354c82da89c612b7ec1e180a48596681a6ad277913b7

Observation 68e85a4b-f1a1-4ab9-8edb-9d8d4677b56d · inbound

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models cites this paper.

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T00:13:30.825592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:13:30.825592Z digest=sha256:41801f6e5d65f6e6db0766a83e2d1a316a1602113547088bb24eeadf004602e6