Pith. sign in

Paper Citation Record · LEDGER

Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.00626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00626 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.061030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:58:53.646539Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d902c91b-1bc2-4c0f-bf22-4402e9b5bbd7 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 294

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:33.969912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:9381f8b6430eff5e902635865ed611523ac9ddad8d657190b2a607f5e754cae5

Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · inbound

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation cites this paper.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.061030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.061030Z digest=sha256:a747c7331e05aec440532db7e7f343145861eb6356874bad002c1070369f1a24

Observation b64cd411-d2df-484b-b9a4-a53bd96f9064 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:05.912718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:05.912718Z digest=sha256:9a4b5887b0f528472cf7fe49ccce4c521fddcf8ca2c60cec0ef7d1a2ea29b76c

Observation f6be0366-bb58-4df4-bfc0-216ecec2d5fe · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.848984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.848984Z digest=sha256:33980a1293317b48350ec3305fdb9df41b4eb006932f74548dd984df05785f0a

Observation ae799f24-33b2-466b-b074-5fb104fdb242 · inbound

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models cites this paper.

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:02:39.623038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:02:39.623038Z digest=sha256:030f398c58ea183508486fb0c917f17a32ae4342127d7eca5132866975f028ae

Observation bae15cad-b497-4d7f-8413-2808442bd699 · inbound

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models cites this paper.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:15.720550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:15.720550Z digest=sha256:84e6e9e446cb0ceaf625a8c98d6c2fce952707bed52309d99fcf20707b28e63a

Observation bc7b464c-7e41-40f4-9d14-a4616f6bb9a7 · inbound

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning cites this paper.

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:38:02.892843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T17:34:10.089555Z digest=sha256:24ea737bbcc7db04803c23368d7c39d6e414ca05ace9233dbac29663f380d139

Observation 4e6ee125-5564-4257-9c71-d84db55ad4d8 · inbound

SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models cites this paper.

SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:57:27.669590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T06:56:10.053418Z digest=sha256:52032026857b17c240b5d8dcdc9d8a48a6dbedd6f159f0f62b72f6eda99b62af

Observation 34b715b1-8df7-46f1-bc91-fa514fce4b35 · inbound

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems cites this paper.

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:58:53.648368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T19:52:11.018335Z digest=sha256:0afab2c7ebe8149498bb7c7e6aa4296147b850df07b5577717afc2c110034a10

Observation 88d15d79-a94a-4377-8555-f4b57c9b6586 · inbound

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices cites this paper.

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T13:02:42.673767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:02:42.673767Z digest=sha256:c69e9db99581b5f2e0055b3da6d8dfd29beaf2705177b4ecf0bc223bf114bf6a