Pith. sign in

Paper Citation Record · LEDGER

Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2401.11170.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11170 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:46:49.497562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:56:14.839778Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e49005a2-0855-4804-b1a8-fc25f761621d · inbound

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments cites this paper.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.365436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.365436Z digest=sha256:e00e635ab8377dbb1fba1c0bb00b1775f1f7da872592cfd39ddf673b6307df2c

Observation 555c05e6-2912-4c4c-82e6-9d5d054b1f52 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.098397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.098397Z digest=sha256:caa061c820315258ec591f00d68d0d4f4434612e3c36659342a90fe000356da6

Observation 466f9a8a-3cd4-4e70-9995-88a966088bb2 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.271117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.271117Z digest=sha256:d1738de705c2b0c50e6b3c3d9e3596046cb07e1ec53f240d0d4c629e30054fc9

Observation 24064e97-9938-4eee-99d9-168219432228 · inbound

Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving cites this paper.

Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:46:49.497562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:46:49.497562Z digest=sha256:d563e714d0474ba1e77f933c138571b1ed2c51b63aa103388a6178fabb48988f

Observation 2d5dcb8c-a839-481d-aa6c-5913328e0b07 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:05.788072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:05.788072Z digest=sha256:05b14f5313cf0aca37d160122708807854ad253d16c6a0c129443f1e78c6b736

Observation e0be8add-2c6b-4905-a3d4-34208dc76d5d · inbound

Throttling Web Agents Using Reasoning Gates cites this paper.

Throttling Web Agents Using Reasoning Gates Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:03.958734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:03.958734Z digest=sha256:ce118d53f7e0a860d91323b56cc0ed07529baaccba7c10e93bea846292abfcee

Observation 571690aa-df5c-4710-af44-476268eb5fcd · inbound

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems cites this paper.

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:58:53.681946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:52:11.018335Z digest=sha256:3281ee8fc4cf2d68235890d6c568c0d528e22944c246e4f1f4ba52be113ad7a1

Observation 720f61de-d195-4e7f-b8c6-198ec751f9f9 · inbound

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces cites this paper.

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:56:14.841505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T12:55:34.631248Z digest=sha256:fd690b9a404f26d5a34b7cc2cbc367ecbac3c1506adc7ee4898667b3eb8c5684