Pith. sign in

Paper Citation Record · LEDGER

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2601.21531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.21531 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:09:25.818449Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T02:57:48.018700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact9
  • verified fuzzy7
  • unresolved7
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ab3c746-1fe9-4f79-9ad6-5ffedd4bbbb0 · outbound

This paper cites Qwen2.5-VL Technical Report.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Qwen2.5-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T15:10:16.503841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:39ff7f4dc7e5ca9de48ed755605baa3945d05918f4dd8a5c40e739d80b5cb194

Observation fa6fc2fc-95c7-4f4f-a501-fdb244ebeeb3 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.499228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:83677022ccfe755b69c1d7348d4b325aa7a9de0ebb95f473cdd1a12b49900491

Observation e49eea97-df3d-4f89-bd72-36e15061fbb6 · outbound

This paper cites Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.531308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:df6282ccb2f3b54f2f3f7a156a28a5c77878098b9e1a6dfa33c0ff0577bd162c

Observation e58a8779-ded4-4e44-bd60-a10d6cf7858d · outbound

This paper cites Attention score is not all you need for token importance indicator in KV cache reduction: Value also matters.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Attention score is not all you need for token importance indicator in KV cache reduction: Value also matters

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.871354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:38fe50c13979470dcdcbcea0ffa125150930412b7746a02566a6b01702fb6a6e

Observation 6b7f4381-3e20-4537-91f1-a7c5dd1323ac · outbound

This paper cites Vision-language-action models for autonomous driving: Past, present, and future.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Vision-language-action models for autonomous driving: Past, present, and future

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.526862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:a1dce3b579ef7d37a311b5588c6826f12186e93431a6fcd7fbe4d5021356df3d

Observation dcff0bad-9c40-47ae-9458-c26b0d523365 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.490773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:e4c9c97bf73e81384a5c9ca0bd93d1821b0ca6ea494550b696fbccfc5f3cf3c2

Observation 3257afd3-1b64-4f74-9467-fd0074446b5d · outbound

This paper cites Veattack: Downstream-agnostic vision encoder attack against large vision language models.arXiv preprint arXiv:2505.17440.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Veattack: Downstream-agnostic vision encoder attack against large vision language models.arXiv preprint arXiv:2505.17440

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.535977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:2ded4f9e569f920f9c367d81a3e62a0c60146288a9c7cb5fd605ec08ad33f9f9

Observation 7a2c1145-ecb1-4827-a1f7-4cc35667a8b4 · outbound

This paper cites GPT-4 Technical Report.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression GPT-4 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.512926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:175ab24dd06c82e65a7b5d979717c2394ee525812ffad1058dcd4016c71c9985

Observation a29f7e4e-2710-4979-b29a-f9bdd6c7b702 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T15:10:16.494936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:c300a1949eaa66a370fa33b9260976637613067c1145b8bec9b9127ed74c0525

Observation d8727c11-feb6-4cbd-bec3-549c82636662 · outbound

This paper cites InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.508674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:f7926ed5bef4435cd579cba83b1012f4c87dfba2aafe94469af1eab87b9f96e7

Observation ad7e7783-28dd-46a0-9d18-a5d965129a95 · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.521668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:a8db0c5be2203f4486ef9418d5016be075ba516d81a847b7126948e3b3a47d11

Observation 0a5eb920-a160-47af-978a-633c1499df6c · outbound

This paper cites DART: Dif- ferentiable dynamic adaptive region tokenizer for vision foundation models.arXiv preprint arXiv:2506.10390.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression DART: Dif- ferentiable dynamic adaptive region tokenizer for vision foundation models.arXiv preprint arXiv:2506.10390

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.517475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:534fe53c6c4581315d6f4c1bb35a7e2182c9ae8f551b94b576cc152023ae099b

Observation a313bdb6-e182-4d5b-adfa-07998bc26ea6 · outbound

This paper cites AppAgent: Multimodal agents as smartphone users.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression AppAgent: Multimodal agents as smartphone users

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.887535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:9b9fb5ab3478cd3368c0906b51dfee26e4eddc47b1ac95e2dd5d25598703b16a

Observation c11f1702-528a-48fd-a5f8-07f9f5dc434e · outbound

This paper cites ② Outer-LLMapproaches perform token selection or aggregationbeforethe main language model computation, treating the LLM as a black box.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression ② Outer-LLMapproaches perform token selection or aggregationbeforethe main language model computation, treating the LLM as a black box

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.868416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:d4c8abdd5e27cf2dc0e564f82812674f0a787808f82e4eccf3908abc5bd74825

Observation d83d6d59-0837-4feb-bb19-383c93bb349a · outbound

This paper cites A” and “S.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression A” and “S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.882328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:7feca883f603d0dbc04116c1c35709e6e05b9dff8a3034a30faa7ea3c2675379

Observation 2be79474-60ac-4954-b1dc-af9c94a63f6c · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.879070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:c85a3efb198151e72d81a6672de1b20816a1f4e357839cf459d96eefe61af2d4

Observation 350ca5d7-013f-4d99-b260-7c4f82e622f4 · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.890142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:b7c3cc86ebf5de9a475f92802a8e92e4f7687c86eb7176d4126b5ef656a03fd2

Observation cd59b761-8129-4356-b5bb-b749c6ac37a7 · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.884875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:7f95115fa5ae5375c8286977c2db70fdc6e5a82e6a6db96359ab68f97d107765

Observation d178112d-52c2-4e50-91f7-d5aff6161765 · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.865543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:27561dc2b032e3d60d5e63902316e449d6cab10eecb825acca82666cb83c3734

Observation 33eeca9c-fc58-40e3-91dc-b32edb71a9a2 · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.892636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:4a0eb1fcef34c3ecc6b63bd5c13d1e68792dfcfa19b1dcb7f57911fe42128e22

Observation c37b96d3-f8e9-498e-8816-9f35a383bbd4 · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.900935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:ba95bd1e4ef97116591dab3cd43642bd2cf3ddeb34908b4dc9f454fd9f266b04

Observation e5f01081-b786-494d-b3ea-d9cd256148e0 · outbound

This paper cites For prompt-diverse tasks such as VQA, VEAttack (Mei et al.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression For prompt-diverse tasks such as VQA, VEAttack (Mei et al

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.895308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:c03263fcd9f459f25b4e4b5ee346efba2232f2678481c0e4e86256e76ec5e462

Observation 7e5a2e41-f3fd-4873-8cbf-f2704946f340 · outbound

This paper cites Compared to open-ended VQA, GQA emphasizes systematic generalization and relational grounding.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Compared to open-ended VQA, GQA emphasizes systematic generalization and relational grounding

Reference 23

Resolution
malformed identifier
raw_fallback, observed 2026-05-21T15:10:16.898296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:096deb5e1ebcf96ba2dadfc28bae9869ce300ea430349f0d25dde6c7c0e8ae6e

Observation 9376ec39-d2e4-4335-b692-55e060fe9cb8 · outbound

This paper cites It utilizes vision-encoder attention maps to estimate token importance, identifying and retaining a compact subset of highly informative tokens.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression It utilizes vision-encoder attention maps to estimate token importance, identifying and retaining a compact subset of highly informative tokens

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.903881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:422c56fc223054210adcf39152543232cfb253431f7b8a87e102bb090be5c3f8

Observation 5aec3930-533a-488f-8b85-9a764d368930 · outbound

This paper cites Importantly, most non-zero λ settings are comparable to or better thanλ=0, indicating that incorporating RDA is generally beneficial even without delicate tuning.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Importantly, most non-zero λ settings are comparable to or better thanλ=0, indicating that incorporating RDA is generally beneficial even without delicate tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T15:10:16.873954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:a12f800cb7c25936e821a4dbdd414220e90912800605efcd06d52685687653ca

Observation a1b9f1fe-a914-4f29-ad06-ba85004cd42c · outbound

This paper cites peakiness.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression peakiness

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-05-21T15:10:16.906565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:1a06dc9bab62772b7ef9ddbfc66a0327292470f99a662a278c79ff25f1deef3d

Observation c8cab823-9227-4959-8360-1aaadf0b426d · outbound

This paper cites an unresolved cited work.

On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-21T15:10:16.876443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T15:09:25.818449Z digest=sha256:236304565b3c694657ec86431786b682c58b9ef3a712b261b49ab914d1069a11

Pith citing papers

Observation 59e09af0-3610-4134-94b5-221551f39e13 · inbound

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models cites this paper.

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:48.018700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:48.018700Z digest=sha256:12a39ba47390ba54e2603b88fc429123ce897f910118562e1a5a55268a4ce89d

Observation 27a920ce-8ccf-4bf7-aab9-20ad969a8772 · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.695997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.695997Z digest=sha256:23e846cc3164b7e807a2cab95473ed8a5de85ddf4aad527958525d7231831dbc