Pith. sign in

Paper Citation Record · LEDGER

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.02819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02819 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T06:51:47.981280Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbd41de3-4dce-4f30-89b3-25ce584ebab8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:5337d7f0c78d9248be91272e48645ba63a227febec2f20842d25a925e5af5a7c

Observation 13701d61-e2a9-4161-90b1-07c087281f44 · outbound

This paper cites A survey of man in the middle attacks.IEEE communications surveys & tutorials, 2016.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models A survey of man in the middle attacks.IEEE communications surveys & tutorials, 2016

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:018174c4543fe14c20f4080ce23fdb8103f39c0e441c33a515066199a281bee8

Observation 4198edf6-ccaf-4ec3-9ec3-9000eb436ac4 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi- modality models.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Vlmevalkit: An open-source toolkit for evaluating large multi- modality models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:355c70d4f57a48589b37e3ab8355cb485efa7a390df69a8392957a8226225a1e

Observation f1d08a01-5161-4dd4-99ec-ce2a8733ba55 · outbound

This paper cites On the Robustness of Split Learning against Adversarial Attacks.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models On the Robustness of Split Learning against Adversarial Attacks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:231138d1f5b945d044e36e5245a5ec8633323b1976dcf71f7bcdbd973921a9b3

Observation cc90b245-fd4b-4251-a84b-0b9cc584451a · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Explaining and Harnessing Adversarial Examples

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:bf46cace95ee12a7bcdf91178185dc407f9cc88cc095e3b44eac701d33584937

Observation 00839f20-6653-4b45-aec9-1f27d19caabd · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:c17601de373b4d64fdd850f9d246fd7cb7329fc8276e79c526bd9ec518e167f1

Observation 0a37f00d-ad71-43f8-99d6-5fe879ae7d4f · outbound

This paper cites Hyperion: Low-latency ultra-hd video analytics via collaborative vision transformer inference.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Hyperion: Low-latency ultra-hd video analytics via collaborative vision transformer inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:232e7a011402c8318b78b56bad1815c01f71d9cdb79961dd64c6efb90f34b37d

Observation 440ba0be-6978-4966-a466-44e1c736b82e · outbound

This paper cites Backdoorvlm: A benchmark for backdoor attacks on vision-language models.arXiv preprint arXiv:2511.18921, 2025.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Backdoorvlm: A benchmark for backdoor attacks on vision-language models.arXiv preprint arXiv:2511.18921, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:568e9abd742f37767dc4086fdefcae35e0c8f304d01f16974f962281cba6e6f4

Observation 6aef72d9-2cbc-4180-baa8-6be91531c4b0 · outbound

This paper cites Distributed vlms: Efficient vision-language processing through cloud-edge collabo- ration.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Distributed vlms: Efficient vision-language processing through cloud-edge collabo- ration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:72748a1830bee344987507dc82fd8c3c0678b1d634cbb9368ca427196ff665d7

Observation 06f97519-dd78-49ec-a3cf-2f48a55bd0ae · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:3142d22c04d001db1be9ebbc7c5180f78d0f83710b6f4f45e9423819665d8049

Observation 4fff846b-1e87-4d8a-9f32-fdf25722dcd2 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:c4feebbf78559b403000238b6c4adb1086b0b9297f23b2b5fea79f82be3282dd

Observation c6608ec5-22a9-4180-b9b0-2ff432585b52 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:54259fd2f586685aa14f06a2ccd8c6785d3c31ba63085f0f10866c6188e99558

Observation b1016408-3487-4f47-a519-f22d23f5e031 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:90b5c45335892a2398a49bfe168541cd1da337c168ff219ea4b316a214cca84f

Observation 58ce6770-6ad8-4f14-b6a0-bfe448e71038 · outbound

This paper cites Prompt inference attack on distributed large language model inference frameworks.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Prompt inference attack on distributed large language model inference frameworks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:fa32d17fe46443945990a5cabd5811136b43e626c2e2f831cecaba8228227ae6

Observation 9df952bb-8095-4e5f-af7e-86a1da53896c · outbound

This paper cites Net- gpt: A llm-empowered man-in-the-middle chatbot for unmanned aerial vehicle.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models Net- gpt: A llm-empowered man-in-the-middle chatbot for unmanned aerial vehicle

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:3ebb79e01c30a600f306eb641778be2a44cf2bc0f4d9895579f38a9397736d45

Observation 7805e635-4fd8-41d2-8f52-e562fa6dff41 · outbound

This paper cites edgevlm: Cloud-edge collaborative real-time vlm based on context transfer.arXiv preprint arXiv:2508.12638, 2025.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models edgevlm: Cloud-edge collaborative real-time vlm based on context transfer.arXiv preprint arXiv:2508.12638, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:f2836fe83e3ec64278129f94745431092853ad24526b0f9fb689f298d10b7985

Observation 746fc0ea-f261-4f38-a0a1-7a97b5f82fac · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:88a465b7892c030a16f13466bbc2b8f85d0699dc01e2002f1b48debfe5b1fbcf

Observation 4e14fb2d-6862-4b85-a702-056c58a87ba7 · outbound

This paper cites vision token manipulation attacks on cloud-edge inference of large vision-language models.

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models vision token manipulation attacks on cloud-edge inference of large vision-language models

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-07-12T06:51:47.981280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:51:47.981280Z digest=sha256:eb70c893edf20df5867a962759051f343eb7e02982aa091ce92d843a22a0601d

Pith citing papers

No inbound Pith citation observations are available.