Pith. sign in

Paper Citation Record · LEDGER

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.00817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00817 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:12:03.631575Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab9a3c74-bce9-475f-beff-8717619f3c63 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.460828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.460828Z digest=sha256:d2263eb4bb7719d671b7e050e9b87973836dd859e26fcb2a5b1feda480f32f9c

Observation e405321d-415a-415e-a4c0-01d13ab22267 · outbound

This paper cites GPT-4 Technical Report.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.466590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.466590Z digest=sha256:f9180a1ddb39636e7bf6fda06822f5fd686b4242d6e8797eb9621b955eacd3c2

Observation d3e09ca4-f9bf-44ba-b2b3-5280009a6fc9 · outbound

This paper cites Qwen2.5-VL Technical Report.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.471197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.471197Z digest=sha256:50c2bdcfc69399e20d74ec004980ebc60c5ea18809fdad4e7f17b023e22b59d9

Observation fd4f7cb2-84d9-42f7-9d58-9953771d356d · outbound

This paper cites Rethinking model ensemble in transfer-based adversarial attacks.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Rethinking model ensemble in transfer-based adversarial attacks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.301035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.475888Z digest=sha256:253acf6b2ffaa779b00a9e8b7b38d9f2ec7df0e5d45ef4909d2b5c29539b372a

Observation 7be22a30-c6aa-42a1-8f01-36546015233d · outbound

This paper cites Gcma: Generative cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Gcma: Generative cross-modal transferable adversarial attacks from images to videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.286592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.481152Z digest=sha256:3805027d25ffc2b3f591b3bfb749f8199f67c0067c86e457de439e12baec8a96

Observation 3d11eb4b-aecf-4563-88a8-77e0ac458ce8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.486303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.486303Z digest=sha256:605be41bd270513f0a7541a6fd228a7c645d854faee6ef5e45be35e938b7d6f9

Observation f7554876-00cb-4579-a646-9fbd67b13ab0 · outbound

This paper cites Parseval networks: Improving robustness to adversarial examples.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Parseval networks: Improving robustness to adversarial examples

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.272326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.491971Z digest=sha256:17f16b67bc556285f4b2ed9f957672e5dffdef788a12fcca14468d5b799ff4b3

Observation 128bc05e-11e6-426c-8930-b4fd367ac9cc · outbound

This paper cites One perturbation is enough: On generating universal adversarial perturbations against vision- language pre-training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs One perturbation is enough: On generating universal adversarial perturbations against vision- language pre-training models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.496616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.496616Z digest=sha256:c99aa74a9262c8624b6db9d32f212d794f2bc542add2b46b20c69f1b3068fafe

Observation 8b69b999-4419-4bb0-b762-2b015be3263a · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.500911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.500911Z digest=sha256:e33948cf864d05948e12c80d76da633a7181ad860acdab05387b535ad46ed7e6

Observation 3081277c-1838-442e-8114-028a50bc5332 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.505036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.505036Z digest=sha256:7adf42909192ca376556bc5956e6ebe0d6dc1b4a1735e160f1bcb512d42ff7d6

Observation f83f7fd2-5384-4c96-bf26-f92b75ca8ffc · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.246709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.509895Z digest=sha256:d8490fd218dc82b17df307237a941ec4f126b0000fecf722319bc6ee4b59ab58

Observation 40c594a2-e467-4b93-967c-63785b6c8830 · outbound

This paper cites Retome-va: Recursive token merging for video diffusion-based unrestricted adversarial attack.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Retome-va: Recursive token merging for video diffusion-based unrestricted adversarial attack

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.230595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.515186Z digest=sha256:03c2611f19997217db1cc81d162cbb21e26c00868f70844f62dd02d14a3bd846

Observation 82d0fffc-fb17-46c4-b2ae-dc9427800d6d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.519327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.519327Z digest=sha256:946f7b68f3e5974a92fa1a72429b970e0a5a54d0539fbf9d5d1c60d2d46feecd

Observation 48a29949-3fa6-4845-b58a-894c68812495 · outbound

This paper cites X-transfer attacks: Towards super transferable adversarial attacks on clip.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs X-transfer attacks: Towards super transferable adversarial attacks on clip

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.215116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.523977Z digest=sha256:2a114c2a0f21e02a9e287034d643468ea7f032b4af775a1505ac5b8a712efea4

Observation 4c32859d-f05e-4ad6-94f8-674b6719cd98 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.528213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.528213Z digest=sha256:fbd56243e2b33de8599c7d4b4e8e16fa35b353f641fb3fb93e0b32c6e9c20d92

Observation eead32a8-c627-48ac-b545-2cbbaea125dc · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.532425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.532425Z digest=sha256:383d1d54aa23b8259e1a293fef921a3a63d32f1c2867ddc6fb184ae41f8f31cb

Observation ebd10513-2b87-41c1-99b6-38a145df829f · outbound

This paper cites Vila: On pre-training for visual language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Vila: On pre-training for visual language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.536504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.536504Z digest=sha256:d92c6ff5b2e8c807d5e2fe4dfa244cde5ebd6e72db89d03e49e0ce18fbc00aed

Observation 51a9d344-5cb7-440f-ac7d-d88c4ffc0315 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.540655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.540655Z digest=sha256:e0ce646f7188e87875d444d83f5f23cf75e97395edc30f7a3080a6d5c5ab64f2

Observation a3ea58a3-74c1-4ee8-83a9-81ddaf24e3de · outbound

This paper cites Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.179120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.544625Z digest=sha256:dde41ee79e42dfdf416576c374f86b08650be220bfa94af0fcf11cbd1762dc0f

Observation b6a1ca85-0bdb-487b-98a7-73c187545f8e · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.549503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.549503Z digest=sha256:36c0b2b23a9358f20925b2d8d26e917ed2febb30a4495616e0c3aa748112147f

Observation 93f61dd9-ccb2-44ed-815b-7fdc3892785b · outbound

This paper cites Univer- sal adversarial perturbations.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Univer- sal adversarial perturbations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.163647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.554064Z digest=sha256:41519f7a978b5bad378c1f9b8d436b92cf0de4f85e7fe3966d103286cda710dc

Observation ab3872f9-c4e8-4123-859f-9d5a7f70b926 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.558815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.558815Z digest=sha256:9fdc4d5bcb162abcc77c70bb66ca14131379d66c58c92c9261f2ff2c7a0b070f

Observation 7c9e3e32-446e-4b8d-a0f7-e5a836791a7f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs U-net: Convolutional networks for biomedical image segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.563178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.563178Z digest=sha256:abd6291eba51235a468e30271a260dfc3aeb3488e2dcad4835995b13df9e7fb5

Observation 47612cf9-a8db-4d85-a995-80f50e49d618 · outbound

This paper cites Imagenet large scale visual recognition challenge.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Imagenet large scale visual recognition challenge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.567331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.567331Z digest=sha256:65b4be29c0b0319d320f285fc85ff4b81612705fd12b8d1a96888e8986d4ab10

Observation 7a0c61b1-3f95-425e-b8ce-27c69effc377 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.571419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.571419Z digest=sha256:2e2775794b0446baf76e580bd296fda4ba654c93c1b2b9b293c037d21310fac0

Observation cc980697-0bdf-4e78-936d-0e3f2ad73a81 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.575394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.575394Z digest=sha256:d64d5637f2e3c0ef003fcfd8f5d51c908294f9a4152d518a32e7291aad016785

Observation 77b55267-56ec-4a69-a202-3b0d4b9db60d · outbound

This paper cites Heuristic black-box adversarial attacks on video recognition models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Heuristic black-box adversarial attacks on video recognition models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.117058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.579468Z digest=sha256:36abc7893b66e104d9025e5a74b3396d97a98f4e18ae96f40d56154544819caa

Observation 90334929-223c-417a-ab9e-70774e542d45 · outbound

This paper cites Boosting the transferability of video adversarial examples via temporal translation.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Boosting the transferability of video adversarial examples via temporal translation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.101434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.583937Z digest=sha256:16ecca9be6a4a894ed24ceb1fb16b1557f55d1f2ea760e24c6a9b93f2a7f35e8

Observation c90db321-d442-49cc-a7b6-74635529a31c · outbound

This paper cites Cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Cross-modal transferable adversarial attacks from images to videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.084862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.587903Z digest=sha256:6c7f5b56fee36edf9a923cde91598a125d1acb10a1fc9aa61e5315431374837c

Observation 11506826-ea51-45e5-baa0-929ff5c98432 · outbound

This paper cites Adaptive temporal grouping for black-box adversarial attacks on videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Adaptive temporal grouping for black-box adversarial attacks on videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.069512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.592180Z digest=sha256:b60e066a62aba1fc71d83d9241731868f4e094c9aab9106d1643384ae4d4f270

Observation 63950faf-6718-4384-b2d9-e190306a711b · outbound

This paper cites Adaptive cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Adaptive cross-modal transferable adversarial attacks from images to videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.054373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.596832Z digest=sha256:f2e1ac2653662908b3f6bd7ef23f3dc2c5b2d423841a4194e16e0068bac68ed2

Observation ef1ee7dd-1f2b-495d-9dc3-6daabfc7e3d5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.601085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.601085Z digest=sha256:08a1d3382c03feb09d91ac261922bcc06d11aa2a0adfa65837d049b500032ce2

Observation d2c84931-efae-493e-b958-0a6fe724b5f0 · outbound

This paper cites Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.038244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.605433Z digest=sha256:b3437eef83be53b6c28300f295f001bb0664c2af8640736ece82aa411fc7e928

Observation cafb8a2f-740b-498a-8a6a-66360466bc4d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.609942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.609942Z digest=sha256:83259661ea928d98e2373778392b1e5174a9a37eb3f6f5a6c8dd468e672d3058

Observation a12e187f-292a-4365-a725-f186be113542 · outbound

This paper cites Towards adversarial attack on vision-language pre- training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Towards adversarial attack on vision-language pre- training models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.021487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.614224Z digest=sha256:2401d334dda656d3f80ce5082a368a3bbdd8801b388d7da5eb4b359e8f774d81

Observation 3917b869-8038-4e06-92c5-d231c4ea92f5 · outbound

This paper cites Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.618948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.618948Z digest=sha256:ad17679efc2b9be0e8a4b3383e17f417cfdf265bc9ffdd296042e4f629cf235f

Observation b3f8ab5d-e1e0-4dc3-a967-4662dd2d825b · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.623119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.623119Z digest=sha256:67fffb080dd468277d0ef2d02a1bc0f20f9bbdcea0be1474712b11766598bd11

Observation 0dcb5446-bb7b-470d-9d49-d67f429a301d · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs On evaluating adversarial robustness of large vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:03.985237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.627271Z digest=sha256:fc465134b8e15ff4f95ed97f6b7d658d2ec9c47c7f2b5245f4b602e2ccd4b29a

Observation 79e3a4b7-3877-468d-9264-e3da8c4e4bf2 · outbound

This paper cites Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:03.968998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:12:03.631575Z digest=sha256:a747467f6652f524b7a65ed0e7da0ca463602d519b03d28d8d142ff3e11a1456

Pith citing papers

No inbound Pith citation observations are available.