Pith. sign in

Paper Citation Record · LEDGER

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.00817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00817 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:12:03.631575Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab9a3c74-bce9-475f-beff-8717619f3c63 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.460828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.460828Z digest=sha256:d2263eb4bb7719d671b7e050e9b87973836dd859e26fcb2a5b1feda480f32f9c

Observation e405321d-415a-415e-a4c0-01d13ab22267 · outbound

This paper cites GPT-4 Technical Report.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.466590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.466590Z digest=sha256:f9180a1ddb39636e7bf6fda06822f5fd686b4242d6e8797eb9621b955eacd3c2

Observation d3e09ca4-f9bf-44ba-b2b3-5280009a6fc9 · outbound

This paper cites Qwen2.5-VL Technical Report.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.471197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.471197Z digest=sha256:50c2bdcfc69399e20d74ec004980ebc60c5ea18809fdad4e7f17b023e22b59d9

Observation fd4f7cb2-84d9-42f7-9d58-9953771d356d · outbound

This paper cites Rethinking model ensemble in transfer-based adversarial attacks.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Rethinking model ensemble in transfer-based adversarial attacks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.301035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.475888Z digest=sha256:b0c9e3c9833e02091a7c10acd86681b44b784314d7fde6f7eb5e6adb42e16959

Observation 7be22a30-c6aa-42a1-8f01-36546015233d · outbound

This paper cites Gcma: Generative cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Gcma: Generative cross-modal transferable adversarial attacks from images to videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.286592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.481152Z digest=sha256:984e617ff7251467b1d2e856a6ed43f97a861bfc8c824081d24e2a9ed2bf5ddd

Observation 3d11eb4b-aecf-4563-88a8-77e0ac458ce8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.486303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.486303Z digest=sha256:605be41bd270513f0a7541a6fd228a7c645d854faee6ef5e45be35e938b7d6f9

Observation f7554876-00cb-4579-a646-9fbd67b13ab0 · outbound

This paper cites Parseval networks: Improving robustness to adversarial examples.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Parseval networks: Improving robustness to adversarial examples

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.272326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.491971Z digest=sha256:bd27dc8ffda232b36589a704c74071c21515e7f69a7ef4ecb6608ad72db7531c

Observation 128bc05e-11e6-426c-8930-b4fd367ac9cc · outbound

This paper cites One perturbation is enough: On generating universal adversarial perturbations against vision- language pre-training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs One perturbation is enough: On generating universal adversarial perturbations against vision- language pre-training models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.496616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.496616Z digest=sha256:c99aa74a9262c8624b6db9d32f212d794f2bc542add2b46b20c69f1b3068fafe

Observation 8b69b999-4419-4bb0-b762-2b015be3263a · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.500911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.500911Z digest=sha256:e33948cf864d05948e12c80d76da633a7181ad860acdab05387b535ad46ed7e6

Observation 3081277c-1838-442e-8114-028a50bc5332 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.505036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.505036Z digest=sha256:7adf42909192ca376556bc5956e6ebe0d6dc1b4a1735e160f1bcb512d42ff7d6

Observation f83f7fd2-5384-4c96-bf26-f92b75ca8ffc · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.246709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.509895Z digest=sha256:bfe25b9eaa99d8919a1c28d6b59085c505bade0599b7ee5c542177591e1b388c

Observation 40c594a2-e467-4b93-967c-63785b6c8830 · outbound

This paper cites Retome-va: Recursive token merging for video diffusion-based unrestricted adversarial attack.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Retome-va: Recursive token merging for video diffusion-based unrestricted adversarial attack

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.230595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.515186Z digest=sha256:6f13fa6da1740fdf0ee88ea06086f062deb1125810694efa13a5cfad616d55f9

Observation 82d0fffc-fb17-46c4-b2ae-dc9427800d6d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.519327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.519327Z digest=sha256:946f7b68f3e5974a92fa1a72429b970e0a5a54d0539fbf9d5d1c60d2d46feecd

Observation 48a29949-3fa6-4845-b58a-894c68812495 · outbound

This paper cites X-transfer attacks: Towards super transferable adversarial attacks on clip.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs X-transfer attacks: Towards super transferable adversarial attacks on clip

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.215116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.523977Z digest=sha256:795bdb512f1a7f14252fdef415a67baa156a28c1a3748728bb95b03ca9547e3a

Observation 4c32859d-f05e-4ad6-94f8-674b6719cd98 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.528213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.528213Z digest=sha256:fbd56243e2b33de8599c7d4b4e8e16fa35b353f641fb3fb93e0b32c6e9c20d92

Observation eead32a8-c627-48ac-b545-2cbbaea125dc · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.532425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.532425Z digest=sha256:383d1d54aa23b8259e1a293fef921a3a63d32f1c2867ddc6fb184ae41f8f31cb

Observation ebd10513-2b87-41c1-99b6-38a145df829f · outbound

This paper cites Vila: On pre-training for visual language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Vila: On pre-training for visual language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.536504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.536504Z digest=sha256:d92c6ff5b2e8c807d5e2fe4dfa244cde5ebd6e72db89d03e49e0ce18fbc00aed

Observation 51a9d344-5cb7-440f-ac7d-d88c4ffc0315 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.540655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.540655Z digest=sha256:e0ce646f7188e87875d444d83f5f23cf75e97395edc30f7a3080a6d5c5ab64f2

Observation a3ea58a3-74c1-4ee8-83a9-81ddaf24e3de · outbound

This paper cites Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.179120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.544625Z digest=sha256:d800d935f0c074c3ebf3bbb2ff7f2cfa54c245ef9046393ea8ea6ca31d49d3e9

Observation b6a1ca85-0bdb-487b-98a7-73c187545f8e · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.549503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.549503Z digest=sha256:36c0b2b23a9358f20925b2d8d26e917ed2febb30a4495616e0c3aa748112147f

Observation 93f61dd9-ccb2-44ed-815b-7fdc3892785b · outbound

This paper cites Univer- sal adversarial perturbations.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Univer- sal adversarial perturbations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.163647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.554064Z digest=sha256:d161891e0a2c06c6aea873c995205b9cb31a95e217690b608d00685b1ced290d

Observation ab3872f9-c4e8-4123-859f-9d5a7f70b926 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.558815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.558815Z digest=sha256:9fdc4d5bcb162abcc77c70bb66ca14131379d66c58c92c9261f2ff2c7a0b070f

Observation 7c9e3e32-446e-4b8d-a0f7-e5a836791a7f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs U-net: Convolutional networks for biomedical image segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.563178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.563178Z digest=sha256:abd6291eba51235a468e30271a260dfc3aeb3488e2dcad4835995b13df9e7fb5

Observation 47612cf9-a8db-4d85-a995-80f50e49d618 · outbound

This paper cites Imagenet large scale visual recognition challenge.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Imagenet large scale visual recognition challenge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.567331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.567331Z digest=sha256:65b4be29c0b0319d320f285fc85ff4b81612705fd12b8d1a96888e8986d4ab10

Observation 7a0c61b1-3f95-425e-b8ce-27c69effc377 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.571419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.571419Z digest=sha256:2e2775794b0446baf76e580bd296fda4ba654c93c1b2b9b293c037d21310fac0

Observation cc980697-0bdf-4e78-936d-0e3f2ad73a81 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.575394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.575394Z digest=sha256:d64d5637f2e3c0ef003fcfd8f5d51c908294f9a4152d518a32e7291aad016785

Observation 77b55267-56ec-4a69-a202-3b0d4b9db60d · outbound

This paper cites Heuristic black-box adversarial attacks on video recognition models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Heuristic black-box adversarial attacks on video recognition models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.117058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.579468Z digest=sha256:02330fa3def2c383ea8760a670502881330be8dd2b7cc60ebdf98a43f1f04cb7

Observation 90334929-223c-417a-ab9e-70774e542d45 · outbound

This paper cites Boosting the transferability of video adversarial examples via temporal translation.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Boosting the transferability of video adversarial examples via temporal translation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.101434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.583937Z digest=sha256:cc447c9ebe7cd6ddbc3cc5d0b22f6d48ec2d67f2c920c3f0c418134416fbb816

Observation c90db321-d442-49cc-a7b6-74635529a31c · outbound

This paper cites Cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Cross-modal transferable adversarial attacks from images to videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.084862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.587903Z digest=sha256:ec7c51ad19c2d194c173ca4d4074d71a111161cf18dd919953889a565fa6948a

Observation 11506826-ea51-45e5-baa0-929ff5c98432 · outbound

This paper cites Adaptive temporal grouping for black-box adversarial attacks on videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Adaptive temporal grouping for black-box adversarial attacks on videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.069512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.592180Z digest=sha256:57a78a216cb2c9d83ab995c17fc374edf06383fec590f8a47f6a497b56e2ec53

Observation 63950faf-6718-4384-b2d9-e190306a711b · outbound

This paper cites Adaptive cross-modal transferable adversarial attacks from images to videos.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Adaptive cross-modal transferable adversarial attacks from images to videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.054373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.596832Z digest=sha256:02bf6ad9d788492b34a39d54a9a6e33c6ea1ab1c901a0e90a7de8cc58a497786

Observation ef1ee7dd-1f2b-495d-9dc3-6daabfc7e3d5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.601085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.601085Z digest=sha256:08a1d3382c03feb09d91ac261922bcc06d11aa2a0adfa65837d049b500032ce2

Observation d2c84931-efae-493e-b958-0a6fe724b5f0 · outbound

This paper cites Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.038244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.605433Z digest=sha256:27d632dc4e06390d4b98953656cedaf626d852660db4521f7a7977925d1beb17

Observation cafb8a2f-740b-498a-8a6a-66360466bc4d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.609942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.609942Z digest=sha256:83259661ea928d98e2373778392b1e5174a9a37eb3f6f5a6c8dd468e672d3058

Observation a12e187f-292a-4365-a725-f186be113542 · outbound

This paper cites Towards adversarial attack on vision-language pre- training models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Towards adversarial attack on vision-language pre- training models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:04.021487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.614224Z digest=sha256:9307120a3e444cd961a329db88e4d1be099d4a326f0a72f9d387265aa39b340a

Observation 3917b869-8038-4e06-92c5-d231c4ea92f5 · outbound

This paper cites Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.618948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.618948Z digest=sha256:ad17679efc2b9be0e8a4b3383e17f417cfdf265bc9ffdd296042e4f629cf235f

Observation b3f8ab5d-e1e0-4dc3-a967-4662dd2d825b · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.623119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.623119Z digest=sha256:67fffb080dd468277d0ef2d02a1bc0f20f9bbdcea0be1474712b11766598bd11

Observation 0dcb5446-bb7b-470d-9d49-d67f429a301d · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs On evaluating adversarial robustness of large vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:03.985237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.627271Z digest=sha256:9159c61c87d0dda51a06bde64d8232da8f014f72f65b4e2ef48353563f026e78

Observation 79e3a4b7-3877-468d-9264-e3da8c4e4bf2 · outbound

This paper cites Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:12:03.968998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:12:03.631575Z digest=sha256:faf08c782e4db298da021710ecce7caf79501820a54f71fada5d1e0031765aa4

Pith citing papers

No inbound Pith citation observations are available.