Pith. sign in

Paper Citation Record · LEDGER

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

As of 5 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 4 inbound Pith citation observations for arXiv:2602.07574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.07574 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:36:43.892124Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:16:46.738649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:19:44.559138Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c941f91f-458d-424b-bfee-57c79342a20f · outbound

This paper cites GPT-4 Technical Report.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.786045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.786045Z digest=sha256:07f177ed62beefb75a93d83e8037322c136cebbecea31c338bbbb100d2377516

Observation 66d283e1-b9c9-4aac-a3a6-5ebdbb55a070 · outbound

This paper cites From llms to lrms: Rethinking pruning for reasoning-centric models.arXiv preprint arXiv:2601.18091,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention From llms to lrms: Rethinking pruning for reasoning-centric models.arXiv preprint arXiv:2601.18091,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.797433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.797433Z digest=sha256:616198ed4dbf533cfc33a40494de456bdae28b7c60d8a2555d22f6cbb299b5c9

Observation e581471b-eaaa-4f14-bc2b-e3565ed59700 · outbound

This paper cites Visipruner: Decoding discontinuous cross-modal dynamics for efficient multimodal llms.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Visipruner: Decoding discontinuous cross-modal dynamics for efficient multimodal llms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.800697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.800697Z digest=sha256:069dd80fb1b5db256312e690ee31a3931a3aef7ef117f2399f8fa99972dcadb1

Observation ba9bde6f-dc17-470b-b9ca-27ae5a9e6181 · outbound

This paper cites The framework tax: Disparities between inference effi- ciency in NLP research and deployment.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention The framework tax: Disparities between inference effi- ciency in NLP research and deployment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.803774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.803774Z digest=sha256:adfe3e78ef528b5cf8512191ceb595ef387503199f1807efb16ba659a9a6de8c

Observation a00e5172-ce7a-4dab-a7b9-6aceeea84fb6 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.806789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.806789Z digest=sha256:a2e5121f6ba27c487f6bb35539022086a22ef9f5393cb4a8ccd4613b2805cf20

Observation 725bcd9b-1c9d-45b9-85f1-f21bfefa4233 · outbound

This paper cites Filter, correlate, com- press: Training-free token reduction for mllm accelera- tion.arXiv preprint arXiv:2411.17686,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Filter, correlate, com- press: Training-free token reduction for mllm accelera- tion.arXiv preprint arXiv:2411.17686,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.813342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.813342Z digest=sha256:b5de553a00d8102bbd8d830614bb7c842849e1ccb351aff14c4e8952d8e19965

Observation e6ac5bc9-df02-4fb8-b581-a2bcfa666296 · outbound

This paper cites Training-free token pruning via zeroth-order gradient estimation in vision-language models.arXiv preprint arXiv:2509.24837,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Training-free token pruning via zeroth-order gradient estimation in vision-language models.arXiv preprint arXiv:2509.24837,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.819419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.819419Z digest=sha256:e552c6f62dd7741ead0b32a486072eee7c55302283d7066f3723997a88beb5bc

Observation 2698dd5f-f538-4696-9fc2-925bfeb145e0 · outbound

This paper cites Crafting papers on machine learning.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Crafting papers on machine learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.822334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.822334Z digest=sha256:f746f86c1abd6d215fddebb06989e35db996d540bf8be2c82d31c90553a16bae

Observation fe6a8878-5f93-4c2b-a242-d594f641f58d · outbound

This paper cites an unresolved cited work.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.885996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.885996Z digest=sha256:49fc5c4bafd4b8b67d182738d92cdd4715bf7ca38c4125f76ad0a8b3229c4a1b

Observation f79e5802-4c2a-456f-93e8-e2229b6e0f5a · outbound

This paper cites an unresolved cited work.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.831132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.831132Z digest=sha256:57bcb3620e2da01e32974d4ae46586fc76c5f3706b53ba696bb4fb2dd49b13f0

Observation 344c15a1-56d7-4455-85a7-9c432869e3f9 · outbound

This paper cites The few govern the many: Unveiling few-layer dominance for time series models.arXiv preprint arXiv:2511.07237,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention The few govern the many: Unveiling few-layer dominance for time series models.arXiv preprint arXiv:2511.07237,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.833966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.833966Z digest=sha256:8d3aa2a3375888f385aa47471f19caec6aca634333f35165bcce5e2b5d8ec7fa

Observation f6a59e34-f22e-4c8d-9df7-618e24ce7a7d · outbound

This paper cites Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.836659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.836659Z digest=sha256:7684f141bd5d9e8a3b85c68d92ab028f37022b03a412fc7bc76995380ec42369

Observation 6f6e8c13-3300-4526-90f7-e31565de98f1 · outbound

This paper cites WeLM: A Well-Read Pre-trained Language Model for Chinese.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention WeLM: A Well-Read Pre-trained Language Model for Chinese

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.839708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.839708Z digest=sha256:809a4cfb5ce3e04b61d314d2add86427c81b08f8a547d5a77b1e4adeef9f82fa

Observation 1fd97f3b-7b79-4986-bc5d-e941bd1876c4 · outbound

This paper cites Unraveling the Mystery of Scaling Laws: Part I.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Unraveling the Mystery of Scaling Laws: Part I

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.842782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.842782Z digest=sha256:093f9717ad4fba87a2fa6a280a9dbb67213ee90de324be4b0d113a5feb4c2228

Observation 2e46ced7-cd4e-4db9-b348-b7ac0ab0b99d · outbound

This paper cites an unresolved cited work.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.845796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.845796Z digest=sha256:9f809f672c06576362e5ad65b55b77f27d9782897b142dda3171a762bcc96c07

Observation 422fd897-39a2-45c6-b90b-6c39a3d0f062 · outbound

This paper cites TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.848589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.848589Z digest=sha256:91b2b103662744a7ae4805128a94f3692a43068b5585034af148552ebe698f10

Observation 3d4548d6-a087-41d5-86cf-b33d5ca5d617 · outbound

This paper cites Flowcut: Rethinking redundancy via information flow for efficient vision-language models.arXiv preprint arXiv:2505.19536,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Flowcut: Rethinking redundancy via information flow for efficient vision-language models.arXiv preprint arXiv:2505.19536,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.851615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.851615Z digest=sha256:cdedb0e654136bc52da11ae546954803b9104d6757f4ef15d80e7551f5a6c120

Observation 77fa743d-5dae-40b2-8c33-67135aaea41d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention LLaMA: Open and Efficient Foundation Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.854536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.854536Z digest=sha256:92709a7eabb50f27c176f0bb7393896a9b2d54a0507283a62768c2115ea98cf9

Observation e4a79494-0980-49de-93da-1059d4310440 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.857558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.857558Z digest=sha256:00d878cf71ef184f22ca668cf884bd6a6631446af0405ec56598e289f885df15

Observation a379abfc-05fe-4aee-bb5f-0ce88744d63f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.860430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.860430Z digest=sha256:265bc4922002e137737edacfbc19db1551cca772e98b0d70bd3cb93efc5bae47

Observation 018b7785-79ee-4344-a710-5e45add4c883 · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.863246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.863246Z digest=sha256:c81ca3255ce607f23991186e02dbcd9efc58a92286765f492e02ad05672cd547

Observation 8d0b4ba3-5918-4c10-af00-09dbfa938a6a · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.866305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.866305Z digest=sha256:d315aaa9601d7c09d75eee9aa19290097e1f52ff0282a4229db997a66b649e2a

Observation 068ed5f8-001a-4c3a-b897-6ec3b1f2a7d8 · outbound

This paper cites Shortv: Efficient multi- modal large language models by freezing visual tokens in ineffective layers.arXiv preprint arXiv:2504.00502,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Shortv: Efficient multi- modal large language models by freezing visual tokens in ineffective layers.arXiv preprint arXiv:2504.00502,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.869360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.869360Z digest=sha256:99e72f9baeeb3656b7885270dfdfb5aed62ea00487181c29718d46f2fb244e42

Observation 70140cca-3983-45bf-b5b3-eec4c05b5e23 · outbound

This paper cites and Shukla, D.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention and Shukla, D

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.871953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.871953Z digest=sha256:21cd0cb8ab0feb940b2d9f7ad5b4811d44ae2aa2c72df44d9240efc874a86365

Observation e035451b-2ac5-4803-9ea6-bea64010fbf2 · outbound

This paper cites Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.874806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.874806Z digest=sha256:e0f9d50440ab0caea04dcb3016c67ce76973b09de4b08746453eb000d91a5205

Observation fe600739-bc70-4b9b-b5f5-0881241d255b · outbound

This paper cites Vscan: Rethinking visual token reduction for efficient large vision-language models.arXiv preprint arXiv:2505.22654, 2025a.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Vscan: Rethinking visual token reduction for efficient large vision-language models.arXiv preprint arXiv:2505.22654, 2025a

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.877803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.877803Z digest=sha256:0307c44402e9cf03e16b198c9407a70625e02cd93c2a9bd3ffdc274951ae1212

Observation 9f8eb748-1179-4bab-a8f9-329a9567700c · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.880446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.880446Z digest=sha256:ecad71823f3a9fa4e838bf9fb62bb03ab521d44a2276671b871b32635f7c919a

Observation 1d029859-003d-4829-b021-a177564b018d · outbound

This paper cites Don’t just chase” highlighted tokens” in mllms: Revisiting visual holistic context reten- tion.arXiv preprint arXiv:2510.02912,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Don’t just chase” highlighted tokens” in mllms: Revisiting visual holistic context reten- tion.arXiv preprint arXiv:2510.02912,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.883350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.883350Z digest=sha256:36ccdc38f6c61389465d5f8180ff2c081ed0875f4c26f570d4790d82fd58d23a

Observation 9c9823b9-2a33-4884-8775-97d688063185 · outbound

This paper cites an unresolved cited work.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.889099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.889099Z digest=sha256:80b7bdecaccac56c5bb9eda0c966bd5dde4a84d8dfd800097daf6f8c4c548bab

Observation eba19160-294e-465f-af7b-e273d149be6e · outbound

This paper cites (b) Our method under eager attention, where visual tokens are frozen via explicit masking and participate in attention only as key–value representations.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention (b) Our method under eager attention, where visual tokens are frozen via explicit masking and participate in attention only as key–value representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.892124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.892124Z digest=sha256:5f4c6aa217c5d51057e49528cda806e05739ccef4ca28cc62f17648d27c64961

Observation e48dae1a-7083-4a06-9cc9-8fe29af1ac0b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention LLaVA-OneVision: Easy Visual Task Transfer

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.825189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.825189Z digest=sha256:f3c211e2492d35cb83acdc063508046b9086f039d83aea439d88738eca0269b7

Observation 2ff8ca33-d940-433b-be82-67ebea37e28f · outbound

This paper cites Informed routing in llms: Smarter token- level computation for faster inference.arXiv preprint arXiv:2510.13831,.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Informed routing in llms: Smarter token- level computation for faster inference.arXiv preprint arXiv:2510.13831,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.810118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.810118Z digest=sha256:c1218b0f8cb6b13f0880bf0580b395846b7fc66fd1da5b6b5c5eeb0955c5936c

Observation e25429e2-c7ad-4414-a600-650fcb048802 · outbound

This paper cites BLIP-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention BLIP-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.828231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.828231Z digest=sha256:933bad9b476db29f0f367ff9ccbedeb19376802fd2b1bc9e3793e384c891dd0b

Observation 91f02fb3-bf7b-443b-a0cb-d07aa6f705b0 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.794122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.794122Z digest=sha256:832b9eb935540bc2efeead050307c9611c0598e17179cbc6980896bcfcecab4d

Observation 42a9394c-6e02-47db-8123-83d92b80befd · outbound

This paper cites Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.816273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.816273Z digest=sha256:f74427620f11f3e82260931f9303c061ab449dec39ded14d4b6a29aae9e72517

Observation ba07eca3-772f-4565-9d43-a4718bd14d41 · outbound

This paper cites Qwen2.5-VL Technical Report.

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:43.790329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:43.790329Z digest=sha256:4d4fa82f1f6a40a0fe469981705a4e00b7ea163842f7af5298dc579de37363bc

Pith citing papers

Observation 36894a2f-169f-4e22-aa11-c9d715c1990f · inbound

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation cites this paper.

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:46.738649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:46.738649Z digest=sha256:a6692a1e786f77d472e414e68004d59bf7a9660df5e9f9f7d00be7d5ef92cd54

Observation 0ff9dad2-ab9e-4d4a-9b72-d7405e708a29 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:29.982411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:dbc2be82db026f961d742519d6fb401b4fbe207aac7cb21c3dda903d7320e3f0

Observation 5477009a-0d09-4044-911e-9ec5cb869b84 · inbound

From Recognition to Understanding: Unlocking Cognitive Time Series Reasoning with LLMs cites this paper.

From Recognition to Understanding: Unlocking Cognitive Time Series Reasoning with LLMs ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.560480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:49:04.883246Z digest=sha256:2e063a5ac2930eb715599df0648f328823145c544bba5c8fc3f9ede0e70f9036

Observation a404a56c-7951-4876-828e-90c1de1ccc84 · inbound

Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes cites this paper.

Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T05:06:55.036228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:06:55.036228Z digest=sha256:1ba41b52785e11b46bb6337d143f3e79dd54e45ed8fbd842c1be58143d0eb523