Pith. sign in

Paper Citation Record · LEDGER

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

As of 18 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 7 inbound Pith citation observations for arXiv:2411.19628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19628 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T06:04:22.130258Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.764054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:28:04.162152Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved45
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7fa0714-7f7f-4c6c-8c14-ea1ed9f12fcb · outbound

This paper cites GPT-4 Technical Report.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.785211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.785211Z digest=sha256:1ed387e1d8e9e4abad47a702220bdfbb5d7f12f826e2eea661fa305c677e0deb

Observation 9e933751-9f44-416f-b4dd-be77566ed270 · outbound

This paper cites Qwen Technical Report.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.790338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.790338Z digest=sha256:22dd337a635d097e6ee6db988e0509c9683a946e50ad0f842dcbc9be6a8a838a

Observation 20d4be6e-e68a-4069-9c1c-db8d9afd5f8a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.795845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.795845Z digest=sha256:bd995fed3abdcba8790efdf48714bb98294c956884868275630d147d69ce2c30

Observation a9bd0c36-ec62-4fcc-9594-8c8cd5dcc18f · outbound

This paper cites Qwen2.5-VL Technical Report.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.800835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.800835Z digest=sha256:65e8c9c88fca96251de562bf4c1b87145067b5fe9ba3038e5a457ecb8dcbd4a2

Observation 1437d00c-11ba-4f2d-a941-c1dcc9780652 · outbound

This paper cites Token Merging: Your ViT But Faster.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Token Merging: Your ViT But Faster

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.805637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.805637Z digest=sha256:b294b8204bdf082811b081ab65cd860bb827319471adff0c872b937b36468883

Observation 8842a2e0-28a1-4ecf-9c35-a9257a9fd0ee · outbound

This paper cites Language Models are Few-Shot Learners.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.810355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.810355Z digest=sha256:e319d4b3117be3201b681823bce3590b8dc35ee33239b7580de238749fb4eca7

Observation 0f5753ea-e6f4-4547-a809-2f18a63c6689 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.815034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.815034Z digest=sha256:65f57185a29bb3b1b51df6399f74da899ab43ac6cb6d6e228ac51de8b753345f

Observation 3f74320d-d2ef-475c-a61f-065ec70a59bb · outbound

This paper cites Diffrate: Differentiable compression rate for efficient vision transformers.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Diffrate: Differentiable compression rate for efficient vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.086789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.819220Z digest=sha256:79eede66d1e6603c4e2eb466e2e59e00593b5917b9f441c4b86882ca2cbea11d

Observation 30eab4de-4038-458b-a653-c6cc1e24f9bd · outbound

This paper cites EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.823327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.823327Z digest=sha256:2d092055f0a3626d8e4543dd50bc3ed88771853c6a82b49b4d51e3aabb9f36b3

Observation 9cb2905a-ccdc-41ca-bbe7-2fabeec38825 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.828092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.828092Z digest=sha256:867d0361a997f71fed55e9d5c20f6f8b720d74e54560f12bc68b8c2ae676d7e6

Observation 2b6f44f4-dbc4-42b1-aa8e-5703f5d01691 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.832091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.832091Z digest=sha256:cb678df9be0a1df2493c5238aa99b46a82cea0e9900bca0f7f6f15296788f0de

Observation 6949731c-6b21-4127-a66a-fe773eb2b603 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.835768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.835768Z digest=sha256:23e9f5601934a0f4a32ecf8d55cda26c32843255499358268753d984ce20c5a4

Observation 5272414e-490f-4617-8024-66431c39a580 · outbound

This paper cites Heatvit: Hardware-efficient adaptive token pruning for vision transformers.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Heatvit: Hardware-efficient adaptive token pruning for vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.074471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.839813Z digest=sha256:49e00468b9d6c8f6550d13d5c78e8935ac4db1c9a6ea11dc61de97f9705ac6b4

Observation 5b6809dc-15dc-453c-aa96-308612bf2668 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.843334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.843334Z digest=sha256:2061994dd9bfce1fdc69b33a6d676568f145342752c659b00ff51b7c1b750b29

Observation 68666eef-9be2-41c9-b6f6-7c7402ca0b14 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision- language model handling resolutions from 336 pixels to 4k hd.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Internlm-xcomposer2-4khd: A pioneering large vision- language model handling resolutions from 336 pixels to 4k hd

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.060866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.847449Z digest=sha256:af251210e30dc32bc27863eb1f5438f46a9c17488e2ab5afdbe240b75a1f7c62

Observation 40a00deb-fdca-48a4-a7fa-e33528c8e097 · outbound

This paper cites The Llama 3 Herd of Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.851308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.851308Z digest=sha256:db458ce87299333ffe778afe86838e8397f64c3ec413cf0d6bf405495d050cf0

Observation 424c65b5-9b19-45bd-820c-d9a27478b5b0 · outbound

This paper cites Not All Layers of LLMs Are Necessary During Inference.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Not All Layers of LLMs Are Necessary During Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.855233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.855233Z digest=sha256:a51688fdc1487b8af1a36439fb8e4e59c690be884eea70897d44fc87d517f75c

Observation 478adf23-ba37-49e2-8642-a7fcb4620345 · outbound

This paper cites Adaptive token sampling for efficient vision transformers.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Adaptive token sampling for efficient vision transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.047224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.860527Z digest=sha256:a98f3174961279718f1c083f55d3b168521a1c4237f56c1712b4ce26148d2eb8

Observation 1e08c585-456a-44b4-aee9-b6f8dea135d2 · outbound

This paper cites Deecap: Dynamic early exiting for efficient image captioning.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Deecap: Dynamic early exiting for efficient image captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.032117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.865239Z digest=sha256:7720f16c9ce4bbd916ff49ed14a2fa303409db50768d3ef8c9cced2c57c51780

Observation df016c43-767a-4315-a598-5f8403efb6a4 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.019433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.869070Z digest=sha256:9dda615d078ccab41b88341b32a6bea679a15542889d0871d3639cb6dcaec8fb

Observation 8940cc22-bc4e-452e-a394-25ff7c716919 · outbound

This paper cites Frameexit: Conditional early exiting for efficient video recognition.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Frameexit: Conditional early exiting for efficient video recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:23.006808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.872850Z digest=sha256:b59155e2596ae9754839cf5064cd377b294cc8dc6b7ca29deb83014b9e5d53a1

Observation a56632bd-b690-4206-aae0-249a999b5b41 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.876465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.876465Z digest=sha256:3dfd3573a35efc5047e59f94cead5ca1dc0bc0b70e38906cc9b0232d1043fe53

Observation 0a9aab22-7e9c-4f66-a838-41f9c525b287 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.985901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.880644Z digest=sha256:1a2c6b6e2e0e060ab3cfadba3c42b5c5062d7d7aac3b543fc9f006ac91277be5

Observation 0c679286-540e-463f-8042-5b10796a3c58 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Categorical Reparameterization with Gumbel-Softmax

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.885376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.885376Z digest=sha256:ca0ba59e5eb2cebf2d35ddada302d888e5d50af8175feb63d9fe1cf97b5ab87e

Observation d8697c61-d8d2-4148-b85b-8079e5afc944 · outbound

This paper cites Mixtral of Experts.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Mixtral of Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.889737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.889737Z digest=sha256:69f5156bc413a2c6091dfab055fd22f1822e9e94241e94da2c187359512ad8d0

Observation 34b38d83-6c62-4c96-b68a-760f1e5e1bee · outbound

This paper cites What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.895045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.895045Z digest=sha256:0532ca7c330ca0749301d5e3e117b3ce876a72d2f0f27d6f695b2fd4138de7f1

Observation 4b0dc344-5387-49c1-9d37-7bee09ec74ed · outbound

This paper cites Spvit: Enabling faster vision transformers via latency-aware soft token pruning.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Spvit: Enabling faster vision transformers via latency-aware soft token pruning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.972796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.900561Z digest=sha256:3a0b4bc6bbd4ff690626981066c59a89621e3e51f2ea86c6d68b235ae4d2f786

Observation 5211d79a-491c-4951-9946-bb5dd6b3947d · outbound

This paper cites On information and sufficiency.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings On information and sufficiency

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.904628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.904628Z digest=sha256:227ee73e06f8bc5a271ee6f630514100479600d69a41e9cb5e285e73cb1e6f2f

Observation 1ab43473-6cb9-4410-80f8-f00ce523c587 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.908563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.908563Z digest=sha256:a390b689877b15086e61e2a3522bf53e56d4e698ddde39c120ff1e435979a0b0

Observation 9b3189d7-0583-4edf-9937-d7f94e36997c · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Seed-bench: Benchmarking multimodal llms with generative comprehension

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.953871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.912633Z digest=sha256:ae82d32b41b9559db3343d55682bb93158ded0e9a68d4c74d6b65c2e28a49c83

Observation 105a0086-71cf-4f37-a1c3-291c38ce80bd · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.916538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.916538Z digest=sha256:63790f7e38aa8ed5b819439feab70415eb78240c8f523e860694c3602a6e9f05

Observation d5a7df4a-714f-4097-893a-1e71631a0529 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.920297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.920297Z digest=sha256:d0ee42cce891aca31686b4a1cad6fff64248f4937ee6b8ad7d63b2d015f29b39

Observation 600d70d4-c2c0-4ec5-af37-6af14ea74c78 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Evaluating object hallucination in large vision-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.934254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.924664Z digest=sha256:3713dda0f6ce4f541768dc06f7140a9f07d930e2c0c69887b3b3ad3821c61a6b

Observation 88c28a23-9dac-4efe-94ec-f9420798bfa3 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.928897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.928897Z digest=sha256:bbee2db214a340b507b92de047eab6f8abaa18d07e082e7c953e5b4288a8c5b9

Observation 43faf140-3256-4be7-8d68-3f5a65ec93d5 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.933315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.933315Z digest=sha256:03a7c425f6f797a3b0f595f3f70b8eabada73c40a540e27a56be72ff9ebec903

Observation a26d361b-458d-46ad-8683-25451e33e88f · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.937545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.937545Z digest=sha256:448e8036d79ad4a7cbe70572b4317fb6e34870cea08324cf5c11609a3b0bd4a1

Observation 650dff22-6f06-47c2-8247-6d95ecf79a50 · outbound

This paper cites Looking Beyond The Top-1: Transformers Determine Top Tokens In Order.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Looking Beyond The Top-1: Transformers Determine Top Tokens In Order

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.941986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.941986Z digest=sha256:7189edf82ee98720964f117f5607d25f0bb793993960c3f956570c2e18f02e6e

Observation 0da02878-3dfa-4a2a-bde0-2093aeff14e8 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.946520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.946520Z digest=sha256:1a83031b141c3a3cf1454e8f9d7372eaa4bc03eed804e01c712d15576246646f

Observation a3ce24cf-b297-4372-9d14-6d3835a357de · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.950889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.950889Z digest=sha256:151c6d8dc66d6db9b1e03327f9968de559ae51062b6440d1d43306606ec256a8

Observation 2a0da20e-0e42-4499-8a25-64c5e48374d4 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? Computing Research Repository (CoRR), 2023.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Mmbench: Is your multi-modal model an all-around player? Computing Research Repository (CoRR), 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.908220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.955025Z digest=sha256:5d37d9b1082c99455024267f33c96ff517a1a81df3a9dfe4c0bcab5b7c5aa372

Observation caf056e2-52c7-46b2-a82c-646e45cc0ffa · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.896753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:21.958986Z digest=sha256:d3604986bd536d27203e2b9a29d314ce3f26ba8176689b9900733709b32feeaa

Observation 90d7aa42-0a91-4048-bd20-d1b3611c86cf · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.963520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.963520Z digest=sha256:e4b47504c44069920b6db56a7d9b5874791e9df07a0fadd349e743fbeaea57ec

Observation 0c93317a-d433-433c-b558-e12423e4ee24 · outbound

This paper cites Infographicvqa.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Infographicvqa

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.968221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.968221Z digest=sha256:821083fa447e768ccb64b75d8535fd535ec6a332d92e58bfc1d23f8a1c358508

Observation 0147851e-7164-4655-a069-d35197dbf8c3 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Docvqa: A dataset for vqa on document images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.972537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.972537Z digest=sha256:741f0e6914c38fefbf0661a7bb81675d7ebc07c4e41a712568c43d2c8611e23c

Observation 9d4e526e-5e47-4348-acca-dfa99769d574 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.977433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.977433Z digest=sha256:701efcdade07532d99481a2b418fc226cf421e5eaff36eab617be01ac23e3c2b

Observation 4739534a-6456-4b2d-b560-a050ad465524 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.981984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.981984Z digest=sha256:59bf284602b472560ed9203b35d202e9da7da25d2e2c3b2c48d9c37201f74470

Observation 520bca3a-9828-4f0a-8aa0-be5507544d76 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.986845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.986845Z digest=sha256:6d68202a9f8b7e7f268f867a99611b8b4d6e1ea367e8a9cede7a13ee38513d33

Observation ff6630a0-aae5-4100-82b5-96e4e0137e79 · outbound

This paper cites Towards vqa models that can read.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.991955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.991955Z digest=sha256:375e22bc0f6fe8552ba3181908e34407098115665de9dea3673bbf571a396317

Observation c8033177-2b91-49e7-918c-f0cce9e83620 · outbound

This paper cites A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.996124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.996124Z digest=sha256:7be61d77f81ae9304bb4fbe76c37876f5c5e7fe0aa648ef7260478566ec71c9b

Observation d9985e08-4041-479c-87a0-897586267db7 · outbound

This paper cites You need multiple exiting: Dynamic early exiting for accelerating unified vision language model.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings You need multiple exiting: Dynamic early exiting for accelerating unified vision language model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.001368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.001368Z digest=sha256:f92740fb0e1bdc4fb7eb67f95e9e1bc1c8c32bd950df50a8843d05a164fdb4e2

Observation 08cf5ee7-7419-4783-ae7d-c1f6fab53703 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Qwen2.5: A party of foundation models, September 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.006511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.006511Z digest=sha256:ed4d48d020775d4198d89d639232545dcbd81b7525c6ad82b6e91d9ff745375a

Observation 79c47bf7-ee05-4909-b592-60bfe0b79867 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.012057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.012057Z digest=sha256:ba99dedcc0ec50d81d7d6cead675e9f01d682c65393a2a8190d3ed83862c602b

Observation d04d3bfc-a572-4ab3-b771-538ee0dbf1fb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.017568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.017568Z digest=sha256:b78676d79a7ef4e6f54c1ab6cf8c0f9b38a3de969ce6c7a3ca4c23473cf6793f

Observation b23538f6-64d4-4fc2-831f-f564543689e8 · outbound

This paper cites Zero-tprune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Zero-tprune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.848211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.022189Z digest=sha256:6151945550aa6a66dc6df42c0312b0e7709f146a2c9ee13fbdd5872e90e5ba26

Observation 23960389-7de0-4cbe-a5aa-d2090aa7c592 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.026877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.026877Z digest=sha256:e40958487fccca434f7a655e707d6fda7787a9b331d5546abcf7b2953a289ee8

Observation 9f9f16fc-2bce-44a8-948c-969f0bdebf39 · outbound

This paper cites Joint token pruning and squeezing towards more aggressive compression of vision transformers.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Joint token pruning and squeezing towards more aggressive compression of vision transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.834941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.032362Z digest=sha256:2de2b2bc2c2e39c3b65a2f57a337595666c8d50e28484234f0f62412a37061f5

Observation 8a20ff2a-cece-4070-bc03-06edff368251 · outbound

This paper cites Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.036925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.036925Z digest=sha256:aa4777d8c5c12912969bb4125eb912b57bae6b548debc6000262c2149ccdb06d

Observation 75827c6d-13a3-48fa-b147-4e8e45f3a131 · outbound

This paper cites Parameter and computation efficient transfer learning for vision-language pre-trained models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Parameter and computation efficient transfer learning for vision-language pre-trained models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.822318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.041663Z digest=sha256:31426a5bc3fa2f6bfde1e1112687c5fc19c3722a849195ab46da708eff857d55

Observation 6b3c2619-29df-49d0-a8f8-2f8fefc4e496 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.048660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.048660Z digest=sha256:4be62ddcafec7e949b558775e4e2954be1e9ea52de8c125ed2b3a3622ba5a6a8

Observation 393329a3-b647-4b1a-8236-92e1c3c38c42 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.053400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.053400Z digest=sha256:6ea1b031f2e9b5410ae463b9389c9fa9b3292c1f1e9729144db6cddb37b72f3b

Observation 5e12da4e-5ad2-4e0a-8d5b-c84464ae0f64 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.809932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.057859Z digest=sha256:cbb8db5966643765f42096b9ab45a254b34e69411697aa62177ee3b978207238

Observation 2cef6003-40d7-47cf-9eae-088364b11489 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.797578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.061798Z digest=sha256:284231d5f12d114d2021993d8ef9c390805d9383acc961750195d3f9b3a70ff9

Observation 0bba889d-0024-422e-a317-4e22f0bcf10c · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-12T06:04:22.065841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.065841Z digest=sha256:484d632529bd55b8cd5059a2862c6c8f2141a3f90293d3688b4d74f986eb6b67

Observation 5d7b68f4-2d66-4079-b56b-29c5bcf58245 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.786131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.071014Z digest=sha256:55949a680d3ddb5f2903c18ccf6c34aa2931b9962435a2d31a72f3ed96e6a67d

Observation 9a82735b-f8ee-485f-9835-1e7060909915 · outbound

This paper cites Limitations.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Limitations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.774526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.075011Z digest=sha256:7ba5ce049ad704d1c36f47173595af2aac6c34c6803be6ffbb3b1e73dad6d321

Observation b6d86275-0ed9-4426-8559-654e80b2ffaf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.762883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.079172Z digest=sha256:275ff7919a75c6eb4bf43db5912e1b955288a5454e1005d9e79d04e1882a3507

Observation c62e8254-8b34-46a5-a549-397f140a401f · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not include experiments

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.751346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.082996Z digest=sha256:854b01fac6d8e4d05a0370bf5ce4c801dcc233f611a819a3603e93f60c8aef35

Observation 1e40ba61-aa22-43df-8c85-bd997ae91d3b · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.739591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.087739Z digest=sha256:723b300c0540242094088a0aa830cf7952dbc48ce309a94566f7cc0b84ee3836

Observation 221209bd-72ff-40bb-be5d-155ee7c16017 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not include experiments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.727652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.091807Z digest=sha256:5e6217034631e6c2fc3978d5ceabb489ef07c8288eab63f29e9cafb876714d01

Observation 9f01c4e8-b7a7-43a1-b1c5-bbe616b58a1c · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not include experiments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.715582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.095530Z digest=sha256:2f80e2143bf521293410b8336e564578eaf69dd75712286b7ca7ad6ee0ee962b

Observation 1cd81772-c7b1-43e2-8fe9-76077bd87df8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.703118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.099327Z digest=sha256:560c87a1f7df7930cdb85e616e91ede2a559c3c21ba815c77cc3532e764d2d16

Observation 6f92737b-a167-4fca-b32c-b8db77214605 · outbound

This paper cites 22 Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings 22 Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.691102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.103286Z digest=sha256:d6eb241f82432f3a91ff4878c0a864166a1b9edcb464e2b1e8c720a2d88f827a

Observation 16c13f32-6630-4ecd-8e4e-307463c81beb · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.677799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.107460Z digest=sha256:68813c246f3a9af18ce2db8ee2505f9304a61693d41757c0a49dbe2772b81430

Observation 2124bd1f-02d6-46b6-9fc1-51a1bc1425ab · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper poses no such risks

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.663835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.111644Z digest=sha256:52167d9f687f363eb062ad3677daf565bbf1b716fd10bee74ca9685ff30e137a

Observation 145f4e33-8a3b-4749-bfbb-917199901370 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not use existing assets

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.651703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.115098Z digest=sha256:5bf35b5d83eaf5490a65e2975efd72f13a73e0564c895fff8b622efb90406505

Observation f25ec9c9-58a6-4300-9235-aa723ca96499 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not release new assets

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.638594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.118870Z digest=sha256:ba702a4243981355a23f402f75f6729738aa9f5897a706826219a1c8fc53e368

Observation a1ac8fbb-c5e9-4718-9ed8-0b305978fb0b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.625993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.122478Z digest=sha256:c6db0e54be72c37c0883bbb5bf5fb3c6565971696cee4d521fac6c5237afbf02

Observation f046147d-d26f-485f-b3a9-4ec0a8b82eeb · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.126180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.126180Z digest=sha256:f252d3c1a538f607f681213405acba0aec4b649ed464e187182061a45ad4f246

Observation 200a829f-cd32-4381-a14e-8127df4d31fc · outbound

This paper cites Answer: [No] Justification: LLM, as a part of MLLM, is the object of our study.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Answer: [No] Justification: LLM, as a part of MLLM, is the object of our study

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:04:22.604187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T06:04:22.130258Z digest=sha256:e2103502156175e36e555172f49efabbaab3da4b3da1f085d9f5524b70b114ea

Pith citing papers

Observation 8b26692e-6562-4150-969d-72e6e6bad9ad · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.805853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:8924fc29c761eaaa6227811ecdc29032a17abbefbbcb3c49a20424d62d6aaa28

Observation 0ed183cc-559e-4e27-a08d-8b033531194e · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.764054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.764054Z digest=sha256:736277ac55a1dcbd835a1f6891855050ae84b953ab46c2b0f26346d12243d022

Observation 673f7275-980c-40d1-98c7-f3809026eda4 · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.628462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:9a97d02b9f7ae3902223cdac34a4c757c73876f8a8bd22bef4773e10a85c1a32

Observation e11b46e0-8a0b-400e-842e-a2919bcf1509 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 21

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:06:03.284956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:26b82183cd9b682d3d8dc79793feda470abe1f406f99ecca213c046e42b98ebf

Observation 86468620-b557-4ddc-86e4-0b67b7a43ae4 · inbound

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models cites this paper.

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.163716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:35:24.118536Z digest=sha256:c6de262caac966ea1740f1f40df24fd6890b756c81b23e4909a7f7b2bda78782

Observation 1d859195-f4e4-4553-8e8e-ade08e6c65ee · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:55.518589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:55.518589Z digest=sha256:40ead2d10435b8f30d40e1c070f753a92ee4715524db18cf8be752590bb03d0c

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:0e6b725fd727cfb5e909927f2905293caf35d709e2a63c6e854b9021789ec588