Pith. sign in

Paper Citation Record · LEDGER

Speculative Decoding Reimagined for Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.14260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14260 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:56.294364Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:44:37.884895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:11:02.504431Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dd385fe-9b0d-445d-b8f1-cc68bc70f85e · outbound

This paper cites Gpt-4 technical report, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Gpt-4 technical report, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.607827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.002187Z digest=sha256:7a4da7b69bfee9846695f8bbb0a4d5487370d0ff6df4a37a45cb8858b5f9b9b1

Observation 7f67eed2-59f5-48e5-991c-cd46b061895c · outbound

This paper cites Sharegpt.

Speculative Decoding Reimagined for Multimodal Large Language Models Sharegpt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.438602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.078322Z digest=sha256:d2d9e5d38aa41d89465ce7c7194bbdcd6ca5f494646a5117e53bfd2e37341bcb

Observation e6927f43-8e81-4274-8bf4-258444046e7f · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.149539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.179436Z digest=sha256:5e2824fd98586860dab0ffc5c192104f00a3b2eb860c087f8f70748b91d47aea

Observation 965a96a4-3d00-4de6-ab69-719b76398cba · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Speculative Decoding Reimagined for Multimodal Large Language Models Qwen2.5-vl technical report, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.929426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.248582Z digest=sha256:b08b60f9c6d9b2c878845c7fdb9fc84483979ed8dc7a8c00872b7d21b627a0b1

Observation 8a5d4887-01d3-4a29-a57d-195b85dab7ac · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Speculative Decoding Reimagined for Multimodal Large Language Models Lee, Deming Chen, and Tri Dao

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.679671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.378487Z digest=sha256:5f7d224c04e0a341f81f2d16f70e275cf9704d47fc343faac47f1b15167afff2

Observation eb9916d6-7160-437e-ab3b-039bc1341d38 · outbound

This paper cites Madtp: Multimodal alignment-guided dynamic token pruning for accelerating vision-language transformer.

Speculative Decoding Reimagined for Multimodal Large Language Models Madtp: Multimodal alignment-guided dynamic token pruning for accelerating vision-language transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.486357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.477020Z digest=sha256:14e1d233bb9979ee5a4c0c2378b5ba590afc3f4ac683c0dc1505131d6adfdd55

Observation 2f28a232-067d-40cb-b824-62a0df9bdbc0 · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:51.603894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:51.603894Z digest=sha256:6b66797fe800aae69a5eb1a3d1efcda1a4c70e2e55bf1bf3ce8368545de9ee3e

Observation d2cb13a1-3cfb-4125-9508-c2cb4afd6868 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Speculative Decoding Reimagined for Multimodal Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.308827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.758861Z digest=sha256:1a192f1408e607ea61173856e6498c650376359309e5b4f7e3cdec788e86b5f8

Observation 8472b815-8778-4719-b53a-e657fa0b0a88 · outbound

This paper cites Diffrate : Differentiable compression rate for efficient vision transformers.

Speculative Decoding Reimagined for Multimodal Large Language Models Diffrate : Differentiable compression rate for efficient vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.090731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:51.856526Z digest=sha256:0e53595bc1de2455fcce256961f2bd4a4557d77a489247589642f9e1a3cd1784

Observation 5971ae9b-2c20-4fff-85ce-4694e2f964cf · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Speculative Decoding Reimagined for Multimodal Large Language Models Evaluating Large Language Models Trained on Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:51.998002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:51.998002Z digest=sha256:4209de53fa9a73b0e377a0c14e7fb919d51e1dfd652addac563a83e0252c45de

Observation cc6f647b-2343-4021-b403-87dbb174dd76 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Speculative Decoding Reimagined for Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.140080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.140080Z digest=sha256:98dc1f8508978e72addf0013534b411029b39883167172f15e93d63ab44feebf

Observation a1a6d5c4-a8c2-4b6b-a810-159e414f5e6d · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Speculative Decoding Reimagined for Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.269586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.269586Z digest=sha256:d5d36e6827f846e3cd8bc101f079b272298edf2fa3f6c39adfac6104f917a929

Observation 35311bcf-9417-48ea-8de5-2a63cfc3d0a1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Speculative Decoding Reimagined for Multimodal Large Language Models Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.384747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.384747Z digest=sha256:0cac9fad25bda023ff5a351de456e6528b61dc8eb3a1313f436673be967e6876

Observation 607f0ff6-625e-4e48-b6a0-4c393f4a9d1d · outbound

This paper cites Flashattention-2: Faster attention with better paral- lelism and work partitioning.

Speculative Decoding Reimagined for Multimodal Large Language Models Flashattention-2: Faster attention with better paral- lelism and work partitioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.902246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:52.520459Z digest=sha256:d5c03c0fe7905d5b4190d0b6e8441433bfafac9cc4c2c7861df8c8e732010154

Observation 2cca68c7-34f7-4400-a9b6-cb1ffa537e3d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Speculative Decoding Reimagined for Multimodal Large Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.688863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:52.654998Z digest=sha256:9b8fc4f64ba796f35b7da574263ff3151b78fd2ad6d66ad22fa4d10f709fbac6

Observation 46f9dfcc-1f41-458f-8b67-45ad1b2ab7f0 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.454182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:52.766877Z digest=sha256:5665c5593ae5695db65741afcb2c093947c582b9708db1ecd1f489081c1c9003

Observation 40c1840c-6422-42a9-9344-4dacfb33782b · outbound

This paper cites Break the Sequential Dependency of LLM Inference Using Lookahead Decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.964039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.964039Z digest=sha256:46fc0c9b428cf4b5ab0620d367705b41b9ea2aae08eb7183f22402d810045aa9

Observation db0aa4eb-b18e-4dcd-9db6-1983100197a2 · outbound

This paper cites On speculative de- coding for multimodal large language models, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models On speculative de- coding for multimodal large language models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.284745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.096692Z digest=sha256:fc95cbef61414781b718aa34d684855742ece508a95060428f015f8f82b8a13d

Observation e8600f12-c8e8-4e6d-bf18-324907bf2c49 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Speculative Decoding Reimagined for Multimodal Large Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.107501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.184615Z digest=sha256:9ea92b7b47e215b3b3f897fb28792ad25bc13e9eb79b6d20a5e3e01ec71d1152

Observation e755fa42-dd18-4317-a2b3-9d9c1240456e · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Speculative Decoding Reimagined for Multimodal Large Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.878228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.273506Z digest=sha256:3255f500bd6e822e41c195803992750a838d706b7fce0b8ee7dd35d0551f20c9

Observation 00150653-8d2e-4abe-93b2-8d56c5ad849c · outbound

This paper cites A diagram is worth a dozen images.

Speculative Decoding Reimagined for Multimodal Large Language Models A diagram is worth a dozen images

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.615171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.386578Z digest=sha256:d5974f3656491a680276a1d40d50937d6f43e09e41df1e256064a5cb6277dace

Observation e581eb29-484a-4097-97be-101792a256b0 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Speculative Decoding Reimagined for Multimodal Large Language Models Efficient memory management for large language model serving with pagedattention

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.306653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.502385Z digest=sha256:54cb938505978b0d10ec9112e5040a09f1ff450c226046b16519ab7579a15dc6

Observation df04763c-b5f0-4827-aee5-b89b7354b870 · outbound

This paper cites Fast in- ference from transformers via speculative decoding, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Fast in- ference from transformers via speculative decoding, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.079645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.602320Z digest=sha256:7d603a7c768f0d998f512dd9aff7349e2f519fce2a8e3eb8278e7b43aec7d315

Observation a7c93abb-242d-4d10-bde5-3d0554f5ef1d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Speculative Decoding Reimagined for Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.815404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.666871Z digest=sha256:f2d12f9c592a969038cd76508817e9bc725cc1efa356d3e48577f0ac5ff0c7c1

Observation 21f14983-b513-4013-aeb5-2932a222da6a · outbound

This paper cites Mbq: Modality-balanced quantization for large vision-language models, 2025.

Speculative Decoding Reimagined for Multimodal Large Language Models Mbq: Modality-balanced quantization for large vision-language models, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.617127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.769305Z digest=sha256:75b5457f6849d6ff8e6de8df704ac1aa726be1739494a633405c7d0fc2a09e9c

Observation 5240c826-d93f-4cdd-824b-a258b39e2219 · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Tokenpacker: Efficient visual projector for multimodal llm, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.352493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:53.976045Z digest=sha256:764c29e7ca080558556fc97dc152be9486a5c6c5edb43baa43dbce0759a1aa6f

Observation 67046b93-e169-4e1e-96d0-0e3727a4473c · outbound

This paper cites EAGLE-2: Faster inference of language models with dy- namic draft trees.

Speculative Decoding Reimagined for Multimodal Large Language Models EAGLE-2: Faster inference of language models with dy- namic draft trees

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.132549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.072661Z digest=sha256:be933566d92c4d03588ca30dfbab4631d37fced287fb5f0fed6e46e00b787df5

Observation 44eecb0f-c787-40ae-a6bb-4606fbfc0afe · outbound

This paper cites EAGLE: Speculative sampling requires rethinking feature uncertainty.

Speculative Decoding Reimagined for Multimodal Large Language Models EAGLE: Speculative sampling requires rethinking feature uncertainty

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.916324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.172704Z digest=sha256:8dce4386b84b644d603c203f80c24b9d904416e6e176098787f43b4e21055a82

Observation e58f0841-c3dd-4168-914c-81caba70ba05 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Speculative Decoding Reimagined for Multimodal Large Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:54.275560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:54.275560Z digest=sha256:b67ceac181337cd719b1837317bac4ef4bca8e7eba55430a70bbdc287c70cc4a

Observation fb07f3c9-5560-4d84-ae63-8a5668afcedc · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference.

Speculative Decoding Reimagined for Multimodal Large Language Models Boosting multimodal large language models with visual to- kens withdrawal for rapid inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.756646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.379762Z digest=sha256:ec615122a5db4eeb49c3fb1a5b0f33076020846c3944dd678ce4f6d5bd68ea6a

Observation 7329108f-1ae6-449f-965e-d2f725a69bb8 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Improved baselines with visual instruction tuning, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.538357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.468070Z digest=sha256:677904aadafd280cbb6c8cdba513b88be86ef396f988eefeb4dbd427855d4e70

Observation 35094642-eb29-49a9-8ad5-976e6191ee16 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.419555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.618384Z digest=sha256:563e8861c7bed63abf0ceca5bf3c03b471798064529e9e0cf9a1dcf8b5149234

Observation 6220cd67-de78-4ac0-81a0-8e3d457447ef · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player?, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.118135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.717220Z digest=sha256:c548334bde96d7ccc6d9838868bfecfbbb48281de8f9d05b1f18ee6e3512ba3f

Observation 986fedbd-0afd-4107-8494-be3dbdc3aad6 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Speculative Decoding Reimagined for Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.857099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.806172Z digest=sha256:ea0cb6b887ae3b0bad54279dfe35eba2dc3100364bca239b307dd33783541412

Observation 207bbc7e-f974-421d-8697-829ca8da4c0a · outbound

This paper cites ChartQA: A benchmark for question an- swering about charts with visual and logical reasoning.

Speculative Decoding Reimagined for Multimodal Large Language Models ChartQA: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.680608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:54.900842Z digest=sha256:81a911bb58ee61ac38adc68a96f455c959ff2d5c2c9984e5a824efe1533fabba

Observation 53475137-9217-4efa-b2a9-87ad353e4469 · outbound

This paper cites Mm1: Methods, analysis and in- sights from multimodal llm pre-training.

Speculative Decoding Reimagined for Multimodal Large Language Models Mm1: Methods, analysis and in- sights from multimodal llm pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.426790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.018726Z digest=sha256:ed19cbb4037ec984ef2b5ed84cae321cc08213b603af3e16a82f03118ed2fb41

Observation e5507623-85a9-4fa1-848a-c3647f6cdb67 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Speculative Decoding Reimagined for Multimodal Large Language Models Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:55.084527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:55.084527Z digest=sha256:96a7281221e71bf99b29aa70bb711b03f021d5830c43f8a5227855d84ab418c4

Observation b7bc08e0-ec7a-4491-a807-8c5a8b5bdafb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Speculative Decoding Reimagined for Multimodal Large Language Models Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.052147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.175992Z digest=sha256:ac7c1f270ba9ac07a6dec5e66ba36353d377634c7c34cc1437e9eccf27f5ed8f

Observation 19239ca9-9e8a-4aa9-9123-15b289b32672 · outbound

This paper cites Crossget: cross-guided ensem- ble of tokens for accelerating vision-language transformers.

Speculative Decoding Reimagined for Multimodal Large Language Models Crossget: cross-guided ensem- ble of tokens for accelerating vision-language transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.783792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.378336Z digest=sha256:74bae4aab2685746df7af8c0b062049cf304b180da1972029b4c81d473649708

Observation f628e393-4091-47fc-9b2e-6495cd4a7187 · outbound

This paper cites Towards vqa models that can read.

Speculative Decoding Reimagined for Multimodal Large Language Models Towards vqa models that can read

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.567691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.483714Z digest=sha256:f93fe6615952d467785f6719ad4ee305777dd9587836afc6d40e1977f917bcf2

Observation 174f5246-7735-4a63-9e22-aa1561093d9a · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Gemini: A family of highly capable multimodal models, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.189988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.579960Z digest=sha256:5e9f58cfd52cef1cb1757660f812f308b2d7fd653f52b59b0d7369f12529c29f

Observation f5c3e04f-90a1-43b1-a083-a396502bbb87 · outbound

This paper cites Q-VLM: Post-training Quantization for Large Vision-Language Models.

Speculative Decoding Reimagined for Multimodal Large Language Models Q-VLM: Post-training Quantization for Large Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:55.672801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:55.672801Z digest=sha256:cbd4909f939a43c281190fb0c843c759ecfca62d9c01c464957146c2aeda7c49

Observation 625280b5-3ea4-4d1a-a6bc-a95f5059fee9 · outbound

This paper cites Large multimodal model compression via iterative efficient prun- ing and distillation.

Speculative Decoding Reimagined for Multimodal Large Language Models Large multimodal model compression via iterative efficient prun- ing and distillation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.935490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.730004Z digest=sha256:4b29f88c6acae9900af274de3b4ad39ab57780135a86d55d472268e618ef3a66

Observation 92345a1f-c526-408b-bc67-86134a9b672a · outbound

This paper cites Ppt: Token pruning and pooling for efficient vision transformers, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Ppt: Token pruning and pooling for efficient vision transformers, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.686430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.802858Z digest=sha256:4354f49393f4c5fdb6a401be4ba741340261c4f178e6c254f48ffcfecd5ec535

Observation 3af64a77-fdc8-47e3-9376-1d68bc731d99 · outbound

This paper cites mplug-owl: Modularization empowers large language mod- els with multimodality, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models mplug-owl: Modularization empowers large language mod- els with multimodality, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.404553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.863908Z digest=sha256:bafa8bd891e110c89269d02407ce42fb14ae4f51adb2fedf6eb7cf63b0f2246a

Observation be0b5308-a3d3-45b4-bce2-be9e6adb24a8 · outbound

This paper cites Generation meets verification: Accel- erating large language model inference with smart parallel auto-correct decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Generation meets verification: Accel- erating large language model inference with smart parallel auto-correct decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.136593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:55.955642Z digest=sha256:257130214873bbdc3f96b4336814a68522a553a4f5cd7b36d1fecc69c138d0f3

Observation a7e61b52-b387-400e-abea-ad29daf30a8e · outbound

This paper cites Draft& verify: Lossless large language model acceleration via self-speculative decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Draft& verify: Lossless large language model acceleration via self-speculative decoding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.950279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:56.015580Z digest=sha256:db555627d2ca3548f0d56db72801c180a34353f493b5d3615ea5cc6447138a5d

Observation da12a79a-721c-4b19-83ec-cd7f8a09eda0 · outbound

This paper cites Learning harmonized representations for speculative sampling, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Learning harmonized representations for speculative sampling, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.802116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:56.084348Z digest=sha256:e45bb41ba8f18eefb92d08664cbe360050e713ea8f5a19d3fc66fd5232741038

Observation 5bee170d-f6b2-4047-acff-1e972a7594c2 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Speculative Decoding Reimagined for Multimodal Large Language Models Xing, Hao Zhang, Joseph E

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.679378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:56.141651Z digest=sha256:bae4801854d5c2bf487ce95b7e6fa63893f3190921263a322690f5f854cef732

Observation 04cc1c5d-e3fa-4e99-b933-ece44387040a · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Speculative Decoding Reimagined for Multimodal Large Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.206898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.206898Z digest=sha256:ac30ff7af376bdadc78cb372bf52cea57585ce6e43bdd9d1ce0f5ee8dbe61e64

Observation 10f9c9d9-75e0-4ad6-bb35-d5d11f6e01c2 · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

Speculative Decoding Reimagined for Multimodal Large Language Models Llava-phi: Efficient multi-modal assistant with small language model

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.539136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:42:56.294364Z digest=sha256:033a4bb39c0103bc62222ed63e217c2406ec239a6ecd7b6cab544b075ef49e83

Pith citing papers

Observation d950a3a3-52f9-4710-9137-a53eeb6e7b30 · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:37.884895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:37.884895Z digest=sha256:cad86454d14da670613f2a9b9102a645420a755d78c443c197fa55a333d71286

Observation 8ef8e9bd-1a19-4d4c-a9c4-16e926bbd613 · inbound

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning cites this paper.

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T19:34:58.789459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:34:58.789459Z digest=sha256:b3f868f804114d2da1eb94e58e07e281c4797bfdd7019a56fd4af2d67a395065

Observation 2212655e-0e7c-4a7f-b6d1-3393eb50a814 · inbound

SMART: When is it Actually Worth Expanding a Speculative Tree? cites this paper.

SMART: When is it Actually Worth Expanding a Speculative Tree? Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:02.508599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:47:57.421156Z digest=sha256:2d2c46b6c741c79175f364b2b921b0e7bd8d98c3bcec916e1811cc6e54e89aaf

Observation a35dae9a-87f5-42d7-8f24-55aeac61e1b2 · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:52.479619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:52.479619Z digest=sha256:6b2dcd0b4a4e19a87b09266174138a8228fafd94c6e98eb1897a022ff60d6868