Pith. sign in

Paper Citation Record · LEDGER

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

As of 21 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 8 inbound Pith citation observations for arXiv:2412.03324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03324 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:35:24.414685Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.887036Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.832612Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b80bf627-b4f5-4c8f-92dd-4d45c5596364 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.131649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.131649Z digest=sha256:d83c46a2198a3891e088e2a10c4c4c4e6ebf583cf64c0dd307533beb346755b0

Observation dffbe916-ce9c-43fd-814d-26e8436859e9 · outbound

This paper cites Token Merging: Your ViT But Faster.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.136453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.136453Z digest=sha256:74dd9a846845e310a6c17387553fcc038179b80c77f1ea88f232ee9f6b8d9da7

Observation 597db8eb-6cfd-4ddb-b0ea-a777332ef748 · outbound

This paper cites Language Models are Few-Shot Learners.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.141236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.141236Z digest=sha256:3187b39b1e367e7c8211a5950b5585f860dddf6f5fccab03cbdef05243a89924

Observation 8bfbba58-3684-481e-96d2-33a833a80c18 · outbound

This paper cites InternLM2 Technical Report.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs InternLM2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.145088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.145088Z digest=sha256:8161d8740278299217362ec2bcceddceb18a904217013b6a1ac572fb5eaff602

Observation 7a680d99-9396-4a95-a4fa-dd313db57ba8 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.150592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.150592Z digest=sha256:d3161f9d7b2959fd24794c34e9794ca011d5670fb9551056af4ca8175778a243

Observation cf64be47-647d-4e7a-b025-8120b4953f05 · outbound

This paper cites Diffrate: Differentiable compression rate for efficient vision transformers.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Diffrate: Differentiable compression rate for efficient vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.598509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.155420Z digest=sha256:a4b881f9290e745b3ea3bfec01d033bab33204ce108eb46efdf5c8b31fdfe88e

Observation 87d881bb-902c-427c-ad94-3bd277521ab6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.159958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.159958Z digest=sha256:dc8b94a51b47855d3722d878c074030965cf08cfbb6cf4086fc7d31cc3686be7

Observation 7bb2ad88-8998-40fc-80cb-a5cbe7647c19 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.587132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.164533Z digest=sha256:037c585482dd95b52ecbbfb1656e2ff1339a67967c142f1941b64ef330c9aab4

Observation 59c7c511-e49d-478d-b72d-f71617a54223 · outbound

This paper cites Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.575168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.168457Z digest=sha256:71c98f8c037fb86641cfca50ff53da1b9f22e6c2b13478be1227e9468302a699

Observation 21089d75-128e-4dee-8e7e-65fe7f6a4461 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.172408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.172408Z digest=sha256:f7235dd5c076b3cf8ce26ca946bae971673d666b9e501a32d684bcbb54393ad1

Observation 9a348b81-a753-410b-8de3-cc47f50e1c0e · outbound

This paper cites Variational Inference and Bayesian CNNs for Uncertainty Estimation in Multi-Factorial Bone Age Prediction.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Variational Inference and Bayesian CNNs for Uncertainty Estimation in Multi-Factorial Bone Age Prediction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:25.017858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.176457Z digest=sha256:37fb52336062f7e79e44bcf2abbd79037ddd0453a3b8ac78db5e6a3599760c46

Observation 9a0bda54-eb31-49e8-a12c-c55887089ee2 · outbound

This paper cites LM-Polygraph: Uncertainty Estimation for Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LM-Polygraph: Uncertainty Estimation for Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.180459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.180459Z digest=sha256:105134c614d5bbca045e5ddb376a5c0e14c126a3e09f139cc706ede07973e4fb

Observation 26ae7411-ae43-4cc6-aad5-088bb7696a03 · outbound

This paper cites Unsupervised quality estimation for neural machine translation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Unsupervised quality estimation for neural machine translation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.563403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.184754Z digest=sha256:366eca13813e87f2aeae136ae1f5e09e47c6ba56be6a46ac3c3d3aa8c704e275

Observation 162d0beb-03bf-4071-9ad6-6e14f8f408a9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.188429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.188429Z digest=sha256:2ef8ad9f953b433e4b4af301e9851de0b5a2d81214c3e41e9bf712bf059be462

Observation f0ba39a9-4c5f-4249-9075-fa8b7cc898de · outbound

This paper cites A survey of uncertainty in deep neural networks.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A survey of uncertainty in deep neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.551305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.192530Z digest=sha256:35615e5961899d06b1f1ae26a964ce13c6272326cf01bfad98f4ea8d93579e75

Observation 8fe0ac03-4f69-4af7-ba62-330c607f4be2 · outbound

This paper cites A survey of confidence estima- tion and calibration in large language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A survey of confidence estima- tion and calibration in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.538211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.196476Z digest=sha256:be6b62c670e6c0daa4966925ca63b7afe951a5acc0945c9a6a0268a37b1f75af

Observation 42afce4d-dc94-4a5b-945d-9aa562181af1 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Model Cascades: Token-level uncertainty and beyond

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.200147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.200147Z digest=sha256:ddfa8237b7f3595710978343cca47a84d5580ce3a28607699dc52dba59857370

Observation 22c5d3d7-51f2-4bda-a0fe-adffab19e0b3 · outbound

This paper cites Dynamic neural networks: A survey.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic neural networks: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.525821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.204411Z digest=sha256:e75a651f2659252bafec9b3a10ea8152e6b688bf4c713dc20575e52c07163bb2

Observation f530a7f7-9811-4456-a029-46dba21dbd53 · outbound

This paper cites Learning to weight samples for dynamic early-exiting net- works.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Learning to weight samples for dynamic early-exiting net- works

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.512291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.208236Z digest=sha256:94704728c5723ef975551fab507a9b063203465f268102252594f932ce12d488

Observation ae31749d-d47f-44f2-bf9f-42032d5546f8 · outbound

This paper cites Dynamic perceiver for efficient visual recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic perceiver for efficient visual recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.499337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.212183Z digest=sha256:3c50f0c827c45c39a7f5e92f58a5a3b0137f8d4b70c4ebfafcb0de9a77e85673

Observation 280d79c8-7b66-44b4-9532-ce7da5ca1a32 · outbound

This paper cites Latency-aware unified dynamic networks for efficient image recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Latency-aware unified dynamic networks for efficient image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.486571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.216022Z digest=sha256:80eafb732ba7de2bbd6e0ecacaa6c441b6712e11ffce8ea521e743df5475d0bd

Observation 547c8ff2-a1fe-4320-9230-035534c038af · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.224741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.224741Z digest=sha256:59df077e46736f43344cc590b71617e509a708cf77207fd66dacc0f6f69545c4

Observation e5a79a57-2da3-41b8-84e2-d768c6fb0dd0 · outbound

This paper cites Mistral 7B.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.228819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.228819Z digest=sha256:c812c4f74fcc61c42a3bbae71ae470181e3304205e11d391b8005ae4d7f78207

Observation ab0e875e-af33-4d1a-aa56-0b50ec8d0b3d · outbound

This paper cites Language Models (Mostly) Know What They Know.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Models (Mostly) Know What They Know

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.232497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.232497Z digest=sha256:61d4b256aa068d03f520ac7d780ab60e1c9fe2160ecf328d8065c8f61af7c5c1

Observation c19f64a2-204a-4c4f-b76b-6669c9e6c06d · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.236297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.236297Z digest=sha256:2f6b8dfe6fd6478780715c6f7b069de77763b12c0dba4ace3c21d74951138add

Observation fabfd2d6-482e-4881-9f79-0f3000040416 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.239987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.239987Z digest=sha256:b1f257cbe55f15cb40335dce6cbf333ded73253a7bd892aa4fb7a75c9a62d5a9

Observation a729fb64-efa4-479d-8806-f2c3748c4c1f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.244102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.244102Z digest=sha256:4268e0d5d83a8dca24666d88b42ef019cc5ab2fe15fbff6bbfd8087a5e145745

Observation f9ace29f-b2c3-4c1a-9975-52a0a6b3eaf0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llama-vid: An image is worth 2 tokens in large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.457111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.248345Z digest=sha256:27c0559c340b351b7af9b83162886cf37b999194d767d71fcd34abee452c82ad

Observation 37753420-09e9-4d16-b5a4-4b970e057cb1 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.251928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.251928Z digest=sha256:57c022a5b3c95cfd160b3519007f32fee3f4906bdc7d146547fa4846cd764217

Observation eacf2643-c1c8-431d-add7-0b837a19ad86 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.255857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.255857Z digest=sha256:5d2bc4abb51ba7dec76b28841ce23825f51a7cec7a6d5b2f37a9d8204b61c2ec

Observation 3c0c87d7-f737-454c-acb4-ab55915478e9 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Improved baselines with visual instruction tuning, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.444753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.260042Z digest=sha256:a999128bce2a81f5e57deb725d9000b113d451986c900671b5d701ab8d0e2449

Observation 7bc46c4d-9efe-4d0d-84c4-536956c22cb7 · outbound

This paper cites Visual instruction tuning, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Visual instruction tuning, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.418833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.263882Z digest=sha256:598735bcdc201bc507025fb99c4da04bced349ffe9f28edc75a99b4af5a783b9

Observation ee878242-9dd9-4d55-ab3b-0d4c9dedd592 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.407046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.267743Z digest=sha256:1c804acb3d0323bd59f7cb5877bb113d821d312a978f26e8982790a28d755b07

Observation 3b44333d-2a1c-4329-9d71-99a573bfff67 · outbound

This paper cites A general framework for uncertainty estimation in deep learn- ing.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A general framework for uncertainty estimation in deep learn- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.393865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.271700Z digest=sha256:acf7ef452fd390e603965129ca91824c4bc7a7c2fe8e6b92d7e3d09213e05fe9

Observation fb6c33aa-7d86-47d8-b5c7-34e380f5c08c · outbound

This paper cites Visual Perception by Large Language Model's Weights.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Visual Perception by Large Language Model's Weights

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.275563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.275563Z digest=sha256:5dde1137abcc50f92f00ae13261eae10ec681884c14ab0c8fb92882578888ac4

Observation 9d0b0378-4832-4928-8740-31b2c9434eaa · outbound

This paper cites Uncertainty Estimation in Autoregressive Structured Prediction.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Uncertainty Estimation in Autoregressive Structured Prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.279509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.279509Z digest=sha256:41a6f7ac16c0bd64facff26b9c7f9ef1c87cb07993c133fb08ffdab6cf32608d

Observation 256de611-863d-431a-a553-3b9a3ad11307 · outbound

This paper cites Generation and com- prehension of unambiguous object descriptions.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Generation and com- prehension of unambiguous object descriptions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.379430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.283433Z digest=sha256:8c72a42232030b04bbc35979f8529de798a9741c885f91e6237cab9bdb7e1666

Observation 80bb03c8-50fa-4f9b-a971-e9f869ebe14a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.287112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.287112Z digest=sha256:264165d8b996bb50124503be62f066cd630237c0c4273c2cd8f7bce67739588a

Observation b65d3200-f144-4f79-ab5c-9f1837cedf35 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Docvqa: A dataset for vqa on document images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.365988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.290804Z digest=sha256:e125e0212b9443d259cbba2d1729148b1c6c298258eca9e8b7300ad118a74e88

Observation 6daa8012-bc72-46fc-bb99-f3ad33ae1f73 · outbound

This paper cites Adavit: Adaptive vision transformers for efficient image recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Adavit: Adaptive vision transformers for efficient image recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.353525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.294063Z digest=sha256:c38747504bf3b7325c0bb6d46b95fb0bdfb6173b16704d0f0461cdd8723ed08d

Observation 750d9100-fad5-47cb-ad7d-a4d9a3e72df4 · outbound

This paper cites Correcting Length Bias in Neural Machine Translation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Correcting Length Bias in Neural Machine Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.297510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.297510Z digest=sha256:86d97548d46d22883b609025036810c6f1b14e6a542e6184a0944d0424e0e646

Observation b315dadf-151e-41d1-b558-947a3f938775 · outbound

This paper cites Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.341116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.301667Z digest=sha256:1e1d194a008c95a722d91fa077afdfceb758d520577e197cff5dd5f0e6125353

Observation 54226b09-ad0e-4586-91f6-30610c57c2c4 · outbound

This paper cites Hermes-2-theta-llama-3-70b.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Hermes-2-theta-llama-3-70b

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.330393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.305277Z digest=sha256:252e3a2ab4774f1a0b6961b6108d928c017f29b495c8a900fc54d5adeaa05c86

Observation 1f7da94c-1103-4914-828a-614fdb2de761 · outbound

This paper cites Nous-hermes-2-yi-34b.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Nous-hermes-2-yi-34b

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.319561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.309019Z digest=sha256:3f260da9fc2dd82ad00c07ab12cce68a9e4db44f7fd6f71a207144446fb58c2b

Observation eaa61918-7a26-4005-9505-0f7e8f136789 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Learning transferable visual models from natural language supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.313519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.313519Z digest=sha256:440be75060acd06bc1657ae74181ffa5b12ac623cf4fbcc58df58f28f42ffa2c

Observation 10c028e6-4ee5-485e-877b-a40dd3b1076b · outbound

This paper cites Dynamicvit: Efficient vision trans- formers with dynamic token sparsification.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamicvit: Efficient vision trans- formers with dynamic token sparsification

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.302495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.319170Z digest=sha256:4a7bed7ce9d1631f3354f8ab025c0d153c3476bc3ee7a8f930716c5d04f139f9

Observation 8048377b-bdad-479d-959c-b7b1e2d4f570 · outbound

This paper cites Out-of- distribution detection and selective generation for conditional language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Out-of- distribution detection and selective generation for conditional language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.288698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.323178Z digest=sha256:664f70f7c521a4c2fa7253a34bf666cc877d8c01c574135cc673727c225a3eef

Observation 5cc36cde-1b41-442f-848e-46e7bdff52f8 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.326720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.326720Z digest=sha256:9998a62314749dae7f3c2a569af880a7e747e50017268a572db9f85f633e351c

Observation 4eee219d-cf0a-4a03-a1a4-1cc5f006f489 · outbound

This paper cites TransFusion: Generating Long, High Fidelity Time Series using Diffusion Models with Transformers.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs TransFusion: Generating Long, High Fidelity Time Series using Diffusion Models with Transformers

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:24.709346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.330001Z digest=sha256:0a89dfb851df60d137fc1e0ac31170417e97ad217e3c6172dec10e7bfe2e4b95

Observation f1d2419a-4c06-42a1-ace2-857c86002f0c · outbound

This paper cites Towards vqa models that can read.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Towards vqa models that can read

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.264339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.333889Z digest=sha256:e505e2da2fdee2eeb5309b558399990667307da29ca5e5096de8ee2e10973f48

Observation a3f622d8-a03e-4e49-a129-718206b33d4e · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.337472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.337472Z digest=sha256:62c96e8cc8be81c6d54907890f7650f4c10adbb67b6bd9cd457d34b47b55c4e9

Observation f410656b-ceb0-47da-8a38-6305159b2b8e · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.341013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.341013Z digest=sha256:777a9502bef9c6beddae7e1eb911e2fbfb45a91096d59790adf3a348c8f3aba2

Observation 06c1507d-9f7e-4a92-99dc-209c79fc2735 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LLaMA: Open and Efficient Foundation Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.344055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.344055Z digest=sha256:80305c107c547a362604415d0bca0713bba980ed2228fa88a61e615866c6efd4

Observation efbd56ec-f1a6-43eb-828f-672efa1fa9e8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.347240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.347240Z digest=sha256:8f68496ee7c817caf241e396cdc17c31448c71c3b9ff3c55ff8e62ae4a201cfe

Observation 8de10cbe-8d0d-4f19-8629-9b34d56b30f4 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Emu3: Next-Token Prediction is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.350738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.350738Z digest=sha256:4e3a01c790ca4b1c405db387f90d800f570e3beb08ec901fcb21dd600ed76cc2

Observation aa046629-c8b7-45d6-a284-e4906f8f7665 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.353965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.353965Z digest=sha256:a79ed02f0bd79ee73bf61f1af96fea8f70e5e39d9477807e29f821275e96cfae

Observation 068bc7e2-f04b-44eb-8bbd-a7b133f09212 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.357686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.357686Z digest=sha256:57a00bd95e183d5211c169145528b72d7851031e94f532fb3c9a983ef3fe31ca

Observation 0a008773-a7cf-4d40-8821-176c5e918590 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.362275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.362275Z digest=sha256:a0f196aec0737d418ef2eb3bca6f14f55e21aaa3a7e7080c668fd6ab8cf45337

Observation 24e3a296-f2d8-4f6a-a149-160a987d8658 · outbound

This paper cites Qwen2 Technical Report.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Qwen2 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.366370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.366370Z digest=sha256:1d5f3456a0951c0c4dd33b8154b14d29c5f5d890d9a6ea412cdd2b53a750e509

Observation 94e33742-a456-418b-8ce9-4b18cac4efd5 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.370458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.370458Z digest=sha256:ffe2e62189c78e54bb40888221c5658e49b13f4fea71211f439ee8cb1c188b4d

Observation 0ee9a5c0-53d0-4abd-9d99-71082969652e · outbound

This paper cites Detection of Word Adversarial Examples in Text Classification: Benchmark and Baseline via Robust Density Estimation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Detection of Word Adversarial Examples in Text Classification: Benchmark and Baseline via Robust Density Estimation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:24.543601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.375218Z digest=sha256:220359e4d99de9fbdd2f729cde144e67ff95f971b9ae6ac09cfc92723074f578

Observation 412e5bb1-1918-49e7-83e3-a5b88d4495d0 · outbound

This paper cites Modeling context in referring expres- sions.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Modeling context in referring expres- sions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.235115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.379188Z digest=sha256:0055d04c68f0b94036ba2fe2402215de2feb90f1676d92609389ce809150014a

Observation 136acb78-8d3d-47d9-8e9b-927cd930b3ba · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.382882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.382882Z digest=sha256:e209c96db2deabd98120e8aa64bdcf212d78018a55575d2457f6b996a2e87547

Observation 67df8694-328f-484f-bca6-b88b44094052 · outbound

This paper cites Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.386895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.386895Z digest=sha256:64a214298886f601ffbffc1d1ec0ccb44370e1b75de730ea93833d9bb3ab0755

Observation 185df56a-4a07-48c2-9de0-536b5bd57a45 · outbound

This paper cites Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.205492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.391458Z digest=sha256:c22cbc35318cb13b3073ae3770db8c8f146655eb08f2c466d038c880b7540c24

Observation 126c116e-66f3-4582-89ce-82949449cbe2 · outbound

This paper cites Sigmoid loss for language image pre-training.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.131544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.395578Z digest=sha256:cd46b14c8551877f36ba0dd46331fb9feb3188b1805e3dd2701035f425760a67

Observation cf0e0db3-b9e4-457f-9d14-79a47e836c7a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.399461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.399461Z digest=sha256:a8d79603a8af2f87d000d60938e80935fbfce4ba1afef77f5ec204aab358d05a

Observation af0d7fdc-e02d-4373-92f8-76bd17dcb453 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llava- next: A strong zero-shot video understanding model, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.403563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.403563Z digest=sha256:d93a850799dff6456e78df83d93e7423cd651b82a2b8ed593f85536ceb141e67

Observation 202604d5-e842-4a4a-a66c-965ad95557eb · outbound

This paper cites Dynamic Diffusion Transformer.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic Diffusion Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.407029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.407029Z digest=sha256:aefe1384faccde1a61d543149c0f1c43eaeac5081db66cf5078e5a6ec345c0b7

Observation 45205573-571c-4b89-b3e7-df6f90924a63 · outbound

This paper cites Dynamic tuning towards parameter and inference efficiency for vit adaptation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic tuning towards parameter and inference efficiency for vit adaptation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.111180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:35:24.410825Z digest=sha256:c4266e72af2501cfc2dec1558e793487bcd5704c753460f103c2935d2a3f164d

Observation 9476b23e-cb47-4439-82a8-3a88b722b234 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.414685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.414685Z digest=sha256:dcd448fd63895a7c25f8f6b935c1009a2be600bc812324815071c4eff9bb670b

Pith citing papers

Observation f9dca469-3ca7-4d88-a195-6e300ad78f29 · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.908074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.908074Z digest=sha256:b0c94363e241c13f185120ad25a9203932f00f5630364a17b1807139aac95feb

Observation dd486c63-78d5-459a-be7e-d5ee4d9da191 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.571714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.571714Z digest=sha256:807418fa6b0fe873444d1907b2b01e7f063ef2dabcf84888544acb1657ba6d78

Observation a5b0afe7-b960-49c2-926c-4f6ae91eaa2c · inbound

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection cites this paper.

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:50:08.595852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:50:08.595852Z digest=sha256:efd38eeaf361f7f6601cbae8ac58db87d9686f6e5405223e33f1afe31c58016d

Observation 21413577-52f0-4762-ac3a-89d7f774f25d · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.835408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:80570585ac9b712b9299ff5265bd64ea470c40fb550446782492f72f9d707061

Observation d1e56ae1-92f8-45ce-a098-69b64aae0a9d · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.887036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.887036Z digest=sha256:75b69056f453bd7d33a1491b78cc8d9d8263c4f078179a344875cae737758c4b

Observation bded14ac-604b-402d-8a64-1dfa3143d598 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.547382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.547382Z digest=sha256:4f14e253bdf984b7657c99ac88395f3d1449d936eb32f6877043c24376992503

Observation 40c8c8cd-cadc-47ae-af50-4b537eae3337 · inbound

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding cites this paper.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.639839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.639839Z digest=sha256:7b0f47df1c98d8f0eab38e7ad0c1146b9434813e73fb59b8be9559e0b60cee02

Observation 69f62ef9-cb4b-4adc-950e-d49c7a52aed6 · inbound

CARES: Context-Aware Resolution Selector for VLMs cites this paper.

CARES: Context-Aware Resolution Selector for VLMs A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:44:42.492230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:44:42.492230Z digest=sha256:957d5ae1f254d619f7fab397698ec7d9e0433d04c70dde313d67847200a2649b