Pith. sign in

Paper Citation Record · LEDGER

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.10548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.10548 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:38.969249Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:15:58.111384Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f32fe38-422a-4bc1-9057-a9cef7fd7d38 · outbound

This paper cites GPT-4 Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:31.969501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:31.969501Z digest=sha256:972478ea02a2734db76f9664c148200a1067f650bb21e6e684139d2c0e7bc6e1

Observation e5e7e21f-8336-429d-9e68-3334dfe51165 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.124176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.124176Z digest=sha256:f7c3ae4741ec9b5c5ef075153d28631271704a0a35db77ed8cf2a7f0413ad90d

Observation 06fd9171-731a-4be9-b22d-45c42578b975 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.253101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.253101Z digest=sha256:5bfc90cddec006c70f4505560970594f166f5c39e57ea1cab0104f91df8ead34

Observation 98031748-f8f0-4cb4-8c1a-d4615879ced4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.364961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.364961Z digest=sha256:4c98ad9408c1daa08408c30dae79f039704a82632f68f143b7cdb657db227310

Observation 0f600248-06ad-45f3-8739-0fa3a07e4d5a · outbound

This paper cites Hallucination of multimodal large language models: A survey, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination of multimodal large language models: A survey, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.516659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.516659Z digest=sha256:5854d9c03a712aaa8b7cc7b054aa547cc4bbaaa2cc2f524aa19f50fd76c62ab5

Observation cc41d65e-3770-4a33-987a-68cb1b75de11 · outbound

This paper cites Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.630525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.630525Z digest=sha256:3d75497b0790f6a82fe16d78a6eb56c13bb8c38e7033e595ca3bb7e2c6b71e50

Observation f19e051c-5c62-4334-8c69-0282bd17e0c3 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.756479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.756479Z digest=sha256:22b2ebc207de8c50f95fd348e6e236853c3f33ac63c570853472eb1cd2de666e

Observation d6d70d8b-6a42-4d6c-aacb-1020432feb1f · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.914509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.914509Z digest=sha256:fba82811c9906ebba7a13a2b1272a8362e73ca3a1d96cd5428b6ed1413df506d

Observation 415163b8-b131-49aa-b806-ec6dac56b0f9 · outbound

This paper cites Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.033914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.033914Z digest=sha256:e0dda24f476796a757076f46d5f98c1610dc6b09c2c9ca726672a58d97dd6de4

Observation 132056ad-50ea-4d42-9024-82dddab03ac2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.154299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.154299Z digest=sha256:8f319013deb4260ed425dc2877a16bb0042032e6dc41f0086e60679313c4dcd1

Observation 76eac7fa-b439-4f1e-adf3-e55ada39e82a · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Spatial- rgpt: Grounded spatial reasoning in vision language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.283044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.283044Z digest=sha256:bdf936acf7ba3c9e6e45fcee98c7dc56a4ea26254227716a43f2662ac469bcca

Observation a29e3834-45fb-4375-8824-f45b17b54b14 · outbound

This paper cites Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.398979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.398979Z digest=sha256:57f0bdd6f796bed4e5bd497769fe0f2238bc74fa528bea61653dbce55a71552c

Observation ed9dac77-1ad7-410a-848f-e63800e634cc · outbound

This paper cites Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.555242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.555242Z digest=sha256:79df43e2880dcc6347f94386cca05a60c2229962278e26138ef743402978a734

Observation 225a961d-d473-4f67-acc4-d3a208aaeed5 · outbound

This paper cites Acceler- ating the super-resolution convolutional neural network.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Acceler- ating the super-resolution convolutional neural network

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.660599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.660599Z digest=sha256:a435abf377d82783dae4c2066f7ede3dbc8a60c83e3e7f65b0de79e92335121f

Observation 6ac304ad-c02f-4cbc-a3a3-d9701695e2d6 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Multi-modal hal- lucination control by visual information grounding, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.746918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.746918Z digest=sha256:829e77a8349eb1bc467bf6abee864c4b7a63097160b6d11494c72da5f4054ab3

Observation c15a0393-fed7-4933-8ab4-c68253dc1815 · outbound

This paper cites DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.923627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.923627Z digest=sha256:65564f599dfba8cc2401f40345dfc116c09e100a871c487134c503870bcc247a

Observation d04eaf8d-4385-4bde-aa91-7c5edb3840bc · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.073705Z digest=sha256:5ee5cdb93272e119ae44c0e91a185b7969dff59ef655ac3f69368a07dce96b20

Observation 8ea7cf24-597e-4089-aa6b-7b1d31388ee5 · outbound

This paper cites Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.244703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.244703Z digest=sha256:e6b4487e617a8685bc45c929d4a44fae808558b42cf37f109d3d647335d38d9d

Observation 26d508e5-4d87-448e-927d-efc006131837 · outbound

This paper cites Beamlora: Beam-constraint low-rank adap- tation.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Beamlora: Beam-constraint low-rank adap- tation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.426101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.426101Z digest=sha256:dc34cd281463fce04b4a37c681deaabf21e706041b557e67f508b07e0a6ed82f

Observation f18c14b3-1ee4-4bcc-88c7-13c64e9d650d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.507754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.507754Z digest=sha256:140deaa67a8dca7083a8b737815f67d08dc77257a9ecc893add1284d90f4b687

Observation cd0fdf4e-338d-489c-9b9b-911e6d7c33d1 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.700795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.700795Z digest=sha256:60d75a6eaf8f41d751f61d7a87dbb35d5d1c268c325ac8aa107947f72160cb37

Observation 57913d8a-3348-4335-b498-5364c074a75c · outbound

This paper cites Hallucination augmented contrastive learning for multimodal large language model, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination augmented contrastive learning for multimodal large language model, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.887592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.887592Z digest=sha256:57968ee0fe58bc58fd8626e8e712a1b292cbe7634b5ac373016f28f854ce6701

Observation 3d7c207d-cb2d-42d1-a6d0-27320cf6055e · outbound

This paper cites Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.029874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.029874Z digest=sha256:fa476e816c77f013f184f78e4fe68055532defbeba6dd64da3f900ebabf5929f

Observation 8df33330-5259-4fed-bac0-3ba32b12bcdc · outbound

This paper cites Accurate image super-resolution using very deep convolutional net- works.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Accurate image super-resolution using very deep convolutional net- works

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.187960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.187960Z digest=sha256:47ae9348dba7d9d813ccf1fb45abacd843eb4d3d8a6c4dedf1a98f25e9a812b4

Observation bb761497-fc53-491e-ac5c-fced096a0003 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.331706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.331706Z digest=sha256:8e294ac25bf3803431e9a7ced431d0d4508edfc53958fcda7511c709d6438e14

Observation 6c64248e-1ef1-4283-96f7-4f6d05e4eb20 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.462066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.462066Z digest=sha256:81e1e7b392fb81d0f6729bda8b3dfdbd827d11f5200f1d24ead82d3cf8e4da6c

Observation ba5dbffd-f91d-4d93-9d2c-da234d9df531 · outbound

This paper cites Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.645936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.645936Z digest=sha256:3bc336c0e7ec1a55193ea4f37b67bbdb361af89fb0d7b967508f5edd0a7c3030

Observation 8ae0dc74-3ea3-484f-a6f7-c87df8108fa0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.790140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.790140Z digest=sha256:c2338ec75f8ee5a8b143dc18df27d0888d83ff9bc9b0c3c60e07b8d045e69ec2

Observation c4cd6922-8f0c-40f3-9f69-f5992716d92e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.981648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.981648Z digest=sha256:adbb7d640f448bd3676acfe04f3f44581a8ee78c3b42c251dd40866c73f9fd44

Observation b88d4203-8a5c-47a8-b28a-b9da0dce8b4d · outbound

This paper cites Microsoft coco: Common objects in context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.135377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.135377Z digest=sha256:502539be4b6e2da9ef7c7a491c382daef614d6f893f5614bec7e29048a4a5758

Observation 263bf065-c9b1-4704-87cc-d0fdd74c2cdb · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.260287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.260287Z digest=sha256:4e4e5f36fe2e2401f9ac8967d494b2a3a93073597e941b24889e63aa7cda2ee7

Observation bf27882a-4faf-48c9-91f4-f0d0a73e5fb0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.441445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.441445Z digest=sha256:59078d6ec61d8b098c57af7ec422d22500572da0f1c51849852cf44ac94e1f45

Observation c24e4044-c131-42e5-8cf8-0d76a396a91b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.559709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.559709Z digest=sha256:2c76bbaa352ffadce34323197395ea3ae578f01bd58afa7110ec0e733ca4b8c4

Observation faba4aad-1c8f-4096-bcf6-b7274439f3bd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.675881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.675881Z digest=sha256:d19acda4e2ac3d974dd3148894df24f2b69abd73841e5c8e6b06fe980ce7cbcf

Observation 114edba4-7afe-4b80-883b-69254ee70302 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.738767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.738767Z digest=sha256:2d64bfe6c3655e8c095b01d6da1fb3a2284e8b7f75f786570cbb251137a59b04

Observation 1b184aa7-8a20-44d6-a001-f40884234dbb · outbound

This paper cites Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.859880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.859880Z digest=sha256:74e86abff8a7aeb57b51b95389d07e656caa6aefcb3317206887f919b4494317

Observation 98f33d94-5808-4029-8137-268054a72f27 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.974350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.974350Z digest=sha256:7ad800deb2cdf308395a6845d691d0a0e31e2fad7db2c5453a2175bd5c5a0474

Observation ce4b50ba-1971-473a-9dc8-5acc4e0474ef · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.107535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.107535Z digest=sha256:4859e7a2ac3d456a07627a3452ff54c72c81449e5b9eb34cba2e5f444e0d8718

Observation ec8c6013-d94d-44a2-aadc-f06a26411e2b · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.236417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.236417Z digest=sha256:8a99b3432cd00d71c905b33a6d6fa25620839457be5dcaffaf42e7c128be9a5c

Observation ac8345c7-e171-4208-a687-284a42c9e5ae · outbound

This paper cites Towards vqa models that can read.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Towards vqa models that can read

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.399420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.399420Z digest=sha256:6b15189c8bad3fdbb6b37e085dbd861c0af217da0217ced67277cccd9749e7db

Observation d6c710be-23b8-4ffa-9a21-abb42453bfef · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.581401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.581401Z digest=sha256:ff75101ed1a0fa96ac0abe3df245e483938f350802d0ddf9ba99696c0f2bbf41

Observation 11c66680-208d-422b-8114-03d3941c5080 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.702564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.702564Z digest=sha256:d10a7ba31cb6ae5d51b9347f4303842c6f68c98898ade1f49b17d59c7e16576e

Observation 5640b0b4-4a97-406b-bbbe-6fce424be439 · outbound

This paper cites VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.841750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.841750Z digest=sha256:07a7ba7be049efc531bf5ba1da11c7c3668f99bc4c85fb1610428d161b830f6e

Observation 2ec4e44d-88e9-4bb4-b788-08e4eeea8147 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Transformers: State-of-the-art natural language processing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.025746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.025746Z digest=sha256:788d7f7e1d7f5d82f46cd64f3451bb9e5e08af89c50702d0771c8a260833f9c4

Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.165356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.165356Z digest=sha256:a99e419a4de872c9c7f65a9c3f1b5ac6e756eba7a88f57664011f66ae5c80237

Observation 8c8e008f-e00e-404c-a8af-c2ad5d00a78a · outbound

This paper cites Grounded chain-of-thought for multimodal large language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Grounded chain-of-thought for multimodal large language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.306215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.306215Z digest=sha256:3759f8102277b9164870fd30d5252f2b1067760ed2e409e20b682939b2ae8a41

Observation 6635180d-9c07-4ea3-8288-e28d31e73a4d · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.421930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.421930Z digest=sha256:720e5418d815d8862cecd4c87bcbe5c654489e66bac75ec34de51eb9fc69748d

Observation 8331d809-d8a9-4c1d-87c5-82b7ec5e5952 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi- modal large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Fit and prune: Fast and training-free visual token pruning for multi- modal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.563170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.563170Z digest=sha256:7b2dd3a78b85337cb3b0026b0c35f699a46f694275dcdbbf5f2b15f0fdba97bc

Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.749150Z digest=sha256:66ecec4b1dba2c7e558baa601761207df035dce6ea483656be0eb826429ac4cc

Observation 1c6baefc-31dc-4d52-91ae-aa6d452ffafe · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.835314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.835314Z digest=sha256:ca7899ae48f7dd6c6567c2dc443ad750fcf406101a67e9851d5f7e8597772d33

Observation 52371a29-5dbb-4375-95d8-72bfd611be92 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.897509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.897509Z digest=sha256:591ecf40985e8e0ce9769cb53ae4236ad7dca63645786112e2b03fa3f34ab00a

Observation d0cae116-48b6-49e8-9be2-be8c25934b68 · outbound

This paper cites Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.969249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.969249Z digest=sha256:c19362db79e0b89b41443e5b545c5e387fe36c425af76f68f83eabd0030576da

Pith citing papers

Observation fa65d30c-7fb7-470b-95d6-9e823029c87d · inbound

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? cites this paper.

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T00:15:58.111384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:15:58.111384Z digest=sha256:1b437a7ef1def065dd8acdee12fc69d4e5c2d4a9352e7ad022850402c6d9cc6d