Pith. sign in

Paper Citation Record · LEDGER

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.10548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.10548 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:38.969249Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:15:58.111384Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f32fe38-422a-4bc1-9057-a9cef7fd7d38 · outbound

This paper cites GPT-4 Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:31.969501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:31.969501Z digest=sha256:bc07117397525a1cded97c4f0a1cabf8ae7fca9ede80dcf3226db515aeba12c3

Observation e5e7e21f-8336-429d-9e68-3334dfe51165 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.124176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.124176Z digest=sha256:577dbdaf8230249619d8702ac514758f5621d3dec9aec892666880a4e75004a6

Observation 06fd9171-731a-4be9-b22d-45c42578b975 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.253101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.253101Z digest=sha256:67a406861b1cc4f3a25e7c1c2f42bdac0cdbe861fd0d78d0e844cd4f633bcf4c

Observation 98031748-f8f0-4cb4-8c1a-d4615879ced4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.364961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.364961Z digest=sha256:753e1643269de495a89b07914201c7e4841ca6babb3028182daf3c19a2ece008

Observation 0f600248-06ad-45f3-8739-0fa3a07e4d5a · outbound

This paper cites Hallucination of multimodal large language models: A survey, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination of multimodal large language models: A survey, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.516659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.516659Z digest=sha256:0e41234349d7231a10c4623c525df1675c0736e56aa97059f329920efb4e1c7b

Observation cc41d65e-3770-4a33-987a-68cb1b75de11 · outbound

This paper cites Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.630525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.630525Z digest=sha256:5452ed79579a6aa340bc2b246a57febf2642a41ff3d4a6986bbdef2030eb40c9

Observation f19e051c-5c62-4334-8c69-0282bd17e0c3 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.756479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.756479Z digest=sha256:399412cece6515174a0a952a7268d476fb3220103c883d51c2a74e56a50610b3

Observation d6d70d8b-6a42-4d6c-aacb-1020432feb1f · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.914509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.914509Z digest=sha256:4fc1c931dc87115353b58cbe2f43aa088dca89c6cad6a2e6944fa729938c5ad3

Observation 415163b8-b131-49aa-b806-ec6dac56b0f9 · outbound

This paper cites Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.033914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.033914Z digest=sha256:8b9f85cd00a2b5d7b52ffafd652e2731d38a70b284acda206eff88031196b3f8

Observation 132056ad-50ea-4d42-9024-82dddab03ac2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.154299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.154299Z digest=sha256:da898348cc9bb69455f0a3bf0e4e566d866849cf20b3b3dd85eb79b3e880ae62

Observation 76eac7fa-b439-4f1e-adf3-e55ada39e82a · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Spatial- rgpt: Grounded spatial reasoning in vision language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.283044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.283044Z digest=sha256:bf0a07d132a1927d47f9d0ae579f5c00e361c492cc5b77d9adf98b6cd032b196

Observation a29e3834-45fb-4375-8824-f45b17b54b14 · outbound

This paper cites Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.398979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.398979Z digest=sha256:d37dc180fc69d7ad1031617b2994b4acf10d8ea19b818374eaef6c9d553f54bb

Observation ed9dac77-1ad7-410a-848f-e63800e634cc · outbound

This paper cites Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.555242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.555242Z digest=sha256:82678cbcc1a9bf161c02def1cc7c5910f04b5b61c365fe232b0c9813097b18c2

Observation 225a961d-d473-4f67-acc4-d3a208aaeed5 · outbound

This paper cites Acceler- ating the super-resolution convolutional neural network.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Acceler- ating the super-resolution convolutional neural network

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.660599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.660599Z digest=sha256:90326cbe2cb62c7f923e021b05afea7b909f75ce0dd74f69b5c7fa13dd2e0e56

Observation 6ac304ad-c02f-4cbc-a3a3-d9701695e2d6 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Multi-modal hal- lucination control by visual information grounding, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.746918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.746918Z digest=sha256:e62ca05666960aeb556ff607761d0d0568a2c4a6e7429a347310528455225863

Observation c15a0393-fed7-4933-8ab4-c68253dc1815 · outbound

This paper cites DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.923627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.923627Z digest=sha256:22d4b9ef575af32255a32f8f834b39f8e0a1f21170f28af8b57ed591ea4ac271

Observation d04eaf8d-4385-4bde-aa91-7c5edb3840bc · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.073705Z digest=sha256:b14a4a78572d7a6fa98fcc8924a516399251b0b5c51e71087d186699e798380a

Observation 8ea7cf24-597e-4089-aa6b-7b1d31388ee5 · outbound

This paper cites Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.244703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.244703Z digest=sha256:ecad54cc7ff68df2ed3ffe049f516f11c62c636b9cfe08edf0b7c06ea6fddede

Observation 26d508e5-4d87-448e-927d-efc006131837 · outbound

This paper cites Beamlora: Beam-constraint low-rank adap- tation.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Beamlora: Beam-constraint low-rank adap- tation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.426101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.426101Z digest=sha256:b4d5ba5965cdd131d3d0054f5efa18afbce873bcb9600a01c22ed85bdc0f206f

Observation f18c14b3-1ee4-4bcc-88c7-13c64e9d650d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.507754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.507754Z digest=sha256:70e1b102249f18371806f883c26729bfd7aeed0a8ea729ff43657b758540cc8a

Observation cd0fdf4e-338d-489c-9b9b-911e6d7c33d1 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.700795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.700795Z digest=sha256:e9977b6bd10910f6bcfdfe57aa43b01fced353bcd4de3f1d04cd73bfadca6c6a

Observation 57913d8a-3348-4335-b498-5364c074a75c · outbound

This paper cites Hallucination augmented contrastive learning for multimodal large language model, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination augmented contrastive learning for multimodal large language model, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.887592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.887592Z digest=sha256:ee6ff4608c202572c242b9c8a9db6a93159f2cbfb0b4abcf58481d007684fdd0

Observation 3d7c207d-cb2d-42d1-a6d0-27320cf6055e · outbound

This paper cites Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.029874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.029874Z digest=sha256:0eca5cf9532827ab80168055b76869721cbe323d0e32f66d54d9bc80c04bb2a4

Observation 8df33330-5259-4fed-bac0-3ba32b12bcdc · outbound

This paper cites Accurate image super-resolution using very deep convolutional net- works.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Accurate image super-resolution using very deep convolutional net- works

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.187960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.187960Z digest=sha256:6f4230bc004b8e5db39d52a59ef539004b96538ee442d2cee7c146262e66c45c

Observation bb761497-fc53-491e-ac5c-fced096a0003 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.331706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.331706Z digest=sha256:55edc4fa2116d01e634c1cd2036a6c3fb609f62b043ef4b081a415b2229d9860

Observation 6c64248e-1ef1-4283-96f7-4f6d05e4eb20 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.462066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.462066Z digest=sha256:41f7a7a565a9d935e4e4bb5c215d0e957d2d27488aef9106b5a3736d86e211e6

Observation ba5dbffd-f91d-4d93-9d2c-da234d9df531 · outbound

This paper cites Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.645936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.645936Z digest=sha256:ac195b215ee201ef82b944f59fe459c5ec326815b80cca81298e88a0c33ab912

Observation 8ae0dc74-3ea3-484f-a6f7-c87df8108fa0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.790140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.790140Z digest=sha256:52252ce1974101413d4bbc60f169e4c54a793982bc54081a48743d64b3416a05

Observation c4cd6922-8f0c-40f3-9f69-f5992716d92e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.981648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.981648Z digest=sha256:f563a611c415410af9bc97ccbbe811466f9973782e711dadcfd8410ac17c13a1

Observation b88d4203-8a5c-47a8-b28a-b9da0dce8b4d · outbound

This paper cites Microsoft coco: Common objects in context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.135377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.135377Z digest=sha256:5285a2701707fbac4e9ef557ddb251deb06cc690feb9ed6b31a37d7ddc67edcc

Observation 263bf065-c9b1-4704-87cc-d0fdd74c2cdb · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.260287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.260287Z digest=sha256:c4e7c476207a27157e9bc25c5ac1ae5f406fcfcb3f071bf191bcdbcb7f00be17

Observation bf27882a-4faf-48c9-91f4-f0d0a73e5fb0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.441445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.441445Z digest=sha256:74af9231795e7d6a49c65192a8c02fb139d3b08c2bb32106c01a3c4c1bcafca1

Observation c24e4044-c131-42e5-8cf8-0d76a396a91b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.559709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.559709Z digest=sha256:f3929664bfa9d973a62565f4271cc3c446d4a6aded45eca823ee0a11e151e6b3

Observation faba4aad-1c8f-4096-bcf6-b7274439f3bd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.675881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.675881Z digest=sha256:92e5aae76e15e5c5b70708397d753c07ac74ba3d50412865ebeb8ca255e6a180

Observation 114edba4-7afe-4b80-883b-69254ee70302 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.738767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.738767Z digest=sha256:ee77893fbb3ea7ebb63f60c661dbba6addfa64986105f2b2f2468bee320bbb39

Observation 1b184aa7-8a20-44d6-a001-f40884234dbb · outbound

This paper cites Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.859880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.859880Z digest=sha256:0f5a93bbaffe3eff8163f68c6f8eb768961bda6746a590b47b59035b248c6b26

Observation 98f33d94-5808-4029-8137-268054a72f27 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.974350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.974350Z digest=sha256:e20992e47d70155fa3a71dbc0d052c2ef605da46f553c37aa7657414be2b97a6

Observation ce4b50ba-1971-473a-9dc8-5acc4e0474ef · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.107535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.107535Z digest=sha256:e6eea34284c7a24f78cd313e6b8a887e5173f4a9b63fd824012b49c3889384a7

Observation ec8c6013-d94d-44a2-aadc-f06a26411e2b · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.236417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.236417Z digest=sha256:bc02546c783445e9309938d44b3529ecdaae9c3392c9f8f1bc3a688870a260c2

Observation ac8345c7-e171-4208-a687-284a42c9e5ae · outbound

This paper cites Towards vqa models that can read.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Towards vqa models that can read

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.399420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.399420Z digest=sha256:fe0398775eb413a2cba6959ac70ee8adc147f3c2c87535f1d1fbc6d0eb225d9a

Observation d6c710be-23b8-4ffa-9a21-abb42453bfef · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.581401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.581401Z digest=sha256:26328555543c0e03dd8ba2108ffc4a013dea6d38038979e8dce390af3066b68b

Observation 11c66680-208d-422b-8114-03d3941c5080 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.702564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.702564Z digest=sha256:e3d96a84c5e064fc0d6b26f2c029ea1d6259c1fcfc4c1748084a375876203384

Observation 5640b0b4-4a97-406b-bbbe-6fce424be439 · outbound

This paper cites VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.841750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.841750Z digest=sha256:6e74c62456bc1b8168b99333d7245c5f1dead17865e670a21b6f73151349ff16

Observation 2ec4e44d-88e9-4bb4-b788-08e4eeea8147 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Transformers: State-of-the-art natural language processing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.025746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.025746Z digest=sha256:ca7791c5363f75009072044b09a7a5e62362a879a767c15ca0b053c1c49a290e

Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.165356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.165356Z digest=sha256:e2cbbbabe7e69e6092ee21ec8b372b526779ae9988a10b9a4112914d7f32c927

Observation 8c8e008f-e00e-404c-a8af-c2ad5d00a78a · outbound

This paper cites Grounded chain-of-thought for multimodal large language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Grounded chain-of-thought for multimodal large language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.306215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.306215Z digest=sha256:d31724b6da67f51e9dd132097cf8221eb3d7e63d04486f3cec43e8a711230666

Observation 6635180d-9c07-4ea3-8288-e28d31e73a4d · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.421930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.421930Z digest=sha256:481cecbd05159447a9ce83a0227f4d737440cbd4d63cb90aefee1220d5305cad

Observation 8331d809-d8a9-4c1d-87c5-82b7ec5e5952 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi- modal large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Fit and prune: Fast and training-free visual token pruning for multi- modal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.563170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.563170Z digest=sha256:f86632ad20366a5fc74d640de8386cf35607ac782a2cc1c16fa3f8daeb42995c

Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.749150Z digest=sha256:3b604065f112c6361d66b4de39241d4d2fb20953cc4720ded06192a6c0982918

Observation 1c6baefc-31dc-4d52-91ae-aa6d452ffafe · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.835314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.835314Z digest=sha256:84104095a36a837d629efab2cf05717f563ae38d79523ffb9af9a18fbe7d7dd4

Observation 52371a29-5dbb-4375-95d8-72bfd611be92 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.897509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.897509Z digest=sha256:832dbcd94f97a275de5daefab0dc0cc44c2b0fb1fcf9f975847bd1ebf279cc6c

Observation d0cae116-48b6-49e8-9be2-be8c25934b68 · outbound

This paper cites Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.969249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.969249Z digest=sha256:0519a67d1cd25aa3a4dfa76e25aee7cb40f836bd7604a72b0b2240fed81c74c9

Pith citing papers

Observation fa65d30c-7fb7-470b-95d6-9e823029c87d · inbound

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? cites this paper.

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T00:15:58.111384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:15:58.111384Z digest=sha256:1debbe25f5e4d51e460abb61e91053dfa5f782599061dc537299dc8abf0d720f