Pith. sign in

Paper Citation Record · LEDGER

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders

As of 13 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2501.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01709 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:27:13.689128Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:57:21.526961Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:57:29.885926Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf6a0da2-a87d-43cc-8679-884247b53839 · outbound

This paper cites GPT-4 Technical Report.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.382925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.382925Z digest=sha256:ab2a31730d509298a0817196ada61f56cbaa8d66c9b5027a0fbea0708225be95

Observation 72bc8d5d-2859-488b-9811-fdd6089aacd4 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.389227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.389227Z digest=sha256:757f987aa8aa54b4f18f8fd831baf669c3ba8938de18afcc5941acb2c7418640

Observation 394a6094-ea0f-44c6-aa33-56a8d98c6341 · outbound

This paper cites Does Combining Parameter-efficient Modules Improve Few-shot Transfer Accuracy?.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Does Combining Parameter-efficient Modules Improve Few-shot Transfer Accuracy?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.394598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.394598Z digest=sha256:46baebf98b6911573ee5bcda0c50a602e545d982c474efb9699fffe22b65af00

Observation 29eadd2a-b321-4116-aaba-f565bcd3dc3b · outbound

This paper cites Qwen Technical Report.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.400410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.400410Z digest=sha256:52034d43fb90bfeb0ad335aa1f3999e8c6fc46d85c7201bd8c72fe9fadc18cc7

Observation c1af4069-2089-434c-a895-c190c413f8ee · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.406616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.406616Z digest=sha256:e6bfa39ff13f4fc13467e333eabaa234bf7d7fd35324b0cf85171b3bb0440537

Observation 2687125c-24e1-48de-ab2e-96b715d75515 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.412710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.412710Z digest=sha256:fdc375754552b7830a466d4aa193da80cff61756b047071c3bf905591a483778

Observation e50f52a4-ca56-4c32-a693-2cd55c8f7e59 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.420168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.420168Z digest=sha256:7f4186d23e053602b4bab61160a326b61ee47cc3d0faaef3c536fb69bd7e5873

Observation f2fa7aa1-aa03-43d2-87b6-742b494381e7 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.427397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.427397Z digest=sha256:2a57aecfa93d4e06e2c15a48cb0849a8ed9818c7b1da8a8f72ac19cd86ac621c

Observation 4034b046-be82-4ccd-ac3f-6a160a373862 · outbound

This paper cites Vision transformers need registers.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Vision transformers need registers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.627450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.436265Z digest=sha256:32e5b7f4592d63bd933a47b09a2f5f684b1a0b8af960a9c6977b0e4d65334897

Observation 8aec9829-59fc-488c-b5c1-76ea0e4368cd · outbound

This paper cites Eva-02: A visual representation for neon genesis.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eva-02: A visual representation for neon genesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.441521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.441521Z digest=sha256:6aa4d92e14daf3b596e74964ebc528437ee9af57eb84647bc0c32f593222f170

Observation 65d36236-5e96-4cbd-9627-66da1a147166 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.447468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.447468Z digest=sha256:e4a4aaff0e8d5b0c914988f00fa43f150e201ff5f8e07c2c0b566dcf840400f5

Observation 578d87cc-5f4c-4c24-8401-dedc8188ec17 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Dat- acomp: In search of the next generation of multimodal datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.452669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.452669Z digest=sha256:1d4092c71acfd5d3642daa5825c2c816a7e3ec2effb0ef2797923c6129ba79b2

Observation daca6a56-0b16-47e8-b833-314ee6d83281 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.458349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.458349Z digest=sha256:2375b4e8dcba76f10b716686102089ee59aa80d3b4e41e954bb4d3208f5664f2

Observation bba2ffc5-ee89-464a-9fbc-86b4d9ca3449 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Vizwiz grand challenge: Answering visual questions from blind people

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.464635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.464635Z digest=sha256:9fb52fafcbafa6df8be47438a5ed17c16a66b3996b8bd5000876be9dba82182b

Observation 488b43f6-0e74-4076-bee2-cd45e28ef5b9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Distilling the Knowledge in a Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.470237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.470237Z digest=sha256:2071508792e468c63b44a8be336dca913de23d37cc72adb409e7df4690ef97ec

Observation 5271c9cd-2ae0-461f-85cc-35d4c849efe7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.475542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.475542Z digest=sha256:4e14156e6c17abf3b2b96e67c44a2918ce25cb670770b993bfd8431222449db6

Observation d5f3b0c7-f067-4180-a8e8-8e2f027fb0d0 · outbound

This paper cites Masked Distillation with Receptive Tokens.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Masked Distillation with Receptive Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.480425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.480425Z digest=sha256:9d34ef09b46ea885e0a828fe30a69eb78d766668334949e54202922bbf3d097c

Observation f62af99d-6761-4c4f-8e45-5d29316f57ed · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.568334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.486669Z digest=sha256:7f51dd8792adcf988eaa3744f6900e9178c0c70f2b5e14b484b01091212318e2

Observation 87fc07ba-5945-4995-863d-0ebb0777647d · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.492983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.492983Z digest=sha256:05133c53f504ccde0d4fe6f80a0881396a104f86ff89e5a0a121553cd1c887d6

Observation 6154a4ab-ccde-4993-8d97-e860da380d1c · outbound

This paper cites Lift3d foun- dation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation, 2024.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Lift3d foun- dation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.540306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.498263Z digest=sha256:f2a36061297818f4ef5d3b5114c291693481433c56e040c0e62abf98f38a082b

Observation abd0a356-d95e-490c-9744-21e05817f992 · outbound

This paper cites Segment any- thing.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Segment any- thing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.507315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.507315Z digest=sha256:1ff06159d6b8964e210ead19c2cb7df5f7fdfd189e946945988fd664aab08218

Observation 6c29da05-1f08-4f24-a27b-e8095cba03dd · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.513574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.513574Z digest=sha256:ec6e2ca976dee8c9c3cc93a83597345034aa7aaa3ff608fde3cd7f6fbce0d561

Observation cb110152-8f13-4916-a61a-3729b7c2f7d1 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.519870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.519870Z digest=sha256:3fb4974f05a60e6ca29925278f1e7add5ccea32e073c6d3fffe54cba10a62bfd

Observation 101b2808-9905-4246-8aaf-4da0ed31e7c7 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.527145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.527145Z digest=sha256:c21842327bc9b2961460c7e30d26200ef09a9010321e887e2468b1da87d433bd

Observation c2d3d16b-74e6-4b1f-8485-73d85e27d9b2 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.532542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.532542Z digest=sha256:6a9b437422a082e0764e969105763bc427269d6004089b503dcc6ca237726827

Observation 3b8063a9-da3c-4c10-9bca-b042bab36ae5 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.538320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.538320Z digest=sha256:5e422e8f0307eb8d16366d5e54f84731c221a382fc48486cc4c94dc3b43b2cbc

Observation 6f68b4fa-2d80-436e-85a4-0f58ef3f9bdc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Improved Baselines with Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.543646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.543646Z digest=sha256:c154647a4cc9ee1f2e42533cb933b0900d38e78f3dff0878a7780c98e3e1e3df

Observation 2974143e-4c3f-48c8-b346-855f5fb5070f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.548511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.548511Z digest=sha256:2330ce03b59971bbf0611c5af43de931cf7e3eda6e9cae3d86c2d205fb14702b

Observation 50f0df7c-440d-4b04-9783-e3aab29c2115 · outbound

This paper cites Visual instruction tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.503141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.553737Z digest=sha256:b3095d88f2dd7c1e6c85f41ff0c7cf39c8eebdadbb8ff16973f446cee3f3f14f

Observation 7eb9eb38-33d2-4e57-b0f5-ef1c3e9b7ff7 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MMBench: Is Your Multi-modal Model an All-around Player?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.558488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.558488Z digest=sha256:711fad7d048d9d4df5c4fb37a466888d6bd52381b60398afcb984cdd9af309dd

Observation cbafa25b-627d-46b0-85b8-438fd0be0cda · outbound

This paper cites A convnet for the 2020s.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders A convnet for the 2020s

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.564407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.564407Z digest=sha256:1db88d7e33f9094ae9bd9250c8f71551b6e015d945bc10efb29a52dc64910207

Observation 205b5671-f1d1-48e7-9646-48d2498a0787 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.570609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.570609Z digest=sha256:d2ccab8c06af47658deaa5abc5c593311351c8b7127d13fb29f81a0c0a90df48

Observation 80498a9c-ada8-4563-95bd-d9a85ce414c6 · outbound

This paper cites Llm as dataset ana- lyst: Subpopulation structure discovery with large language model.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Llm as dataset ana- lyst: Subpopulation structure discovery with large language model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.460442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.575854Z digest=sha256:b49d58c332a43080ca5d7c0765534789bec9e6fa9486b89bb4bd8e5260d765bd

Observation 2607c897-9f4b-40d7-b82c-a561a42ab248 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.581901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.581901Z digest=sha256:9eefec8c6eda1f18b290996b8b9f4e8629adaa88805e2ef4e8c51b9bf818092a

Observation 12b5d433-b490-4924-927a-714bfe6a51f1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.442625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.589164Z digest=sha256:6e4115c20479bbd0e5ebab9bbe31708a960984b49f70859903b4437bc1003022

Observation fdfdba85-9aec-4478-9b2c-c5f8a7ad38ee · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.419028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.594849Z digest=sha256:fad58c7ae0b96ac57a915ac0f96ae912e635255a65b65585a7afba8b808ba5c5

Observation b8b51a70-6c67-4186-bf91-29f9eda3a3ca · outbound

This paper cites When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.397989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.600081Z digest=sha256:f21a213ce937e685716b83e3227d9abacc599c979c83827a4cff6a6acb8919dd

Observation cea5852e-cdf8-4cd4-90d7-0adf3091d588 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.606238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.606238Z digest=sha256:840f77de563df34aa009de305d691826c2f6c93378584f5bbc85793d72685d2c

Observation 1fb87df2-bf6a-4d24-86d2-b06f143106d1 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.611917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.611917Z digest=sha256:7fb723e0848cc3922f0762b7d5d562964f6ae6e44637b2bba8e82bd7e76ebc2c

Observation feab3b34-00d0-4463-af1b-f2373b6e3400 · outbound

This paper cites Towards VQA models that can read.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Towards VQA models that can read

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.376638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.617983Z digest=sha256:4c68a7aa51cc54eb999e1743ce6ed5c3a60b3326a469cf8f071bce02ebb7a643

Observation 537f842c-2f51-4b1f-a934-71619aa3f266 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.623266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.623266Z digest=sha256:92e9d1ffa4d1f6f619826e78fa268a9c374a70e5eedbcb11115f2c94675d8da6

Observation 06dddca7-3899-4452-a179-9dfe1bb073dd · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.629155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.629155Z digest=sha256:8175630068031d7664bd3f8f5c6ea02f3608b83dd2f4f635ba0233157d30f13b

Observation cc810408-1203-4a73-b879-54a5a7f7a875 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.634602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.634602Z digest=sha256:95078f84efaa90546f17802f1fc7e1cb577110ac3be154c6822db7520518aa74

Observation 6fd74cd9-03fe-46e4-b28b-161e108eac19 · outbound

This paper cites Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.640466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.640466Z digest=sha256:ea559c6f61f56a5e522cd88f168f33adb63fb9a0c47d0417504f57a3b90b7ee2

Observation f461571d-9743-41b5-a747-724556a9ebd8 · outbound

This paper cites Beyond full fine-tuning: Harnessing the power of lora for multi-task instruction tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond full fine-tuning: Harnessing the power of lora for multi-task instruction tuning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.358827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.645663Z digest=sha256:1805226752a3110159c43721c33d3448c5e8993f55d88670691d5744e2a14dbb

Observation aeb84c7c-b67b-4163-8095-7d3114642c16 · outbound

This paper cites One Student Knows All Experts Know: From Sparse to Dense.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders One Student Knows All Experts Know: From Sparse to Dense

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:27:13.840003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.651129Z digest=sha256:18dff55a3e5d7225330c1a070f7e95c3d8d07c8e0643a90982798ffb13e06ab4

Observation 360e5dd2-430c-497c-b871-c338e1d837f6 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Baichuan 2: Open Large-scale Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.656119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.656119Z digest=sha256:8b171e989e1822e5c0e83d6ae9fa938825e2d6143530bd35d99860a83e439445

Observation 739e32ef-aa33-4050-b254-3065d0eade26 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.661456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.661456Z digest=sha256:bd4012e91a048766462b73fddee6fb8aedf559904370ea9fb12c4de790b77117

Observation 70e96a49-6bc9-4699-b345-d73b97b094ee · outbound

This paper cites Learn- ing from multiple teacher networks.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learn- ing from multiple teacher networks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.341911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.667871Z digest=sha256:e5a56cfc13a5299443f4bf568fe9ed0f605ed5e948c1e19190e39ea9ec8564aa

Observation 99d02520-c2a3-488b-b4a0-cf37515833be · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Yi: Open Foundation Models by 01.AI

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.672970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.672970Z digest=sha256:cee6bff02de57dfb15a465af485a9e19bd018560137dcd96992675791cdbf9dd

Observation 44a85115-9640-4250-aeab-06739e43206d · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.678455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.678455Z digest=sha256:08e44b6d4411c42739163f9f9f38ad9bf92fb8159b9138baa42f8f803a613b4e

Observation df1d01a3-4cf9-421d-9494-3812d1d04ec7 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.683823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.683823Z digest=sha256:3c705e5c06f356ff7a1e22969f81d4ac956f7b7720693ef44d8ceab425db6003

Observation a5c23181-1923-488d-a002-443d8b833862 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.322850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.689128Z digest=sha256:f5549cec7be48a15e031e4d91c8babead75b0122efe64ac8268272c91a4583c8

Pith citing papers

Observation 67cb66a0-1f7c-4564-97a1-625fa7ed7c47 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:57:29.890113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T23:57:21.526961Z digest=sha256:aca1766f1cfc6f9358a23c0dd0170c12246fc41c7cce507afb003df802563734