Pith. sign in

Paper Citation Record · LEDGER

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders

As of 13 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2501.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01709 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:27:13.689128Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:57:21.526961Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:57:29.885926Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf6a0da2-a87d-43cc-8679-884247b53839 · outbound

This paper cites GPT-4 Technical Report.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.382925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.382925Z digest=sha256:2c4a84507491922a68ab71087b4753f54618d049dfbe8376422fcc02da64416f

Observation 72bc8d5d-2859-488b-9811-fdd6089aacd4 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.389227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.389227Z digest=sha256:8dc9ff4ffc94b9a5876d59a5cc741ea48c87ef648c27df9897f2256e8733d240

Observation 394a6094-ea0f-44c6-aa33-56a8d98c6341 · outbound

This paper cites Does Combining Parameter-efficient Modules Improve Few-shot Transfer Accuracy?.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Does Combining Parameter-efficient Modules Improve Few-shot Transfer Accuracy?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.394598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.394598Z digest=sha256:da4b33170c8e8333c2874dc31ba92e0d37c54d8894cd1a1fb15f9a03d45aca63

Observation 29eadd2a-b321-4116-aaba-f565bcd3dc3b · outbound

This paper cites Qwen Technical Report.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.400410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.400410Z digest=sha256:74f634a7b15d097be76df2cb32d86bf47ce00c336a5f9775b3b4409c32011fc8

Observation c1af4069-2089-434c-a895-c190c413f8ee · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.406616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.406616Z digest=sha256:724730a82c4739a611479ce9ba99f254e56f1e7a973efd05eef4c10c353b957b

Observation 2687125c-24e1-48de-ab2e-96b715d75515 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.412710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.412710Z digest=sha256:5d60e61f30a7e670802aefbf8a073fdc6d23524c73b1b453908ce96496393519

Observation e50f52a4-ca56-4c32-a693-2cd55c8f7e59 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.420168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.420168Z digest=sha256:ce19f3f40a67e2544dcdfbefc651c437c6becac9f11397ec59e54d6862269a6f

Observation f2fa7aa1-aa03-43d2-87b6-742b494381e7 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.427397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.427397Z digest=sha256:e0507e8b251e308ef088c344835686b26846eca0c95abf1993ec3b775bd40fe8

Observation 4034b046-be82-4ccd-ac3f-6a160a373862 · outbound

This paper cites Vision transformers need registers.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Vision transformers need registers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.627450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.436265Z digest=sha256:7e80d5fdbf8bab40aa373104514c315f1b64687f8f4a67366f8086975758dcee

Observation 8aec9829-59fc-488c-b5c1-76ea0e4368cd · outbound

This paper cites Eva-02: A visual representation for neon genesis.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eva-02: A visual representation for neon genesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.441521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.441521Z digest=sha256:390fcf634323c9f8453fe5146559ffc8a5174729acadb0772756a73920653832

Observation 65d36236-5e96-4cbd-9627-66da1a147166 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.447468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.447468Z digest=sha256:39a1385cc59ecd93d2d5f11401f97b43108bad62b0d9ff1c8f15bcf5bb25a165

Observation 578d87cc-5f4c-4c24-8401-dedc8188ec17 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Dat- acomp: In search of the next generation of multimodal datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.452669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.452669Z digest=sha256:26cdc8ae2de2bf52b1d8374aad75e412ca39dc603956c5e327f31ffe95b0a938

Observation daca6a56-0b16-47e8-b833-314ee6d83281 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.458349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.458349Z digest=sha256:054b5f47f714e61a88a618fc3994d8d75fc822a23adce923ceb6e77d456e5fc5

Observation bba2ffc5-ee89-464a-9fbc-86b4d9ca3449 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Vizwiz grand challenge: Answering visual questions from blind people

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.464635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.464635Z digest=sha256:15865ddc5e1400c0245aac54c9699685b51273b008fe8bf7509e169ba7cd287f

Observation 488b43f6-0e74-4076-bee2-cd45e28ef5b9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Distilling the Knowledge in a Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.470237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.470237Z digest=sha256:87027596c64f1e40bc7896277cc21e0d1451d1e88c3b09b37ac1732cc72be266

Observation 5271c9cd-2ae0-461f-85cc-35d4c849efe7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.475542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.475542Z digest=sha256:4ac8219727724db3b54b43b3c1c953f605105980a7783611e6ce21a29f3b9ba3

Observation d5f3b0c7-f067-4180-a8e8-8e2f027fb0d0 · outbound

This paper cites Masked Distillation with Receptive Tokens.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Masked Distillation with Receptive Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.480425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.480425Z digest=sha256:414adfef0e67a3cead39ce259921d8413f31e9618156acb41fb82fcad9606a27

Observation f62af99d-6761-4c4f-8e45-5d29316f57ed · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.568334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.486669Z digest=sha256:ea5462ab1d98d8e44b962168263e3ab203161c075fd56fca82ea7688e0f3c059

Observation 87fc07ba-5945-4995-863d-0ebb0777647d · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.492983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.492983Z digest=sha256:46d89acbf0c6ad2158e052259e816859e212c62cfc0b90ebf857ca3e1346f29c

Observation 6154a4ab-ccde-4993-8d97-e860da380d1c · outbound

This paper cites Lift3d foun- dation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation, 2024.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Lift3d foun- dation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.540306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.498263Z digest=sha256:dd3bbb9da8598a6569edd553cbdae575d389a382cfc4a4c9f1135ccd088654f9

Observation abd0a356-d95e-490c-9744-21e05817f992 · outbound

This paper cites Segment any- thing.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Segment any- thing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.507315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.507315Z digest=sha256:b3ba9796b49d449cd90c3cea49cb7f2813fa4f78f2eba0a30a23dbc0fe0373da

Observation 6c29da05-1f08-4f24-a27b-e8095cba03dd · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.513574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.513574Z digest=sha256:1199af7cee2a2ce767961d02c0a8dd795f550c8fc440db862298da359576c1db

Observation cb110152-8f13-4916-a61a-3729b7c2f7d1 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.519870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.519870Z digest=sha256:4d7018537f49eca7ac653757094849eec3fa2338bf6d4902e01c0caa9895702d

Observation 101b2808-9905-4246-8aaf-4da0ed31e7c7 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.527145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.527145Z digest=sha256:126f00ea8f8b7b6a0cad82526bfc522789a6d77f6b92c165a9ff6ee5d50e74c8

Observation c2d3d16b-74e6-4b1f-8485-73d85e27d9b2 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.532542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.532542Z digest=sha256:e0add9581ae8e4a52bcc9d8bb769d290752d42378092180863159fc1d76dd3e9

Observation 3b8063a9-da3c-4c10-9bca-b042bab36ae5 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.538320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.538320Z digest=sha256:68e8228a5111ff17a8ea834d354da1370a56b51df711eae2c62244d4cd0406ff

Observation 6f68b4fa-2d80-436e-85a4-0f58ef3f9bdc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Improved Baselines with Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.543646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.543646Z digest=sha256:9e2b02784fad9060bbc5e59a71896a9a17b72dc8c0aae60abbb789ae01b60e12

Observation 2974143e-4c3f-48c8-b346-855f5fb5070f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.548511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.548511Z digest=sha256:ca2a5ff361de09e2982ec55de21f78aa311f6b59c7738d95beb525ca1de88428

Observation 50f0df7c-440d-4b04-9783-e3aab29c2115 · outbound

This paper cites Visual instruction tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.503141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.553737Z digest=sha256:a475c9d1cdacae4fd164e89763ad1853c594c8bf8dd742947c7d340ff5ae20a8

Observation 7eb9eb38-33d2-4e57-b0f5-ef1c3e9b7ff7 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders MMBench: Is Your Multi-modal Model an All-around Player?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.558488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.558488Z digest=sha256:05f68dbf68108d4d321de718a52c2dd9ebbfc1c0a174d2ac453a22c9377b7583

Observation cbafa25b-627d-46b0-85b8-438fd0be0cda · outbound

This paper cites A convnet for the 2020s.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders A convnet for the 2020s

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.564407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.564407Z digest=sha256:0612dd81f3d02b414c28fc3a1f5f798cb39609114b6d81b8fa62865c840e8ddb

Observation 205b5671-f1d1-48e7-9646-48d2498a0787 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.570609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.570609Z digest=sha256:58b0285a8547f3a8f6f9213da5ef75d0c42c1ebd8dfb983b735bcb265f491822

Observation 80498a9c-ada8-4563-95bd-d9a85ce414c6 · outbound

This paper cites Llm as dataset ana- lyst: Subpopulation structure discovery with large language model.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Llm as dataset ana- lyst: Subpopulation structure discovery with large language model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.460442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.575854Z digest=sha256:4238f5f5b90b7ba300a00eaf42f81d885398f794c96ad0fb186cfef460b009c0

Observation 2607c897-9f4b-40d7-b82c-a561a42ab248 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.581901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.581901Z digest=sha256:ca0eee01f6d32a03daa5aa7e701ebae8446d46d23feec1f5a01278f8a35340c4

Observation 12b5d433-b490-4924-927a-714bfe6a51f1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.442625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.589164Z digest=sha256:9a45d755dd168b06b0e4b7d8af3b9dec86e6cbde1051a36729291953743837da

Observation fdfdba85-9aec-4478-9b2c-c5f8a7ad38ee · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.419028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.594849Z digest=sha256:a0371ceb33e2aa2758f8cd0d65761a0b219c40834b2d085b0bfba1bd23e8cb08

Observation b8b51a70-6c67-4186-bf91-29f9eda3a3ca · outbound

This paper cites When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.397989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.600081Z digest=sha256:73fbf695eb2635690d35015caff0b4fd3c28348c7a1c194738aaf2d608b88b25

Observation cea5852e-cdf8-4cd4-90d7-0adf3091d588 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.606238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.606238Z digest=sha256:0fef5fd1803e548dc8cc4449ea55817dd74f24411106f677ce168f7b26c7005f

Observation 1fb87df2-bf6a-4d24-86d2-b06f143106d1 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.611917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.611917Z digest=sha256:a080cca5197fb642e7c6120bf053e943081c66347220b740efcc95e53d489af3

Observation feab3b34-00d0-4463-af1b-f2373b6e3400 · outbound

This paper cites Towards VQA models that can read.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Towards VQA models that can read

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.376638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.617983Z digest=sha256:dd12f5e300feb0477ad6869c5f08eb565480a4da14fb7eeda9873c3ee33f44c0

Observation 537f842c-2f51-4b1f-a934-71619aa3f266 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.623266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.623266Z digest=sha256:c714c5ebd658856fe8f76f4d1ac9532cd039adef7ae27d7022000fcc6762e5a8

Observation 06dddca7-3899-4452-a179-9dfe1bb073dd · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.629155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.629155Z digest=sha256:8eb4a965de7ff0e928325d314600a53099131f234bfcd20273bfce01d811da88

Observation cc810408-1203-4a73-b879-54a5a7f7a875 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.634602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.634602Z digest=sha256:157292af40e212ae543816bba08cd7ee461ade2a066ed918f98ac0230cc8ee6e

Observation 6fd74cd9-03fe-46e4-b28b-161e108eac19 · outbound

This paper cites Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.640466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.640466Z digest=sha256:a798e1c7bbec5fc1f92232c51c949673de801a7ba760eab9c2259e7018b3af27

Observation f461571d-9743-41b5-a747-724556a9ebd8 · outbound

This paper cites Beyond full fine-tuning: Harnessing the power of lora for multi-task instruction tuning.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond full fine-tuning: Harnessing the power of lora for multi-task instruction tuning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.358827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.645663Z digest=sha256:3b40bafb0811ca2fc29515565baaf1a9afe7b16c6185cfc2a0e2f386c7c2e163

Observation aeb84c7c-b67b-4163-8095-7d3114642c16 · outbound

This paper cites One Student Knows All Experts Know: From Sparse to Dense.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders One Student Knows All Experts Know: From Sparse to Dense

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:27:13.840003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.651129Z digest=sha256:edaca338e314e5fe6582ff3269a4afbd2778bd7103f6ba9d5265c8fb40e06d5e

Observation 360e5dd2-430c-497c-b871-c338e1d837f6 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Baichuan 2: Open Large-scale Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.656119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.656119Z digest=sha256:1898db52e2d9b31784f8c2e0a77f2067b79738b1e551db4085fb8d44c35da888

Observation 739e32ef-aa33-4050-b254-3065d0eade26 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.661456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.661456Z digest=sha256:342f31c4a33a48a014c0f4538f6d4678dff875514ef11e7042e801982ec058e3

Observation 70e96a49-6bc9-4699-b345-d73b97b094ee · outbound

This paper cites Learn- ing from multiple teacher networks.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Learn- ing from multiple teacher networks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.341911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.667871Z digest=sha256:7f7663c6bfff8f307b349cd1f44988751fd9b3810e63305f195dc35244fe2b90

Observation 99d02520-c2a3-488b-b4a0-cf37515833be · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Yi: Open Foundation Models by 01.AI

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.672970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.672970Z digest=sha256:c200829449ca985729aedc28ef238038190e2c3246fdd1beeae89d0020d9f967

Observation 44a85115-9640-4250-aeab-06739e43206d · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.678455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.678455Z digest=sha256:2aadf541bb66c3a2dd0969549e01e06d8d6ed3c3337006cba566d06c8fb4b3b4

Observation df1d01a3-4cf9-421d-9494-3812d1d04ec7 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.683823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.683823Z digest=sha256:612bf917f570773c9dac37cba2537986ace52813b71db14819273997d951154c

Observation a5c23181-1923-488d-a002-443d8b833862 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:14.322850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:27:13.689128Z digest=sha256:24b619ca197c257f0be7d34b866f65574caba10aa24410ef17cc7cc63f683fef

Pith citing papers

Observation 67cb66a0-1f7c-4564-97a1-625fa7ed7c47 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:57:29.890113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T23:57:21.526961Z digest=sha256:0eab9642181e8d71ebffa19efee82d1aeb11b6c43129b5ca6e676581639a5c22