Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.480788Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.01289.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.480788Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:40:00.510742Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T16:27:09.401525Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e5020a5b-2820-412e-922d-8a01dcd07960 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llama 3 model card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 579021a5-be40-4b14-b897-eb5b2b030ef5 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f85b83-c205-419c-af0b-3cd0ddf64249 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84bf1a3-e5d7-4fe6-a94e-c2360674bb86 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Token Merging: Your ViT But Faster
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb5b431-b27b-4176-90e4-5d67fb9a08d4 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Honeybee: Locality-enhanced projector for multimodal llm
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 578e2d59-2d4a-4796-888a-adb393df60e7 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f70929-9fb2-4836-8a7c-a31a4843be58 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3fd7e438-4b00-4406-ab26-73712da16cb6 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d157bb-c303-4343-a09c-d07f1ad997d5 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Fusing finetuned models for better pretraining
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c46c07-88a7-4f99-a008-8230dc87b68f · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1961de4f-e606-4566-904a-1bb5a3654107 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dc888e9-7e0c-426d-ba46-4932d7dcc54f · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine Translation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3736654f-5dd5-4490-a3db-223588607acf · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 432abb0b-d338-4279-8de6-313248a04d2b · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932ccbd9-163f-4711-8a14-80f66a95fb4a · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0765a90e-1a92-4eb8-8f00-fd0f3af35a10 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0014ba9e-0cad-4347-bf10-699e3e4c6b3c · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vizwiz grand challenge: Answering visual questions from blind people
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 36af12b2-109f-46e6-84e8-0ac063c7138f · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Onellm: One framework to align all modalities with language
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83921769-1cbc-4a52-a3ad-df5b70759392 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 11de87cf-3771-4388-8e95-300b97321891 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9d214fd1-cbe8-4607-b7d8-87a6f45498ff · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802b26ba-707d-41ff-8c6b-aaeabe7dd343 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Dataless knowledge fusion by merging weights of language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cd60cc8-56bf-4ac8-a293-760a515e2c89 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A diagram is worth a dozen images
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332759f4-3009-434c-955d-cdd30957e077 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion In- troducing idefics: An open reproduction of state-of-the-art visual language model, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e9980917-69d3-430b-9f95-603a1244ca12 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e915f62-38f3-4445-8129-ccf5dee9a37b · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f797669-e749-44a5-bdc2-08af6022129b · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2998fcce-700e-4121-b0b2-7cd9735a91ba · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031776af-ef46-435a-b03c-2f7c4bfe9e85 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vila: On pre-training for visual language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8bc5a57e-f4ea-4e02-8cd2-d4d1b5d11ab2 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Microsoft coco: Common objects in context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d2cefd-2f49-4e5a-a952-a525211d076b · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mitigating hallucination in large multi-modal models via robust instruction tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e449a05d-4bb9-40ba-b126-78193aa34fc0 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Improved baselines with visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6d7f0d48-151c-4fb3-a4e6-ce8c87379131 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0073817c-eee9-4e9e-bb05-9a1b029e7253 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visual instruction tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 300eb4da-51a2-4cc8-b6dd-493e0edca866 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion BitDelta: Your Fine-Tune May Only Be Worth One Bit
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d3f12d-3706-4a72-8c75-50782528f299 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MMBench: Is Your Multi-modal Model an All-around Player?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df3300e-0170-4a34-a7f2-adac24513433 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97a5506-0c4f-4577-87f1-af88335367f4 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c2f7b5-19f5-46c5-aa8e-c34aa7e9cb31 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Learn- ing transferable visual models from natural language super- vision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6cc9117-80e7-49c2-b2af-c2da7fac13b2 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82fd9ad-d49f-40bb-b243-f5d0d3416c85 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Animating rotation with quaternion curves
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4d0d2f25-67e1-4bf3-bd8f-22e13d25057e · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Towards vqa models that can read
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4bce915-5b0a-4100-9b43-9b89da4f8f55 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visualizing data using t-sne
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 174fae8c-ad11-4a6c-895d-cc5a8550ca61 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Knowledge Fusion of Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d56766-116b-49ec-b54f-7c8135a8682b · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Grok-1.5 Vision Preview, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d522b4b3-3f60-491e-a3bd-ff52039a38aa · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6429dc0b-6a57-4882-b451-55182ee5fb1a · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion TIES-Merging: Resolving Interference When Merging Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bd0b00-1e4d-4529-9caf-f9b31182c94c · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Law of vision represen- tation in mllms
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2605e8c-5450-46e0-a93d-1563c818a0da · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a448751-c2b9-4eca-96e4-baba4f97c04c · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A Survey on Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5486dcd6-2b10-486f-9432-c81a8ef32a2d · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Language models are super mario: Absorbing abilities from homologous models as a free lunch
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba3966b2-3d3e-4c1e-922a-3b6331e1f0c1 · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 558f1a50-6dac-4dea-b57e-712b593cedeb · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2293816a-e9c4-45b7-9785-2041aec0a79d · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af8029f1-3b68-4ea6-857e-e5fb67d7643a · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion SVIT: Scaling up Visual Instruction Tuning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fc70fc-8ca2-46a9-8ed4-77b774c000ba · outbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81cc936d-1ccb-4918-be52-a8ad35fac335 · inbound
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d395fd3b-2d8b-4ac5-9735-002f65d9c99b · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.