Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:43:41.879700Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2412.11694.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:43:41.879700Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:53.953513Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T06:00:27.821608Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b095287b-cdc7-4af5-8189-c1dbf3fd4c9a · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6015470d-8b7d-42e3-b7b4-cca96d21a289 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Textbooks Are All You Need
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfa550e-ead2-42fa-a8d2-ac8358eb7bfd · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LLMs Meet Multimodal Generation and Editing: A Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb9eb8e-87e8-4f5a-8150-19ea002bba58 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5aa68bd-f30b-4e65-858d-7977993ceef8 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8172c3f8-a721-4b5e-8508-f550d531a49e · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Zero-Shot Text-to-Image Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad8012f-d2b8-4168-b533-1dbf434585cb · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In 2021 IEEE/CVF International Conference on Com- puter Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 12159–12168
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 33b339dd-fd64-463a-9ed4-2f4dcbc55f24 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936a8bfd-94c7-4b20-bbb5-a12fe8f43a86 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63acddee-8ffa-40b1-b0b9-95846f211d63 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Improving Image Captioning with Better Use of Captions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d4ccea-e6ad-4a11-beed-c4faca960180 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Audio-Visual LLM for Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8d910e-f776-45d3-b739-4fd24cac4519 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities SNAC: Multi-Scale Neural Audio Codec
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354bc140-7518-4a8d-9e2f-28348bb231a1 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8354e47-abb4-42aa-9939-458d6491b53a · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 16354–16366
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c85f9204-f626-49d7-bbc9-ce286fc6e683 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2f0c67-bdcd-4a91-a1e2-b65f44880e02 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 9637–
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation de52a44d-27fa-441d-a937-1ec31f5bc019 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feadd281-1466-4ae1-96e5-880b312e6605 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e71699d6-ca2e-4f4a-830b-6e0c309a4701 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Video generation models include Zeroscope (Cerspense, 2023), VideoFusion (Luo et al., 2023c), VideoCrafter (Chen et al., 2023b), and ModelScope (Wang et al., 2023a)
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feaf6af2-0313-457d-b292-c6f11801bf44 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 576
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795a0c97-bf2a-43b1-9a13-ebaa2a2f75fb · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 1464
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3237d09-3b9d-46ec-9dc0-50da29f93d73 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In Advances in Neural In- formation Processing Systems 24: 25th Annual Con- ference on Neural Information Processing Systems
Reference 2011
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 223f260c-c00c-4627-9ef4-1277e5a1070e · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be402c9-2cd2-47a7-843b-f576b414bafe · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9daa699d-1640-4a67-8399-4500d6a02ed0 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b8c00e-b17d-4aef-9056-860bc1537b8d · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Jukebox: A Generative Model for Music
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14dedcb-cd1d-49fe-9acf-919d218b88b4 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e724bc1e-fd5e-4edf-b490-cb9ac53cd8b7 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74a0feec-5154-4dfc-a652-0bf5f0cc89f2 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b897068a-707c-4498-8b53-e96c1b68727e · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2a0404-5fe8-453c-8c4e-3c059110fc3b · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Licai Sun, Zheng Lian, Bin Liu, and Jianhua Tao
Reference 6121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e45a59d-d3cf-4ddd-8bd8-dfee93308a3d · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Reference 8298
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5c3553-7261-4844-a9e9-5a631ee13d56 · outbound
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities OPT: Open Pre-trained Transformer Language Models
Reference 9662
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df313483-fe6a-4964-a8f7-6a58ae9e0d70 · inbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7c9cf0-32e4-4045-b016-693b24c42168 · inbound
Sample-efficient Integration of New Modalities into Large Language Models From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.