Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:41.586652Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2607.18666.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:41.586652Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:50:07.266379Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:13:54.329150Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57684541-261b-488e-9973-dbc2176ff3ac · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c50d2d-29da-42ee-90ed-d6ba7bb606b0 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242164de-1c8d-46b7-a021-cd9b27f28338 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Cross Modal Retrieval with Querybank Normalisation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf783d6b-2f12-4a0f-a51e-a82bd4abfd61 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VGGSound: A Large-scale Audio-Visual Dataset
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296b60f5-0f1c-4cfb-9137-57dcd02e2971 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Debiased Contrastive Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4202d1d2-ba6c-46eb-92f2-b77b47d1e43a · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46ce0a0-409c-4b8c-bad2-cd897e7b0403 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Scaling up masked audio encoder learning for general audio classification
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb70593-9ea8-46d8-870c-9f531e446cea · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Clotho: An Audio Captioning Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2afe84-bda0-4101-a002-d9f0b655fa38 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb4a74b-5154-47f9-9d28-fbcc3f8f59a8 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio FSD50K: An Open Dataset of Human-Labeled Sound Events
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4e6d03-8add-4756-a62d-18245f5a083d · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ImageBind: One Embedding Space To Bind Them All
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3493da3-643c-4812-a2c6-07c005deb742 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a70e26-d842-459e-8545-be39bd557b7f · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df098999-84fb-4e24-a248-f25fdd37b765 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11bb537f-801d-4aba-a5db-69bbd204d499 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d425ea7-cb0e-463e-9cdc-f29fdaf1eb04 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Matryoshka Representation Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69113afc-ea52-4ef1-975b-5bf0ac9db095 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97cd49ab-0e2e-4030-b8c0-f60474ab909f · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ea5032-9e87-416a-a7da-f8a923c320ff · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1505e678-0dd3-497d-9131-d24882b32d28 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae42b79-20c9-403c-94f3-ca2738838678 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Learning Transferable Visual Models From Natural Language Supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5afa41be-7b07-4e50-a3a3-3bc5122dab85 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Robust Speech Recognition via Large-Scale Weak Supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21cdf96-475e-4b8e-99ad-456da980eb74 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Contrastive Learning with Hard Negative Samples
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1722764a-e47b-4f7a-b2c5-3704b6d0fc5c · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Deep CORAL: Correlation Alignment for Deep Domain Adaptation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a516517-397a-4f39-9f64-924ee457cd8b · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16fb1421-1626-4ff0-8819-e61a07bb8cec · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c244156-7a6f-4b8b-a694-fb7c9f93d971 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d92c79-17be-4057-992c-9775f701c4a2 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio LLaMA Pro: Progressive LLaMA with Block Expansion
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 466884e3-e9a4-4cb0-8fe1-c4c759f3ddd1 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bdb6fd7-8c7f-4bb0-a7d5-6875b20e7d87 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2ca9a1-f9c7-4202-b1aa-55997cda9d8b · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5044bf-9698-4589-9779-b99ceea62a69 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Cacophony: An Improved Contrastive Audio-Text Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd0d16ba-1e07-4306-9b19-f2682e1fc1be · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 38th International Conference on Machine Learning (ICML) , year =
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa801a38-2d2c-4f50-a062-a5a306221f7a · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599f4eb7-17db-44f4-9c7b-faad7eaed3ea · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e781a5ab-81a0-4df7-b7fb-4e3e8fc0e30b · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfa1100-0031-4ca3-82c9-ecc7068e01b8 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv preprint arXiv:2503.22104 , year =
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061103b7-6c54-45fa-bbc7-636c46071a40 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 952714c6-b091-4fd4-8dfb-bffe40bc37f9 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7a9ca5-6e72-48a3-bfbd-4b9b9141d9ff · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a98c961-34bd-4f77-af8f-8f28281cc7d2 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2415df8d-c687-4351-aed6-2560df305a3d · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio NeurIPS 2024 Workshop on Audio Imagination , year =
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1289a50-85c5-4696-8087-93709f687896 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL) , year =
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d312b185-47de-430b-9a07-43a04289dc71 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv preprint arXiv:2602.18010 , year =
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc09cc0-077c-4eb5-8efc-3d8d06d458b3 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f6a33a-30a6-4d31-887f-7209a011635d · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2026 , howpublished =
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ff43d4-5367-4828-9ccd-f27d986b2e10 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b6ff79-97ed-4a18-be6c-0d86169c30eb · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7259845c-4ffd-4294-8ecc-1215f35304e2 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e71f96-af87-4432-8e3c-dc9f73061587 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2025 , note =
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08b9b34-db73-45bd-83e1-18c7d7791da9 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1e66ad-1c0a-4d6b-99a7-3319471a8826 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Qwen2.5-Omni Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb53797d-d1a0-482d-93c2-7f5783ddd66a · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 40th International Conference on Machine Learning (ICML) , year =
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f793620-dc88-486f-bf46-aeba4f56aec4 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04eb5247-40af-4f8e-af5a-04d69deb0837 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2024 , note =
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ac5066-cd7d-4ce6-9123-5feb62406751 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bc642b-c236-4b50-857b-a291c80e89dc · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Representation Learning with Contrastive Predictive Coding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb58c2a-665a-4d20-a1d9-1552299344fc · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio European Conference on Computer Vision (ECCV) Workshops , year =
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbaca78a-f0da-416d-88ef-0dbfeeaa8cdf · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of NAACL-HLT , year =
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b90cb9-a1f3-4862-8d5d-132ec57dfad7 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2025 , note =
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f89ec496-6970-4802-9775-eb2cf4c13327 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c95594cd-9c5d-44f8-92db-ec7b4d5b23aa · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694240ea-f39b-454b-9f16-0dc740280031 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef6708a-a80b-4f44-b2d0-7c83bda471da · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 23rd ACM International Conference on Multimedia , year =
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ebc7435-2115-4e6b-bde4-3403d1f32732 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6359bc2-20b6-43c5-9391-5569c67d2768 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65fe257b-aab3-4299-b080-e6cddc37614b · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac82362-830d-4c79-b817-39cecad1d3d9 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871750c1-3fb7-4a76-9732-f124b8e0bf12 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c46df3-0d70-4dbe-89ee-082ef7a5542c · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b437d22-4816-4edf-84e9-98204630fe69 · outbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Machine Learning (ICML) , year =
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21b671d-365c-4e1d-b228-490938aed096 · inbound
Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e8ce1989-4146-47c9-8c32-b3bb1f51097e · inbound
Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.