Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T18:59:42.949686Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 1 inbound Pith citation observation for arXiv:2604.18360.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T18:59:42.949686Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:37.452149Z
A source-named dated measurement, never combined with another source.
Source: cited_works
4 of 4 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 24e758f3-f4c4-40d2-915a-9144ccc93b36 · outbound
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Is- mail, and Huaming Wang
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a739d04b-8478-412c-ad99-7947c3d3b5cb · outbound
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 10274–10284
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a46be4b-f81e-4621-98ff-aa683d63ea23 · outbound
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Do Audio-Language Models Understand Linguistic Variations?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c92d7c9-e623-4437-89c9-a362addbd4a7 · outbound
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bdb6fd7-8c7f-4bb0-a7d5-6875b20e7d87 · inbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.