Pith. sign in

Paper Citation Record · LEDGER

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

As of 19 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 1 inbound Pith citation observation for arXiv:2604.18360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18360 v3

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T18:59:42.949686Z

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:37.452149Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

4 of 4 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24e758f3-f4c4-40d2-915a-9144ccc93b36 · outbound

This paper cites Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Is- mail, and Huaming Wang.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Is- mail, and Huaming Wang

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T18:59:42.949686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:59:42.949686Z digest=sha256:3138825ab628532ce28ff94d453d304da30a626dc7ed1dc4484933004717f232

Observation a739d04b-8478-412c-ad99-7947c3d3b5cb · outbound

This paper cites InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 10274–10284.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 10274–10284

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T18:59:42.949686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:59:42.949686Z digest=sha256:36cd92d3bd479a9bf6c4ac372df8567ac4ed064c6a7460ce96fd3e8eacdd7541

Observation 4a46be4b-f81e-4621-98ff-aa683d63ea23 · outbound

This paper cites Do Audio-Language Models Understand Linguistic Variations?.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Do Audio-Language Models Understand Linguistic Variations?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T18:59:42.949686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:59:42.949686Z digest=sha256:b32ac653b449716223252595405050ef87691d7a0a9926807dacc755466dc02e

Observation 8c92d7c9-e623-4437-89c9-a362addbd4a7 · outbound

This paper cites MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-07-12T18:59:42.949686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:59:42.949686Z digest=sha256:e127848f600705d78797c15541051b63bd6a8d8310be97b905a40b3628599e05

Pith citing papers

Observation 6bdb6fd7-8c7f-4bb0-a7d5-6875b20e7d87 · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.452149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.452149Z digest=sha256:a2b593dea241f388fa3b143750df69355b0a6172aae7da6881d94e4bc2062114