Pith. sign in

Paper Citation Record · LEDGER

Natural Language Supervision for General-Purpose Audio Representations

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2309.05767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05767 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.280500Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:06.204617Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b97ce92b-dd08-458a-a942-8ba4a2e442af · inbound

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models cites this paper.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Natural Language Supervision for General-Purpose Audio Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.280500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.280500Z digest=sha256:f9495e91d5edfd303da154ce9da426960c630a52f26e054c3f30fb93e845ee4f

Observation 56e5fd14-df51-4958-83a5-9fbf72c1ac10 · inbound

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval cites this paper.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Natural Language Supervision for General-Purpose Audio Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:16.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:16.501373Z digest=sha256:bd9fa7a32477742d1558ac98cbd23b898a6bda697583076cc92effa458a83445

Observation 8c7295bf-8c98-4fc3-92bc-ed14beb24e5f · inbound

Assessing the Alignment of Audio Representations with Timbre Similarity Ratings cites this paper.

Assessing the Alignment of Audio Representations with Timbre Similarity Ratings Natural Language Supervision for General-Purpose Audio Representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:58.841041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:58.841041Z digest=sha256:e7097131335caea67a52396471e5397209835eecdd7138ea58e09fc9aada32e2

Observation d0df0562-3026-4658-90ea-b6d7ec5705da · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Natural Language Supervision for General-Purpose Audio Representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.732238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.732238Z digest=sha256:aef85bbfc01c4c6366007c480417c1b9b00b146b90a52b9b55e1e7b78d028c88

Observation f6cfd9a2-ca3e-476d-a1b4-21d2210b9d49 · inbound

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models cites this paper.

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models Natural Language Supervision for General-Purpose Audio Representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:37.651853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:37.651853Z digest=sha256:abb0561e323d2a87efdab35753d871693d26563840de352b875aa7f4d943df8d

Observation ac7a9a48-9cd1-43a8-ae44-9dccf9fa0056 · inbound

Reasoning-Aware Multimodal Fusion for Hateful Video Detection cites this paper.

Reasoning-Aware Multimodal Fusion for Hateful Video Detection Natural Language Supervision for General-Purpose Audio Representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T19:00:14.134519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:00:14.134519Z digest=sha256:a644fc0b649e3b93e3920218cf2943ad9e8a1feaf6b1dbae220a0aefe244c44d

Observation e94647d9-cd93-430c-936c-48719fd21b48 · inbound

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models cites this paper.

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models Natural Language Supervision for General-Purpose Audio Representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:40:30.841192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:39:46.434831Z digest=sha256:69306d296f16a4c4ed0051dc3f01a211863254a7158efda03c2d080f48b3af7a

Observation a1fde89a-d60c-4ce4-a1db-b92907dcca09 · inbound

FIGMA: Towards FIne-Grained Music retrievAl cites this paper.

FIGMA: Towards FIne-Grained Music retrievAl Natural Language Supervision for General-Purpose Audio Representations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.206175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T23:32:09.401023Z digest=sha256:d831a1148077962de375f4aa49dfdff6033ee721004077d1c696057db44dd3fa