Pith. sign in

Paper Citation Record · LEDGER

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2303.17395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.17395 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:36.312770Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e00eb565-7121-43ac-9e4b-f5ca2226e330 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.641261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:95cb585345ca422deef7808c1bb842096e8c69bd19d5fb5bdfaeae23a778140c

Observation 4d86672d-158f-4db5-897e-765259e9911e · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.229271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:6b561df1acbf71d46f5944d8213fc1003930d5a66dc0f7f6a143815d7fe2d706

Observation 06a0a052-5378-44a7-8f70-e2a218a0ab74 · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:29:46.317296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:a92414e7c703162dd386d7f251ba89f0c6d00c2b37c69e3795e22091d6e5d557

Observation 6a592bd0-e2d4-47da-a507-a6e6606fd5df · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.586093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:dd8b15cba1fee7c44f1939f228260253df8cb772c912080a2f3bf9c3b8bbb393

Observation f421ccc4-8eac-4592-a2c8-eb53ca8900b1 · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T12:42:08.965410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:feff87b5af40624f2e2107c1d8a854cf6d4f16a4f621a776b89a68356ef60c47

Observation 9d8406a0-425f-411c-8591-7d69b41a268f · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 287

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.644127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:cc6a43137adc8781ebe9988adf7202f9706af745883459f25e35e110a8c41758

Observation fbf7f9ca-b8b5-41f9-bfeb-b9bb79684f0e · inbound

Quantum Computing : A New Frontier for Science and Society cites this paper.

Quantum Computing : A New Frontier for Science and Society WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 191

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T17:06:21.742250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-09T17:01:49.230563Z digest=sha256:17a5a1abed4ace3f36949e507c445bc00a79ee3add6e1817c900e7c6e96fd896

Observation 97cd49ab-0e2e-4030-b8c0-f60474ab909f · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.312770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.312770Z digest=sha256:d50a51be90dd1eb2912e48c3d930fd5e5e1368922fc49e0f2b7ade376b2f157f