Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.00800.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:33.023609Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:01:36.150100Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c26a34a6-401c-444b-b56b-9d4426b7316c · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22ad47f7-076f-4c97-94c5-2ddeb94a5f13 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b899848-a6e9-43d5-a3ab-5cb4e692de14 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9178af3-636b-4d95-b1ab-9fbb316c1c21 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24]
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9cbd217-355f-43af-8165-8c4a49217fbf · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer semantic-rich and discrete
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5b382c7-cdf9-45c2-ae30-53407d95080f · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b8053d3-a6a5-4e22-bdb8-2f33fd944383 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning with recurrent neural networks,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21aaba24-c546-4ab3-a3fd-5ecc41bf560d · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning: An overview of recent progress and new challenges,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 327d26fd-73b7-45c0-9581-f4d75e449ba1 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c32a6fbe-9650-4e07-b6a3-7baa2371cee5 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6605451-fc0a-494c-abe4-d485ddffa251 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio cap- tioning by fine-tuning bart with audioset tags,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 405f177c-82a0-4e4f-b9b0-fbffb039535e · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b04eacd2-2fe4-4413-a3e5-caf3b85c35fb · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd3a5ab3-c46d-4b97-8b81-c063f3e49c89 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0199daa7-7f7b-48a3-8f11-3e7babeef12a · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Recap: Retrieval-augmented audio captioning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95fd2992-8ce8-463c-ad70-4910f955bce7 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enhancing automated audio captioning via large language models with optimized audio encoding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5fb0d28-9ed9-4735-b5bf-fa0b242bd93f · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BEATs: Audio pre-training with acoustic tokenizers,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c71dbdf-29d1-4713-b39f-5cc7e576fecd · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Set: An ontology and human-labeled dataset for audio events,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e361f16-f008-4518-b797-c9ae4fd2cf9d · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 849d4c47-fa88-4064-a327-d07ebe4e913d · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer High fidelity neural audio compression,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39e2a501-5016-464b-ac77-b9882674171b · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7244c56a-c1e5-4a44-bd02-48dcc4a1123e · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1daff668-6f25-4c57-859f-d70c503c2e27 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 825707c1-2a9f-4374-806e-0fcdaa537847 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dde80387-82ca-484d-9e41-15118b8e9377 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SoundStream: An end-to-end neural audio codec,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cad56588-70cf-4805-afe2-cb9b317a6268 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer vq- wav2vec: Self-supervised learning of discrete speech representa- tions,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dacc8b28-076a-42f1-999d-7607c3e29081 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer How should we extract discrete audio tokens from self-supervised models?,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a310985e-a820-492c-96f7-f08ca33af77f · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audiocaps: Generating captions for audios in the wild,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b35c6c1-310e-4ee6-8a4e-c1c384c400d2 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Clotho: An audio cap- tioning dataset,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78e4e4d8-f205-4695-a62a-f2b7e767aa50 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32cd370-503b-4519-ad3a-18545a2f37e6 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Meteor universal: Language spe- cific translation evaluation for any target language,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61ddb5b5-50de-4892-bba4-d68103ef3633 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CIDEr: Consensus-based image description evaluation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89e7d13c-cac2-43b4-b452-0bba65e4a275 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SPICE: Semantic propositional image caption evaluation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a78f4136-0a21-428d-a152-04d6e1cace00 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improved image captioning via policy gradient optimization of spider,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffd1cc54-6e51-4411-8c84-438a46ef7573 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Can audio captions be evaluated with image caption metrics?,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e90ee26-cb6f-4015-84ee-ea6ba9f10b91 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a1131d4-b1e1-4a29-91fa-1082570c3239 · outbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · inbound
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.