Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:41.089739Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.14989.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:41.089739Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:37.637530Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:29:41.315494Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · outbound
Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33ad1828-4fd5-4d8e-bb50-0923888c62ad · outbound
Discrete Audio Representations for Automated Audio Captioning We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cefd1b20-81c4-4ea7-a9c9-c6e2de374566 · outbound
Discrete Audio Representations for Automated Audio Captioning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72d3ae5e-6dc2-4d14-8e49-086a70a44d25 · outbound
Discrete Audio Representations for Automated Audio Captioning Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 603fc8e4-6430-4c5f-b5f4-a5dffa6819bf · outbound
Discrete Audio Representations for Automated Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8825f6b-0b1c-4c98-b2d7-44a980073a55 · outbound
Discrete Audio Representations for Automated Audio Captioning Audio captioning based on transformer and pre-trained cnn,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59b5a379-4928-4a17-ad1d-929c748ca906 · outbound
Discrete Audio Representations for Automated Audio Captioning Automated audio caption- ing by fine-tuning bart with audioset tags,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8b710d6-2282-4147-9d68-64b5259cbd75 · outbound
Discrete Audio Representations for Automated Audio Captioning Leveraging pre-trained bert for audio captioning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83fc5ec3-ab4e-4553-bf30-3d2aac66599b · outbound
Discrete Audio Representations for Automated Audio Captioning Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6243da2-8ba2-4ee1-ac41-5b5c5bb07e3f · outbound
Discrete Audio Representations for Automated Audio Captioning Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35d13838-760a-4507-a964-a3d39b0c42e4 · outbound
Discrete Audio Representations for Automated Audio Captioning Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3ef9d7-fdaa-49fb-a8eb-c94b7d2dcecf · outbound
Discrete Audio Representations for Automated Audio Captioning Adapting a convnext model to audio classification on au- dioset,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 46b92cf6-81f9-48cc-933a-4c058438af93 · outbound
Discrete Audio Representations for Automated Audio Captioning Beats: audio pre-training with acoustic tok- enizers,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b9661e5-d984-425b-8ea9-a1c8e8aff9c9 · outbound
Discrete Audio Representations for Automated Audio Captioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97881d5-1890-40a5-8d5a-44f7c8b214a5 · outbound
Discrete Audio Representations for Automated Audio Captioning BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faeb6c81-bccb-4221-a163-554f51667a31 · outbound
Discrete Audio Representations for Automated Audio Captioning Language models are unsupervised multitask learners,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c960552-04af-4b3a-ba53-1ebb82abe3a5 · outbound
Discrete Audio Representations for Automated Audio Captioning STAB: Speech Tokenizer Assessment Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf8414e-bbe7-412c-926d-3c551e889850 · outbound
Discrete Audio Representations for Automated Audio Captioning Soundstream: An end-to-end neural audio codec,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604426d6-34a9-45e0-8b98-d9a7518c4e3a · outbound
Discrete Audio Representations for Automated Audio Captioning High Fidelity Neural Audio Compression
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a635e2e1-ee8c-4774-a573-a3dd8fd5abe8 · outbound
Discrete Audio Representations for Automated Audio Captioning High-fidelity audio compression with improved rvqgan,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db345d99-1cc7-44c3-b2e9-cff7d83fdb29 · outbound
Discrete Audio Representations for Automated Audio Captioning Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d440acd7-51cc-45cf-99e7-d54ad8af7e6c · outbound
Discrete Audio Representations for Automated Audio Captioning Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820f0797-cf0f-4bd5-aad6-9be4af9109d9 · outbound
Discrete Audio Representations for Automated Audio Captioning RepCodec: A speech represen- tation codec for speech tokenization,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · outbound
Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1079c0-39a2-4f82-bd04-ba21a7d587b4 · outbound
Discrete Audio Representations for Automated Audio Captioning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6434d90-b764-4b0a-b87d-a5606774bcdf · outbound
Discrete Audio Representations for Automated Audio Captioning Exploration of efficient end-to-end asr using discretized input from self-supervised learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9031b10f-1abd-4b2c-9429-6d768514a162 · outbound
Discrete Audio Representations for Automated Audio Captioning How should we extract discrete audio tokens from self-supervised models?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f071cc86-68db-4404-ba0a-947998171708 · outbound
Discrete Audio Representations for Automated Audio Captioning Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e716f219-6b97-4506-b963-09661381826f · outbound
Discrete Audio Representations for Automated Audio Captioning Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 559d8f9d-1118-412f-8dc3-f2ee8e0538a6 · outbound
Discrete Audio Representations for Automated Audio Captioning Neural discrete represen- tation learning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca75e92-a442-4ac8-9b6b-6902dbdd3bb9 · outbound
Discrete Audio Representations for Automated Audio Captioning Investigating lo- cal and global information for automated audio captioning with transfer learning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a11eb655-7160-47d5-9313-fe082c58d128 · outbound
Discrete Audio Representations for Automated Audio Captioning Prefix tuning for auto- mated audio captioning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa09ffca-3051-4b66-b7ac-c9a49cbceec3 · outbound
Discrete Audio Representations for Automated Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7a4798-49e0-4a63-8e6a-5c18735ef4c8 · outbound
Discrete Audio Representations for Automated Audio Captioning Clotho: An audio cap- tioning dataset,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca949e47-509f-4b35-9b34-a44501a28463 · outbound
Discrete Audio Representations for Automated Audio Captioning Improved image captioning via policy gradient optimization of spider,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e43b8e5-c476-4d03-a339-29dcf1715581 · outbound
Discrete Audio Representations for Automated Audio Captioning Can audio captions be evaluated with image caption metrics?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · inbound
Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.