Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:13.173892Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2511.16757.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:13.173892Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:28:49.998236Z
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ce739042-061b-4cb0-8cc6-c22fbdc75922 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation YouTube-8M: A Large-Scale Video Classification Benchmark
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b5de18-f890-4206-93bc-4bd68033e922 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Common Voice: A Massively-Multilingual Speech Corpus
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dca371f-5f41-4521-8e81-5b79e200dd2f · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757193c3-f9c7-4b35-b7af-7d9caf4bc02b · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Qwen2-Audio Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f57623f-60c2-4210-ab8a-0d3b36eb8654 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Midashenglm: Efficient audio understanding with general audio captions.arXiv preprint arXiv:2508.03983,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5f3008-b2be-4e24-ae04-917548a25f26 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Scaling rich style-prompted text- to-speech datasets.arXiv preprint arXiv:2503.04713,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73e955d-b0cc-40f7-b1be-502aa5bbf469 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Clotho: An audio captioning dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d4e6e2c-d6cf-4df0-b13f-13feba227ed4 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Threshold independent evaluation of sound event detection scores
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a23a6c-3396-45dd-ae43-fa5e3796ae36 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Clap learning audio concepts from natural language supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0f3a52-89c0-4f58-866f-3797e15bbf98 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b5af39-f331-434e-8f89-d572b3de60d2 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Ast: Audio spectrogram transformer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1f7c7a-b55a-4d1e-b348-b32c898633d5 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Contrastive Audio-Visual Masked Autoencoder
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3fc791d-3b12-4f09-a897-e58df262fbf4 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation The benefit of temporally-strong labels in audio event classification
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe928ae7-8709-4edf-a97e-74d858049082 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Audiocaps: Generating captions for audios in the wild
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c57a7e6f-9813-476f-a697-ffe5175a3ab7 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Bootstrap- ping language-audio pre-training for music captioning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb5d8aa-187c-4b0c-b7f8-9a9eb6fec207 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation A Diversity-Promoting Objective Function for Neural Conversation Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ce15a0-f12f-43a9-9ab7-1316e2bdc2ba · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Music understanding llama: Advancing text-to-music generation with question answering and captioning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd121fca-7daf-49e3-b6ae-7a1c0ac820ab · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c171446b-c51c-4493-9a76-50842f872cfe · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615dafb6-a46e-4309-9fdc-cd1f6b58f6fd · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Expresso: A benchmark and analysis of discrete expressive speech resynthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed604b87-016f-4410-a5d2-8b5f8682cee0 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96523a0-8922-43e6-9e09-2bc2a24d2d48 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation M2d2: Exploring general-purpose audio-language representations beyond clap
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7017ae55-494d-4018-bbba-7d7c1f050101 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Representation Learning with Contrastive Predictive Coding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc2b25d-b38f-466c-9474-26e972a603c6 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e47b155-47ef-4a4c-9af1-7c81fcf94e81 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40cd07d-aa01-47c4-b07d-6516381f5823 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Ears: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537edd05-764f-4232-9357-53816e6f6c68 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0168ded6-1ca9-410c-b702-a6f097fbf190 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9652705-8b00-4560-b9a1-ad4bf3f5b834 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Hear: Holistic evaluation of audio representations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb0f9fa-5e98-466b-86ef-b27020d59182 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Towards learning universal audio representations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f39bf62-dda9-404d-80e7-c689903a62ab · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Wav2clip: Learning robust audio representations from clip
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a645b7a2-ed46-446e-a966-78c3b632861d · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5290700a-c99c-4190-b100-73ddc28758a2 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Qwen2.5 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d11406-f61d-48f6-8d96-a004bf39d094 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Speech self- supervised representation benchmarking: Are we doing it right? InInterspeech 2023, pp
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f80b2d-5844-41fe-98f7-cc8dcdbdee82 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f31a174-8633-4e17-9659-dab533bdfcf6 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e2a49a-a2bb-48fa-8a53-f1bb106aefca · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d305ed88-57b4-4a52-ac45-b4292756d9d8 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a311440c-b7a8-416b-9802-b704963ea926 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ead360e-496e-41d7-aacb-c35a78841280 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation The architecture employs a U-Net-inspired design with six Transformer stages that process sequences at multiple temporal resolutions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa230fa-7631-42be-b74e-14a8343967dc · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0edf5fba-f34d-41b8-a3ae-311d3e6b0463 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation ATST: Audio Representation Learning with Teacher-Student Transformer
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff260b9-0970-433a-a3d5-522553d1d05d · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation MusicLM: Generating Music From Text
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf63c450-d9dc-4046-9236-8c7c6b994720 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03dc1531-4719-4696-9910-0733bb3d7b61 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Look, listen, and learn more: Design choices for deep audio embeddings
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bdd87a5-184b-4737-94c4-aecfa217fecd · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Scaling up masked audio encoder learning for general audio classification
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae167f7-1436-4d5b-ad13-6f556a13f876 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99ea17d-78c8-4096-b2c1-476e33bc9d56 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Vggsound: A large-scale audio- visual dataset
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad4edd1-5800-4821-a31b-d0a2090f0f3b · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36034fc8-1f4e-4a03-8780-8f0aa5dbafb4 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Efficient supervised training of audio transformers for music representation learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcab7841-040c-4567-820b-9a2a5c149269 · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation VoxCeleb2: Deep Speaker Recognition
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d977bf-7ef6-491a-bfea-26e008035d6d · outbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b7fc15-ce07-44c9-9380-e6ffe9713782 · inbound
SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.