Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:05:24.550981Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2506.17815.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:05:24.550981Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:05:24.211908Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T19:05:24.861716Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 945757d7-709f-4ca5-b503-788b01396021 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining without negative samples for Music Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa2f6545-a71d-4258-a1f0-5baa9419fb60 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be0f5391-5e66-4fab-aefd-d19ba2d0c3f7 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 58de1b19-7975-4724-b952-a98bf556c73f · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f7bb58df-5525-4316-a783-368fb19742db · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9744a747-551e-47a5-b147-e0f2095c5661 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 91ca2359-368c-4540-83f8-22e7ad6a372d · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding A blaring metal track with stompy kicks and distorted chuggy guitar
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation afabab7d-3803-4981-86a4-33abc1d99628 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding We train SLAP on an internal private dataset of 260,000 pairs of full-length production-quality music tracks and professionally annotated captions (PrivateCaps [16])
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c7b920fd-2159-4528-99d7-b2dd06023bfe · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding {}”, “{} music
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f5616719-b1f2-438e-81ec-35640d08c600 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP out- performs contrastive models on tasks including text-music retrieval, downstream probing, and zero-shot music un- derstanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 631c1df4-7a69-480e-beb3-f62240b755c0 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccd665f-024d-473b-850e-032f8e69229c · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Learn- ing transferable visual models from natural language supervision,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 89da1510-9bae-444b-adef-8960d29bc426 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Clap learning audio concepts from natural language supervi- sion,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df08d519-5a8c-42bc-8102-31bbfe6aaba2 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Collap: Contrastive long-form language-audio pretraining with musical temporal structure augmentation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 15106905-80fb-48d3-92f3-a055809b817a · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Aligned contrastive learning for text-to-music retrieval,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4374bfe1-e3dd-42a5-b1d0-ccf7c7d7bea6 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding T-clap: Temporal- enhanced contrastive language-audio pretraining,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 07d97a30-41b9-431c-ad79-5f35c95334ac · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Augment, drop & swap: Improving diversity in llm captions for efficient music-text representation learning,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 97a464e5-988b-4a00-8767-b313357e443e · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9495c76-aa8b-4b13-812f-6e449c6f50f7 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ddc9f865-b153-47df-95c5-900ab4587148 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Combined scaling for zero-shot transfer learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c029e01-3fd9-4173-8dd9-ece654f14937 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Bootstrap your own latent-a new approach to self-supervised learn- ing,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a18bf2c1-f291-4d79-b334-ca4e7d6a5c09 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Byol for au- dio: Self-supervised learning for general-purpose au- dio representation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d61d5dd1-1a3c-4cb9-9b38-08ae5fe39dcc · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding A simple framework for contrastive learning of visual represen- tations,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0cf88e36-9dbf-480f-9ff6-c39149484708 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Contrastive learning of musical representations,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6826ed26-a5ff-489b-b229-30370edcb540 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 625a80eb-1f82-48d2-af54-b99714e76f00 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Look, Listen and Learn
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 380390a5-1668-4e61-b3f6-575d85ed12e8 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Contrastive audio-language learning for music,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fa512335-9c60-49af-a599-8568ed8baf10 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Mulan: A joint em- bedding of music audio and natural language,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6225267b-8f9e-4fe1-9f72-6b41394be120 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Cacophony: An improved contrastive audio-text model,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e6a34537-ed9d-4269-ae02-aa94b4261791 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audioclip: Ex- tending clip to image, text and audio,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1874300-58e1-453f-b2b9-49a303d9e593 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Imagebind: One embedding space to bind them all,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 87c41252-fa3b-46f8-b91f-ec1ec878895b · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Gramian Multimodal Representation Learning and Alignment
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138d1298-3f3a-4fe5-9822-e27a6f0a05e7 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c22b2a8-4bc1-4dba-96fe-da2d5c35d0ab · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Flamingo: a visual language model for few-shot learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b388096b-0255-4771-9f4c-dd02bafebac8 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16537c1b-d7fe-400a-b61c-465c069346fc · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Reclap: Improving zero shot audio classification by describ- ing sounds,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c4ee1d7-d1c7-44e4-b32b-15379c0562f3 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Drcap: Decoding clap la- tents with retrieval-augmented generation for zero-shot audio captioning,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb449c45-2402-4a55-a59d-0f9c2f44720a · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Recap: Retrieval-augmented audio captioning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 010c2f34-a4ad-41e4-8701-939b67aeebe1 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Fast timing- conditioned latent audio diffusion,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75347fad-e324-4f3d-993d-ebbaa0b9a3d1 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Long-form mu- sic generation with latent diffusion,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a8f40e6d-2d35-4ddf-8d00-9defa906477b · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Diff-a-riff: Musical accompaniment co-creation via latent diffu- sion models,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa88b59e-227c-4205-85dc-45b104eb117e · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding AudioLDM: Text-to- audio generation with latent diffusion models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78ad5ae0-3bc5-4205-8067-3413369db525 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding MusicLM: Generating Music From Text
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4664656b-87b3-4ace-9538-c84d79341393 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Sigmoid Loss for Language Image Pre-Training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b252c074-0b58-46eb-90aa-f515f7c84626 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding The Hidden Uniform Cluster Prior in Self-Supervised Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c38c8a-af32-41a4-a645-1762d26f27d6 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Understand- ing the Modality Gap in CLIP,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cdfd19ea-2375-4fe8-a822-2311b576cb34 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Exploring simple siamese repre- sentation learning,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5ea40e81-1b05-4726-bba1-475d6c8374d9 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Self-Supervised Learning from Images with a Joint-Embedding Predic- tive Architecture,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e47592c-ceed-4a5a-bc6a-2691f300d9e9 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Understanding self-supervised Learning Dynamics without Contrastive Pairs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25373c4b-0f58-44cf-a6f8-02cbdff7392a · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding DINOv2: Learning Robust Visual Features without Supervision
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281b6adb-5128-4ffd-b925-6983e8fd76cd · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 13929d01-3aec-4eb9-bf95-1f0cd71ecc35 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding ATST: Audio Representa- tion Learning with Teacher-Student Transformer,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c5fcfb1-286f-46e9-b62a-946cf7dacdd8 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Masked Modeling Duo: Learning Representations by Encour- aging Both Networks to Model the Input,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 946ec3e9-912e-472f-b2fb-45c1033238ef · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Masked latent prediction and classification for self-supervised audio representation learning,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db21b2ac-c6c9-4217-b219-16f089a894ca · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding The song describer dataset: a corpus of audio captions for music-and- language evaluation,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb626762-365e-4cee-bb7b-250fd95ebdbe · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Musical genre classification of audio signals,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734ae11d-8940-409a-829e-942bf798e60d · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Evaluation of Algorithms Using Games: The Case of Music Tagging,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 21779266-27ef-444d-b195-564795cad794 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Openmic- 2018: An open data-set for multiple instrument recog- nition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a8d1bfb-64a2-483c-88ca-6afd3997d80b · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Au- dio set: An ontology and human-labeled dataset for au- dio events,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6dfeda40-79a7-494c-9694-38b51eee4358 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Large-scale con- trastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 96f531be-00da-4d83-9902-c66d243b1e7d · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Hts-at: A hierarchical token-semantic audio transformer for sound classifica- tion and detection,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d30969b4-bc60-4a9c-aac8-adaaff1a3403 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1127da-c30f-40a1-b275-7cea13d8890c · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Specaugment: A simple data augmentation method for automatic speech recognition,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9e64b67c-6dea-40c0-b0c1-56c26a9e8c97 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SampleMatch: Drum Sample Retrieval by Musical Context
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ebfbe100-1941-40cf-af50-81acce11737f · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Ef- ficient training of audio transformers with patchout,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 152b4d69-45a9-4465-99c3-c2914abb1c21 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Codified au- dio language modeling learns useful representations for music information retrieval,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e3fd9e40-625b-4cff-81c8-f805c4b88283 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Beats: audio pre- training with acoustic tokenizers,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8dd8a42b-a280-4777-8b9c-6a106a4dc75e · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Improving mu- sical accompaniment co-creation via diffusion trans- formers,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae633985-3e54-47d2-aaaa-ddaf82aaf8f8 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding On the Language Encoder of Contrastive Cross-modal Models,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9daf461b-eab3-46d9-8e7c-7a68f1a3d7a1 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dbd537-398e-4096-9456-24be9624bae7 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb24102-8f43-4882-b82f-cc24cc63e492 · outbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Available: http://arxiv.org/abs/2310
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f7bb58df-5525-4316-a783-368fb19742db · inbound
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.