Pith. sign in

Paper Citation Record · LEDGER

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

As of 22 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2506.17815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17815 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:05:24.550981Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:05:24.211908Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T19:05:24.861716Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact3
  • verified fuzzy47
  • unresolved19
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 945757d7-709f-4ca5-b503-788b01396021 · outbound

This paper cites SLAP: Siamese Language-Audio Pretraining without negative samples for Music Understanding.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining without negative samples for Music Understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.431204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.196688Z digest=sha256:5961c41bb86f70c94402c74ec8a7fba17b9680cdb09358d05b75544ee906f13a

Observation aa2f6545-a71d-4258-a1f0-5baa9419fb60 · outbound

This paper cites an unresolved cited work.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:05:25.421612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.201207Z digest=sha256:7f5a4b77438fa187341a3135c14025f2ef376e0869762f47dd048af2490acf6f

Observation be0f5391-5e66-4fab-aefd-d19ba2d0c3f7 · outbound

This paper cites an unresolved cited work.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:05:25.412006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.204732Z digest=sha256:995a8afa24c530f1e11d15f0eeee6352570b26b715b3622b91ba8bbca63b1471

Observation 58de1b19-7975-4724-b952-a98bf556c73f · outbound

This paper cites an unresolved cited work.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:05:25.402224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.208515Z digest=sha256:dcd412768fb2cf019e59ebc05ad947e7e3ad7d01fea0ae33e996e395dda8e863

Observation f7bb58df-5525-4316-a783-368fb19742db · outbound

This paper cites SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:05:24.865750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.211908Z digest=sha256:eb30f0d4a39f5fa32f41b1a96b7047fbc14516d1cbd2423ec7d0704e09cea3de

Observation 9744a747-551e-47a5-b147-e0f2095c5661 · outbound

This paper cites an unresolved cited work.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:05:25.391813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.215913Z digest=sha256:4697c7edaf8a405615a102a68b5ef86bb8a9e54882f8b284dcb61d1dc53b185f

Observation 91ca2359-368c-4540-83f8-22e7ad6a372d · outbound

This paper cites A blaring metal track with stompy kicks and distorted chuggy guitar.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding A blaring metal track with stompy kicks and distorted chuggy guitar

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.381645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.219886Z digest=sha256:c5fa7f99d8da93476fd94301c5d155e7ed8fc57662d9587e162e570c6c7bbf7f

Observation afabab7d-3803-4981-86a4-33abc1d99628 · outbound

This paper cites We train SLAP on an internal private dataset of 260,000 pairs of full-length production-quality music tracks and professionally annotated captions (PrivateCaps [16]).

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding We train SLAP on an internal private dataset of 260,000 pairs of full-length production-quality music tracks and professionally annotated captions (PrivateCaps [16])

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:05:25.370907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.223759Z digest=sha256:85798a243a8b36b535e66460993890a90c1c7169946a39e68e638b3f91870679

Observation c7b920fd-2159-4528-99d7-b2dd06023bfe · outbound

This paper cites {}”, “{} music.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding {}”, “{} music

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:05:25.360384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.227475Z digest=sha256:eccd78965439f60d70514e5409f331a4665f1f966454ef31b8926b9397857fd4

Observation f5616719-b1f2-438e-81ec-35640d08c600 · outbound

This paper cites SLAP out- performs contrastive models on tasks including text-music retrieval, downstream probing, and zero-shot music un- derstanding.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP out- performs contrastive models on tasks including text-music retrieval, downstream probing, and zero-shot music un- derstanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.349552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.231401Z digest=sha256:a740da7e0cc62cd132aba8b046e810406296ddca6396c6ec7b501ddceb447730

Observation 631c1df4-7a69-480e-beb3-f62240b755c0 · outbound

This paper cites an unresolved cited work.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.235099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.235099Z digest=sha256:cb1753788567f664458aa293f2832c5fd7b2eab85506c7f2507997824ba94e21

Observation 9ccd665f-024d-473b-850e-032f8e69229c · outbound

This paper cites Learn- ing transferable visual models from natural language supervision,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Learn- ing transferable visual models from natural language supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.331702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.238742Z digest=sha256:537980e71c07f2a9e2f9776aa65c9f3b08dcfc4ebbc5528e279638687a717dc7

Observation 89da1510-9bae-444b-adef-8960d29bc426 · outbound

This paper cites Clap learning audio concepts from natural language supervi- sion,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Clap learning audio concepts from natural language supervi- sion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.321236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.242477Z digest=sha256:1461e916e86e10043890bb78b19de87f6a3638ce347258279c478247fa43bc57

Observation df08d519-5a8c-42bc-8102-31bbfe6aaba2 · outbound

This paper cites Collap: Contrastive long-form language-audio pretraining with musical temporal structure augmentation,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Collap: Contrastive long-form language-audio pretraining with musical temporal structure augmentation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.310759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.245965Z digest=sha256:35348c4b1c8a82f91a8386cb04bb7b0b4782bbc077cf0305027ceb220abe416e

Observation 15106905-80fb-48d3-92f3-a055809b817a · outbound

This paper cites Aligned contrastive learning for text-to-music retrieval,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Aligned contrastive learning for text-to-music retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.300145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.249648Z digest=sha256:cacefbcd5945e9549dabedf766bd248909097f26a6d540a1763513904a2968a7

Observation 4374bfe1-e3dd-42a5-b1d0-ccf7c7d7bea6 · outbound

This paper cites T-clap: Temporal- enhanced contrastive language-audio pretraining,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding T-clap: Temporal- enhanced contrastive language-audio pretraining,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.289096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.253342Z digest=sha256:e489edcaa590868652e954694e8be2e07db71f158ea90ab642080566f6b4935b

Observation 07d97a30-41b9-431c-ad79-5f35c95334ac · outbound

This paper cites Augment, drop & swap: Improving diversity in llm captions for efficient music-text representation learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Augment, drop & swap: Improving diversity in llm captions for efficient music-text representation learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.278979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.257131Z digest=sha256:3886a0679ea8e1fec1ec093a53daf83228085cf7253ab4da36b9e4ea6bdd8ce5

Observation 97a464e5-988b-4a00-8767-b313357e443e · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.260513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.260513Z digest=sha256:a627bd7dbc41d83b7da200109ea3fe8a648c6bec012db93f7301cd83faff564c

Observation e9495c76-aa8b-4b13-812f-6e449c6f50f7 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.268934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.264269Z digest=sha256:766199d4f43e05607376f1cf394faeb083ea051a24ba1fa32ba0281c5a3f1fc8

Observation ddc9f865-b153-47df-95c5-900ab4587148 · outbound

This paper cites Combined scaling for zero-shot transfer learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Combined scaling for zero-shot transfer learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.258892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.267955Z digest=sha256:a9967378087c8933934cb10577471d3035eee3f27da7f67e54e1ac6f9883ad32

Observation 0c029e01-3fd9-4173-8dd9-ece654f14937 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learn- ing,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Bootstrap your own latent-a new approach to self-supervised learn- ing,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.248293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.271272Z digest=sha256:76ea9de9186510b5b7e05674f8337859df8b9c047d246d28b084ccd5271e3169

Observation a18bf2c1-f291-4d79-b334-ca4e7d6a5c09 · outbound

This paper cites Byol for au- dio: Self-supervised learning for general-purpose au- dio representation,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Byol for au- dio: Self-supervised learning for general-purpose au- dio representation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.237525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.274618Z digest=sha256:0bded8984aa1103f165570b036c3bf8d5111501c34fbcbcbd5bf5ca555784ed4

Observation d61d5dd1-1a3c-4cb9-9b38-08ae5fe39dcc · outbound

This paper cites A simple framework for contrastive learning of visual represen- tations,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding A simple framework for contrastive learning of visual represen- tations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.227417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.277965Z digest=sha256:d05694b0f7f4f0a55e9325cd70a571d96aebaa1bbd82df1e6be997787a189ff8

Observation 0cf88e36-9dbf-480f-9ff6-c39149484708 · outbound

This paper cites Contrastive learning of musical representations,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Contrastive learning of musical representations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.216904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.281374Z digest=sha256:25fecc6971bd00caf5cabe9ba4036aeb7b35899d62acf04f32475bca4e112af5

Observation 6826ed26-a5ff-489b-b229-30370edcb540 · outbound

This paper cites VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.284761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.284761Z digest=sha256:f1ed85877d48a727d9659615befbe1527de2ca27ce7d8017473c761d0b0ba60b

Observation 625a80eb-1f82-48d2-af54-b99714e76f00 · outbound

This paper cites Look, Listen and Learn.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Look, Listen and Learn

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:05:24.825876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.288371Z digest=sha256:bf2372a31f9d9808b0db2f946084c79cd94b19b5545fa2b259056f05ecef389d

Observation 380390a5-1668-4e61-b3f6-575d85ed12e8 · outbound

This paper cites Contrastive audio-language learning for music,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Contrastive audio-language learning for music,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.206326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.292101Z digest=sha256:4b3abf46a891c33187a42a6147a42e7379d8ca6227ea66c783e81f8d71f9633e

Observation fa512335-9c60-49af-a599-8568ed8baf10 · outbound

This paper cites Mulan: A joint em- bedding of music audio and natural language,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Mulan: A joint em- bedding of music audio and natural language,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.195528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.296065Z digest=sha256:7189bd94b1ff57a57025250ce67a90e09bad3e912d8646f9d6f0562709d564d9

Observation 6225267b-8f9e-4fe1-9f72-6b41394be120 · outbound

This paper cites Cacophony: An improved contrastive audio-text model,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Cacophony: An improved contrastive audio-text model,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.185141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.299437Z digest=sha256:06ba82be289eb5f1c663e1bb1168a49fc9ce7aa57b2066cfa828b39ad92013c3

Observation e6a34537-ed9d-4269-ae02-aa94b4261791 · outbound

This paper cites Audioclip: Ex- tending clip to image, text and audio,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audioclip: Ex- tending clip to image, text and audio,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.175300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.302611Z digest=sha256:73ce38880b059f78498d4125f8bd75c0c31cfa88ee199e844b8b9bd1becf3d4a

Observation e1874300-58e1-453f-b2b9-49a303d9e593 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Imagebind: One embedding space to bind them all,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.164987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.306012Z digest=sha256:4e13c01752c56b3e66bb2abc4eda482faba339034ecffea4873fd5bbfddaaabf

Observation 87c41252-fa3b-46f8-b91f-ec1ec878895b · outbound

This paper cites Gramian Multimodal Representation Learning and Alignment.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Gramian Multimodal Representation Learning and Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.309376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.309376Z digest=sha256:921e885e6464f014983cf28a20bd2ee3622df49af92086a072968fea7bcab683

Observation 138d1298-3f3a-4fe5-9822-e27a6f0a05e7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.312829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.312829Z digest=sha256:b6f9d446a7dcc50ab881109bf661ea9cf6e6c3597a29fcc2d9e8a42b2fd86407

Observation 4c22b2a8-4bc1-4dba-96fe-da2d5c35d0ab · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Flamingo: a visual language model for few-shot learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.154280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.316256Z digest=sha256:47fceaff02d74c4bacba0deaea984dc501120d7baa3d9acc2395fb716afac08c

Observation b388096b-0255-4771-9f4c-dd02bafebac8 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.319505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.319505Z digest=sha256:e9d45bdb7394d258dca8a79435400142422cf657de628da0d69772387e11c1be

Observation 16537c1b-d7fe-400a-b61c-465c069346fc · outbound

This paper cites Reclap: Improving zero shot audio classification by describ- ing sounds,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Reclap: Improving zero shot audio classification by describ- ing sounds,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.143420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.421463Z digest=sha256:643eb37976c04549cc9c510b012778fbbaf6f8b7916a96be1a52aac6052a9b47

Observation 7c4ee1d7-d1c7-44e4-b32b-15379c0562f3 · outbound

This paper cites Drcap: Decoding clap la- tents with retrieval-augmented generation for zero-shot audio captioning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Drcap: Decoding clap la- tents with retrieval-augmented generation for zero-shot audio captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.132958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.425203Z digest=sha256:c109598c0eb3991b1f0ca553293ebccfcf97c2848abe7ff63b1808a65322c3a6

Observation eb449c45-2402-4a55-a59d-0f9c2f44720a · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Recap: Retrieval-augmented audio captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.122683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.428547Z digest=sha256:8e650ec2c436bad8e7a0778e25509b8f4d20d1782602af6ded7c1201d735bd33

Observation 010c2f34-a4ad-41e4-8701-939b67aeebe1 · outbound

This paper cites Fast timing- conditioned latent audio diffusion,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Fast timing- conditioned latent audio diffusion,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.113015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.432313Z digest=sha256:f6baec76cfa539d6df661b12a789813527cabff725bb11944eb1509612317743

Observation 75347fad-e324-4f3d-993d-ebbaa0b9a3d1 · outbound

This paper cites Long-form mu- sic generation with latent diffusion,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Long-form mu- sic generation with latent diffusion,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.102835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.435618Z digest=sha256:66aa65def1329b397acd8ff4da12de3be4d5ed296c758d3f34b98d8f648ff5de

Observation a8f40e6d-2d35-4ddf-8d00-9defa906477b · outbound

This paper cites Diff-a-riff: Musical accompaniment co-creation via latent diffu- sion models,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Diff-a-riff: Musical accompaniment co-creation via latent diffu- sion models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.093007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.439126Z digest=sha256:d22bb41b5a552fa9fc7a19c71edec2ec7ccb687f9f28a2264ffc23507c2549d6

Observation aa88b59e-227c-4205-85dc-45b104eb117e · outbound

This paper cites AudioLDM: Text-to- audio generation with latent diffusion models,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding AudioLDM: Text-to- audio generation with latent diffusion models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.082850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.442809Z digest=sha256:b107385cbeeaf277893660702dbdd2911ea0bb420be304337f6696c4b430208b

Observation 78ad5ae0-3bc5-4205-8067-3413369db525 · outbound

This paper cites MusicLM: Generating Music From Text.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding MusicLM: Generating Music From Text

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.446062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.446062Z digest=sha256:0c6aef3975f3db23e7533dd437bbbf9e7bc90e3a8a400155e47b68d78e67aa89

Observation 4664656b-87b3-4ace-9538-c84d79341393 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Sigmoid Loss for Language Image Pre-Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.449881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.449881Z digest=sha256:f0edda565370bf8586c2553ffaf4b47205dad271686e5e667e5aa4727d6232b2

Observation b252c074-0b58-46eb-90aa-f515f7c84626 · outbound

This paper cites The Hidden Uniform Cluster Prior in Self-Supervised Learning.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding The Hidden Uniform Cluster Prior in Self-Supervised Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.453582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.453582Z digest=sha256:2047118c354bd84c1c5fd10c9b47d516808728bc819b9ad035c580b6afe1409a

Observation 77c38c8a-af32-41a4-a645-1762d26f27d6 · outbound

This paper cites Understand- ing the Modality Gap in CLIP,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Understand- ing the Modality Gap in CLIP,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.072739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.457481Z digest=sha256:0d1ba28d0c6b5bdc86814495303f8380ff848b2bc6cbcf4b9aa5cccde01d2706

Observation cdfd19ea-2375-4fe8-a822-2311b576cb34 · outbound

This paper cites Exploring simple siamese repre- sentation learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Exploring simple siamese repre- sentation learning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.062331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.460863Z digest=sha256:600afb3e7f4e4cc61c0a0561d52e101d73dc8f9fa800a820893e944e0ea8662a

Observation 5ea40e81-1b05-4726-bba1-475d6c8374d9 · outbound

This paper cites Self-Supervised Learning from Images with a Joint-Embedding Predic- tive Architecture,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Self-Supervised Learning from Images with a Joint-Embedding Predic- tive Architecture,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.052075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.464284Z digest=sha256:95a728dfd233e250f10a4567b58ae0f91dcfa627edc78aed7a16516f74d5e28d

Observation 5e47592c-ceed-4a5a-bc6a-2691f300d9e9 · outbound

This paper cites Understanding self-supervised Learning Dynamics without Contrastive Pairs.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Understanding self-supervised Learning Dynamics without Contrastive Pairs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.468014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.468014Z digest=sha256:c6a61e69d41322f539f6ef0a6f7a8a1675690ee1bd0b376c7c5c597441e89f4a

Observation 25373c4b-0f58-44cf-a6f8-02cbdff7392a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.471760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.471760Z digest=sha256:28b732498f1f465e9e59a8d228807ac67af22c4b9fd8c0606f36ab96a067e1b3

Observation 281b6adb-5128-4ffd-b925-6983e8fd76cd · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.041997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.475376Z digest=sha256:1a19610fac7e1a2b2686c04de30d233c32de5d7b253a597d4879900d7dcaa436

Observation 13929d01-3aec-4eb9-bf95-1f0cd71ecc35 · outbound

This paper cites ATST: Audio Representa- tion Learning with Teacher-Student Transformer,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding ATST: Audio Representa- tion Learning with Teacher-Student Transformer,

Reference 52

Resolution
verified exact
doi, observed 2026-08-15T19:05:24.583861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.479341Z digest=sha256:5d0bbeb6f32292964dda042164915b8b396ebe3509fcc261f510e659c870189c

Observation 4c5fcfb1-286f-46e9-b62a-946cf7dacdd8 · outbound

This paper cites Masked Modeling Duo: Learning Representations by Encour- aging Both Networks to Model the Input,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Masked Modeling Duo: Learning Representations by Encour- aging Both Networks to Model the Input,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.031411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.482915Z digest=sha256:c45281f0331ce6fc05408567b469b0ae763c54866f530010bc1517091b4155c2

Observation 946ec3e9-912e-472f-b2fb-45c1033238ef · outbound

This paper cites Masked latent prediction and classification for self-supervised audio representation learning,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Masked latent prediction and classification for self-supervised audio representation learning,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.020250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.486602Z digest=sha256:8c4df27c2bea2f4c35cb1e7a45bf2d4e8ee46efad1cf0d6fd1ee7382640e00e5

Observation db21b2ac-c6c9-4217-b219-16f089a894ca · outbound

This paper cites The song describer dataset: a corpus of audio captions for music-and- language evaluation,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding The song describer dataset: a corpus of audio captions for music-and- language evaluation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:25.008673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.490109Z digest=sha256:7da6fab53711c821b55c0590c992ed22a8c3effe4fdbd54ada102cc96c10a819

Observation bb626762-365e-4cee-bb7b-250fd95ebdbe · outbound

This paper cites Musical genre classification of audio signals,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Musical genre classification of audio signals,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.493800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.493800Z digest=sha256:cce24755c26f0f8a10f74e8700399da0643639706e64260f6e2ac49f7101e102

Observation 734ae11d-8940-409a-829e-942bf798e60d · outbound

This paper cites Evaluation of Algorithms Using Games: The Case of Music Tagging,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Evaluation of Algorithms Using Games: The Case of Music Tagging,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.996794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.497806Z digest=sha256:993a679e4f170d3fc8f455d1ff120ccee6f054133ec0ab0041ca8a9f8cbdcd2d

Observation 21779266-27ef-444d-b195-564795cad794 · outbound

This paper cites Openmic- 2018: An open data-set for multiple instrument recog- nition.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Openmic- 2018: An open data-set for multiple instrument recog- nition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.986528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.501689Z digest=sha256:1444226a6436686e70d53e3a44090bf9a60a4013d7613a712b78c5cd73e4fca6

Observation 5a8d1bfb-64a2-483c-88ca-6afd3997d80b · outbound

This paper cites Au- dio set: An ontology and human-labeled dataset for au- dio events,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Au- dio set: An ontology and human-labeled dataset for au- dio events,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.976548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.505425Z digest=sha256:21a6df54148b2d3a79e9b9d9e1be2d19f8e004cbae10fa28b9e4a0c9cbf9f3c3

Observation 6dfeda40-79a7-494c-9694-38b51eee4358 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Large-scale con- trastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.966603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.509144Z digest=sha256:bbecad641f836971265b290cefe48c55ae835adcdfbbd4b21f4bba4aa34ed965

Observation 96f531be-00da-4d83-9902-c66d243b1e7d · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classifica- tion and detection,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Hts-at: A hierarchical token-semantic audio transformer for sound classifica- tion and detection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.956689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.512532Z digest=sha256:66afafb0d1f3f014ed4e265a5aca786bb97dd9990572e8f7d4f89a3d5bdb0bb8

Observation d30969b4-bc60-4a9c-aac8-adaaff1a3403 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.515729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.515729Z digest=sha256:af0b6cb4b274f6265fcc4cb282786000abc3c881d335c74fc4708c34541d5083

Observation 8e1127da-c30f-40a1-b275-7cea13d8890c · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.944957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.519228Z digest=sha256:63014ad94e3361069ffecdfe209a5a9f2b99d693bffe2b42b8e8b16b2532369e

Observation 9e64b67c-6dea-40c0-b0c1-56c26a9e8c97 · outbound

This paper cites SampleMatch: Drum Sample Retrieval by Musical Context.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SampleMatch: Drum Sample Retrieval by Musical Context

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:05:24.627520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.522813Z digest=sha256:6fbaf58b91a850ac14305b8b0f4ab3f35c164fac46685ccb50849879e44abbe6

Observation ebfbe100-1941-40cf-af50-81acce11737f · outbound

This paper cites Ef- ficient training of audio transformers with patchout,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Ef- ficient training of audio transformers with patchout,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.932660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.526185Z digest=sha256:d5d566985b3f9eb23ffaae0def0c3eb7eb6c91e594ae21b94de785778563c1e4

Observation 152b4d69-45a9-4465-99c3-c2914abb1c21 · outbound

This paper cites Codified au- dio language modeling learns useful representations for music information retrieval,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Codified au- dio language modeling learns useful representations for music information retrieval,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.920316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.529632Z digest=sha256:c8718e4156b4280c445eb23730eac7e789328d1221115df2dbc34e1f72c175e5

Observation e3fd9e40-625b-4cff-81c8-f805c4b88283 · outbound

This paper cites Beats: audio pre- training with acoustic tokenizers,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Beats: audio pre- training with acoustic tokenizers,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.908283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.533197Z digest=sha256:3717e0f158d37deff67fff1d564b8a847cc87d8758eee1736cd1e598c9a064cd

Observation 8dd8a42b-a280-4777-8b9c-6a106a4dc75e · outbound

This paper cites Improving mu- sical accompaniment co-creation via diffusion trans- formers,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Improving mu- sical accompaniment co-creation via diffusion trans- formers,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.897654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.536755Z digest=sha256:3b46f25eed42f07a738fa9e5cb0f0fb1d23d84eb3bf2e692219b065ce80663a9

Observation ae633985-3e54-47d2-aaaa-ddaf82aaf8f8 · outbound

This paper cites On the Language Encoder of Contrastive Cross-modal Models,.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding On the Language Encoder of Contrastive Cross-modal Models,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.886541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.540240Z digest=sha256:cee9cbd0038d248a7ab4a5cce6c2e354b413d030561090602605f6f38c9f5460

Observation 9daf461b-eab3-46d9-8e7c-7a68f1a3d7a1 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.546999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.546999Z digest=sha256:1562a0e39c9d9692bced43ef313e4010a3d494264711b6599df7291937b0f560

Observation 53dbd537-398e-4096-9456-24be9624bae7 · outbound

This paper cites MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.550981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.550981Z digest=sha256:4db507bc329ac0ada2d5478aff3dbbaf72c5515ebdbacd1ba1a4882468d79509

Observation dcb24102-8f43-4882-b82f-cc24cc63e492 · outbound

This paper cites Available: http://arxiv.org/abs/2310.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Available: http://arxiv.org/abs/2310

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:05:24.876052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.543409Z digest=sha256:0260e36110f45687b470f28bc3e254d48c6bf1a514a46e9a2e67095c4eed5dc7

Pith citing papers

Observation f7bb58df-5525-4316-a783-368fb19742db · inbound

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding cites this paper.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:05:24.865750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:05:24.211908Z digest=sha256:eb30f0d4a39f5fa32f41b1a96b7047fbc14516d1cbd2423ec7d0704e09cea3de