Pith. sign in

Paper Citation Record · LEDGER

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2506.12222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:05:45.841972Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T05:52:55.818877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:39.910208Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4184e5d6-aaaa-4143-8a3c-35a947f6d657 · outbound

This paper cites In our initial experiments, we observed that this approach yielded worse performance compared to SSLAM on the AS-20K benchmark (39.9 mAP vs.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes In our initial experiments, we observed that this approach yielded worse performance compared to SSLAM on the AS-20K benchmark (39.9 mAP vs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:46.536287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.841972Z digest=sha256:58b51743a73a962b11193cc53554af919a4adbd9664048fec672faee86b757b9

Observation f6bfd562-bbc0-4c35-a419-d4457273e968 · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:46.767689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.741916Z digest=sha256:6ec027898d4a3754c354cf30920e59c8596d547d5684d9fee734dc6b515e2ec3

Observation c4f98f10-dac5-48fe-9f54-7f3e2be380a6 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Efficient Training of Audio Transformers with Patchout

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.433337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.433337Z digest=sha256:e1ad36c84e079f2fe45a71be8db7348862dcfc72494b979e2a805e914a1ba83d

Observation 0fcc80e2-feda-4148-88eb-e0baa5622bb2 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.507719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.507719Z digest=sha256:a6fdf5662a58afdc9fa10d634be72e5574c99c402e801798b10b91c62340613f

Observation 2e3f4533-e133-4720-834f-0c0df286f6dc · outbound

This paper cites Decoupled Weight Decay Regularization.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.590534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.590534Z digest=sha256:edbcee1a61eb2e4b6c1b902e7551c5c17fcdb209c6625617904f94a496b702a7

Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.701187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.701187Z digest=sha256:60e65d1580eb75f7baff7038f763a0604bb1beea3aea3997b16e4e0a3f7e63cc

Observation 85231792-3c9a-4dfc-b166-7954c14aa931 · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.764826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.764826Z digest=sha256:02b7989b8c754a8bf4a9c5254f74ede04ddc07dc646c15d1b29515b6bf4ee42b

Observation 964ad6b8-1c77-4194-bd9c-710b25a01b72 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.173925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.173925Z digest=sha256:56f7eeefcd2a1426f3139746fbf45efc2921c3a2ea0c173e69df9e55f307e64e

Observation e5832718-e9b9-496d-b2ae-a1ff8d684bc3 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes mixup: Beyond Empirical Risk Minimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.237444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.237444Z digest=sha256:6966ebdf3e9d06691512d3214312d19c55d25565ba316def527471084163f493

Observation 424e8316-0769-489b-a34f-997025e07cd3 · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.325056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.325056Z digest=sha256:7352ba711b33a385f98dd02e1ab16a376e21e466139e912c0312bfdba3726e20

Observation c89dbd5a-c2c2-4157-838f-f2f7eae8348f · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:47.683309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.373288Z digest=sha256:033e030e6ef9b833ae1c99d6ef0456903cb01d1ee96fbe2245c84c08a7df3643

Observation a23beb56-6b39-4f66-944c-119197902a80 · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:47.449606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.491708Z digest=sha256:6b8850d58c99e147a55a97ab4267cec5b22d70f937902937f91801c25342e990

Observation 6fe2436a-4f32-4d0d-8e0f-2730d18735e1 · outbound

This paper cites Each recording is annotated with a single class.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Each recording is annotated with a single class

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.187219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.570518Z digest=sha256:549e1dcf5ea3dcd5077d7b3f7163d48221d171a20e422f4777239807582b6baf

Observation 39447296-8e86-4078-9738-54bc04c8d365 · outbound

This paper cites multi-label.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes multi-label

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.004569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:45.645235Z digest=sha256:6ef61c3398a19332b5d6bcc5fcae1662d209985f3b4505dc486ec226b9346b11

Observation 828e21e4-3bb1-4b55-8a3a-0590664e0a7e · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.096815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.096815Z digest=sha256:87fa0444efb221851dbde6d8d80826a4761597ed412638dc23bba9caf2b7adb4

Observation 807d1e59-79a6-4c6b-a0d5-dc87f44db323 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Dropout: a simple way to prevent neural networks from overfitting

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.906560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.906560Z digest=sha256:1236a5850dfc323674649e6031857de1ac880a7bb0b6c1a7d712f18711168b1e

Observation 08a10811-8325-40d1-b1bd-37c2c3582a20 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.009230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.009230Z digest=sha256:740f50a31fb91cb48dc0a108bd32dba9345f21fbf007d015c6bd90927a0a3ae5

Observation d54b3c52-6a9c-42b0-9fd1-4079138a73a4 · outbound

This paper cites Scaper: A library for soundscape synthesis and augmentation.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Scaper: A library for soundscape synthesis and augmentation

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.893714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:44.843781Z digest=sha256:1291f2c01f7aa0b16f2e8d17ca832c12d910bf7ccd5fc528aa8b9c9e2e03b0ca

Observation aa434d91-4f9b-4397-8fd1-5e44eef43cdd · outbound

This paper cites Masked Autoencoders that Listen.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Masked Autoencoders that Listen

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.339061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.339061Z digest=sha256:2fac82df1d99e802e65726de50649b55cd949c350b5c93f03bfe489e2014a99e

Observation 61012109-f531-4814-87f9-741cb0120e8d · outbound

This paper cites AST: Audio Spectrogram Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes AST: Audio Spectrogram Transformer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.293487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.293487Z digest=sha256:6830f64a9e1f112ffdaf55f02140129325682233fd0ce9b91f75a357e475d28d

Observation 94486ff4-66e9-43fc-8068-e98be5dde473 · outbound

This paper cites MAE-AST: Masked Autoencoding Audio Spectrogram Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.036582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.036582Z digest=sha256:a03343d34a835a723847fcb7d19375772ac3cf709018fb6b437d21804a6e0d92

Observation 794bdb0a-bd06-4ee3-9d74-ceb86def731c · outbound

This paper cites BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T01:05:46.163420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:05:44.628445Z digest=sha256:2e987ebfd766092880c2cce86da9d2398dd39db05f6df6a2f75e5f6d798c12ea

Observation 0cb4f1d3-b331-46e7-8bf8-69e15104e12f · outbound

This paper cites EAT: Self-Supervised Pre-Training with Efficient Audio Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.124592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.124592Z digest=sha256:e57c72a062370d327f0ca889c042bd9dc216f68cc6da1124da8c706bca34c48b

Observation 8712af36-8e93-42c1-9bc3-40c4f46c2d24 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.165961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.165961Z digest=sha256:9fe5563ff8e4a60d768484c0516ffb00725eb8ac7f6c91cf7061bf00dc178c25

Observation 4e0cbe5f-1dd4-4ca3-97cd-d682b4c8e7f0 · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.234503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.234503Z digest=sha256:46661cb5497a12ad01725543281e200c67139df493b8a80e685d7521c5846312

Observation 5031b894-4e5b-4ec3-bdfe-42f319320ced · outbound

This paper cites URL http://dx.doi.org/10.1109/TASLP.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes URL http://dx.doi.org/10.1109/TASLP

Reference 9304

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:43.994900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:43.994900Z digest=sha256:a726b487b30bf73a1ff68c78e3307eeb8c0950c98796cc7cf3b7226b436accf6

Pith citing papers

Observation d3e344d7-7269-4e10-b44c-5fcf151a06f9 · inbound

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge cites this paper.

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:47:52.074583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T23:45:03.154037Z digest=sha256:53af1daf2143b37ebef6bf67971adf629d422eefad3675c5183e87c75c897ece

Observation 2fc140d0-23bd-4548-9c8a-8b924dd0d98c · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.911902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:023a4f2e80cb3e97d31bdd9d85ca249790064a855eab5d86f0323a7a29c20fdf