Pith. sign in

Paper Citation Record · LEDGER

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

As of 7 August 2026, this Paper Citation Record lists 100 of 166 outbound references and 0 inbound Pith citation observations for arXiv:2607.00387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.00387 v1

Coverage vector

measured 100 of 166 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-02T05:52:55.818877Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 166 outbound references displayed

  • verified exact27
  • verified fuzzy69
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15fb7307-b67c-45a7-a1d2-53679810ab05 · outbound

This paper cites Deep learning.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Deep learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.354732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:2808fe782ad6ce2afd8dec4fa0de89496073eeafba4c5f2896f6317cb149ad3c

Observation b09b9182-b6bc-4a2d-bb4d-4c52cbef2beb · outbound

This paper cites Deep learning for audio signal processing,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Deep learning for audio signal processing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.356882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:fbed820d7424b96f0a40c15400e0d89a817cc7c1afc41fe180d10bc3dad3f306

Observation 5d8eb1d8-ca86-4937-9f54-136a52d8dde2 · outbound

This paper cites Audio signal processing in the 21st century: The important outcomes of the past 25 years,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audio signal processing in the 21st century: The important outcomes of the past 25 years,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.257547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a9e246ef08b749b901f6753e7f77b6c15b8d23700e41cb91ef8357daef015cd8

Observation d54d9a1c-6fde-49e0-89a0-2fca90ccf7d1 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audio set: An ontology and human-labeled dataset for audio events

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.220665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:57db7fede0414f53c456f45f7cdde2d6a9070facc96e4c9ccbd382fc447eb0a4

Observation f5a27b21-9c24-4a90-9855-b47d18d45180 · outbound

This paper cites MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.884944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:5daf79d7ad6295fd3b776358640b478167f5ad09fdec4263746f6e33feafad26

Observation bc9d630a-b491-456a-bbe1-b662b2733a51 · outbound

This paper cites Taming data and transformers for audio generation,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Taming data and transformers for audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.237091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:cc4541b071a438f5ceb0f123b398263d68c7459c5cf90f8b4025d8e943c42d78

Observation d9698c30-4990-4f27-ad34-fff468b6eebb · outbound

This paper cites Deep convolutional neural networks and data augmentation for environmental sound classification,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Deep convolutional neural networks and data augmentation for environmental sound classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.359039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:886992b2312ea8bfea9ab48f64b6660720d68cedef78de525b39756ae2079971

Observation 1f28c6f7-96b4-4c15-93cf-7e9832a83473 · outbound

This paper cites Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.881972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:7c4d74c5061501578f004e8e2d4cf1ede443f5420d83be05ce6c940708c973fe

Observation c0593eb1-4fd6-4ead-8a07-8034fbfc99eb · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Explaining and Harnessing Adversarial Examples

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.887680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c716e81de7874b88be2c4161f23dff706312c3383e9b8d62b6282eabaaad70fc

Observation f3e5354a-08f0-4a63-ba49-0c7f789b9133 · outbound

This paper cites Intriguing properties of neural networks.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Intriguing properties of neural networks

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.876297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:377a8b36178384cd9039f4aae51a9bb4613e1928f0252b4caf0766ce90e6c830

Observation 2ea35bd5-ebd5-4fd9-adea-31c56068fcaa · outbound

This paper cites Understanding deep learning requires rethinking generalization.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Understanding deep learning requires rethinking generalization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.922586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:21aaf0567891c1f3da32f1c94725e757514730131d189f61fbde5509d783f6c0

Observation d26ce9db-c44b-433a-8ffc-a124754bf397 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.839281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:585ad36b1e99c78254dbfc49941273a7e3cdf29b51fbe33e1efe27dc1f55366a

Observation 7c5a3410-c7d9-4c56-ae47-80359d2b001c · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.865166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:8926829886cbafcf9aeb2f61ab8a8872a73cc2838acdfe0e4048854d6dba69ae

Observation 9db5be48-9abe-4f4d-a24f-f9960962f9f0 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.370879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:ee98f679bf3013932745b95161bb17882d0fce282573d3648b06c757c83adef4

Observation 7e034ff8-c32f-4c80-a6c5-71d03422ef6f · outbound

This paper cites Musical genre classification of audio signals,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Musical genre classification of audio signals,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.201772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:0684cd3f0470486366dc774338b137af22b963ae3ec9ac505133f46c0f312072

Observation c71a135f-dc0e-43d6-803c-67325fd54f4e · outbound

This paper cites Dynamic attention-asymmetric perceptron network for overlapping sound event detection,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Dynamic attention-asymmetric perceptron network for overlapping sound event detection,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.241172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:f7145ed83375492e6db9f04c64b5fc19778b29a92dd76af3117e89d4a36d78fa

Observation e3ba9944-2a80-4eff-95bd-156255123b72 · outbound

This paper cites SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.893677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:230387c50e645589ce1f9368d36574e4785b87050e6c9d11da37fa43bbf5f5f4

Observation fe046f72-72e3-4569-a496-4df20a5c0450 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Scaling up masked audio encoder learning for general audio classification

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.845086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:dfc17b57f2255e567a4dc751c89801bb7118376efb337d1e74216ebcd0a2ea21

Observation 481eff0a-cddb-4d9e-84a4-70be59d97455 · outbound

This paper cites A survey on contrastive self-supervised learning,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning A survey on contrastive self-supervised learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.172411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:32003f1644082c2f3d1c085f5b2b3ce2073ffb47b4b48529f87e14573fa094e5

Observation f1b9320e-9f57-4860-9cb3-00449f42f143 · outbound

This paper cites Audio self-supervised learning: A survey.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audio self-supervised learning: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.223855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:837c4c3e084d6536811fa59245db512cbeba169c610bcc051aec5fe97f8bb565

Observation 4a149a3f-1bf0-4a07-ad24-a4cf9d6518a0 · outbound

This paper cites Scaling bioacoustic signal pre-training with million samples via mask- modeling,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Scaling bioacoustic signal pre-training with million samples via mask- modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.234892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c42f23ce8797d22d87f6015942de5f35915dc6e67423aef1a4e09b59f87bdcd6

Observation 72e0138e-a7e1-4d02-b026-8ba09aaee080 · outbound

This paper cites Hearsay benchmark: Do audio llms leak what they hear?.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Hearsay benchmark: Do audio llms leak what they hear?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.913510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:0ad70f9382e3b591bd0e4624c186a6e7cf4cac050ca3d372424b31f40f85115c

Observation 81fd931b-6c9f-4d44-942a-f7d388e79ac4 · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning A survey on self-supervised learning: Algorithms, applications, and future trends

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.180032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:bc1885c8bd73eb49ba719393690f768f24a5ef6af9005ed02c3fbf51f4d7f99b

Observation 347f6342-5f94-4917-9acb-daf95f526d97 · outbound

This paper cites Self-supervised speech representation learning: A review.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Self-supervised speech representation learning: A review

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.156645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:b35d74acaced31aec0c7b77b3806c8cbb09729a5a99c2b5e7d6d8041ec83d6b9

Observation 22ebae97-e555-4fdd-88c2-754292e57721 · outbound

This paper cites Ssast: Self- supervised audio spectrogram transformer,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Ssast: Self- supervised audio spectrogram transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.170725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:0874bb1b71b9281c7fbb2a77bb9c5130f68d5ebade913bf9bb57307b80703496

Observation 84be8aab-8c1e-49b8-8a5a-32cf5c7c9888 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Unsupervised feature learning via non-parametric instance discrimination,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.247623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:3dbe826df2ead8639e29dc8f982f15ba8f6c58b9ed6c8709780096a86d625cde

Observation 0bd43875-1f31-4841-88c4-bd1ee41e19b3 · outbound

This paper cites Unsupervised representation learning by predicting image rotations,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Unsupervised representation learning by predicting image rotations,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.253621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:15b623725c07e4932851c14d8ae52528a0d0f3635eb348b97fed1fe843783c0b

Observation e9dad6be-bad1-42f4-9e46-3f2e68bad82b · outbound

This paper cites Unsupervised learning of visual representa- tions by solving jigsaw puzzles,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Unsupervised learning of visual representa- tions by solving jigsaw puzzles,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.228459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:f68882460a2303d0d7489abf23a6a243b7efc0d63c9f1d76df12f9ef25b462e0

Observation 21ce1088-e49d-4434-96c7-172ec12c59b3 · outbound

This paper cites A simple frame- work for contrastive learning of visual representations,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning A simple frame- work for contrastive learning of visual representations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.428012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:d5799fc66f7544c420d0a62f294b78c15c39adb3df5afaf96b0b9c821b25985f

Observation 7ffbe064-3445-4cc0-9618-82661c251fa2 · outbound

This paper cites Contrastive learning of general-purpose audio representations.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Contrastive learning of general-purpose audio representations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.425492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:92643ea69a6eb09ed624bb93ab51124afa7d78ec11541cc247b468a6f15b6f50

Observation e94cc65b-c540-4630-8523-ea79a20a5e46 · outbound

This paper cites Masked autoencoders are scalable vision learners.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Masked autoencoders are scalable vision learners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.430050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:336840a65ee7764ad3f494b1a045dfd28a98f7fa79dc48fa605a4f70cc5dee0e

Observation be3db07c-2e93-427c-b69e-798f0e3653a7 · outbound

This paper cites Masked autoencoders that listen,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Masked autoencoders that listen,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.441242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:29a185d4018ff794ceeccb39e89d2d0c03916b0c735831669286fd17dc3a1a2b

Observation ec4b0e66-dafe-453b-bfad-19c16677cfa4 · outbound

This paper cites Masked spectrogram prediction for self-supervised audio pre-training,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Masked spectrogram prediction for self-supervised audio pre-training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.361149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:534defbbece4fd44ce353d9ff7956438465a82971e7fa4455dd06b3cdc470712

Observation 7447cbed-c823-463c-8491-8f2ee74bfad9 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Recent advances in discrete speech tokens: A review

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.454836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:099213949897953139a2b553618f974bdadf99f1b822071278f69bcd54cc7ffd

Observation f7a930f7-f2ed-43c5-8098-c40fa06e921e · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.917414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:7099b7ab1d922d998f69bd85143f1da0adc31d36e49bcbe8c110c6013325a6d7

Observation c37c3c9a-4f24-469b-9b7d-13e2f7b3faf3 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.415039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:929967ada995c4711b6a5eac0909a14206c5d6db2e771194ce2af5491f136eed

Observation 4de403c4-307c-4cdc-9964-1f5a62be1a00 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.452408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e0eb20894c49315e921436af65f8ddeb7f4dd3b89bfd8434cc0312b11303fa1d

Observation e32121c7-b816-4b15-a5be-086d5824a4ef · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.443320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:fcc9a8e2f2b6ae2360b66f6f900127946c9315e668483e01614c513c8fe5c020

Observation e99942a6-4d46-4b9d-9083-5efb883c2094 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Clap learning audio concepts from natural language supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.407150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:333413c3350674dec69d446ef6a1c5b58af77a7d8b78f61e23f117f3276c5eba

Observation f5d8c9f3-b88a-4343-96dc-dc6752ba8468 · outbound

This paper cites Pre-training audio representations with self-supervision,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Pre-training audio representations with self-supervision,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.437006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e06d7bfe0fe25b126e8704c3d1fa25e70e44e36e3e92115cf518e4859f99f8fc

Observation 94c3c24e-d8b9-4854-abb9-db81a73bc0cb · outbound

This paper cites Shuffle and learn: unsupervised learning using temporal order verification,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Shuffle and learn: unsupervised learning using temporal order verification,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.439193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:b894217dacbb00eb9a58cfa87e29ae01cbbdecc2a030cd2a97ce4d2db12e2ec6

Observation 950485b2-41d6-4fe1-83b8-9a37ce1cdcdf · outbound

This paper cites Self-supervised learning of audio representations from permutations with differentiable ranking,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Self-supervised learning of audio representations from permutations with differentiable ranking,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.400585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c1385cc0946d87bdbc07e8f80d454bd8f0829ec6d2fd7f0c7ba5be89cf30c867

Observation 2e338edc-21fc-4abc-9e48-9aeed74f2cd7 · outbound

This paper cites Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.896395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a76d0d737ee32c9e38569b861603710b76f697e150e05adee28e5a1218071201

Observation 539ae367-2a03-44ff-b666-3652f12aa5f4 · outbound

This paper cites Multi-task self-supervised learning for robust speech recognition,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Multi-task self-supervised learning for robust speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.404619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:5553e1e6ee865197851b692d5c69483e64906692748e133433b63ecc5b4333b4

Observation c2f634c1-db28-4357-94bc-91db13fe66b9 · outbound

This paper cites Clar: Contrastive learning of auditory representations,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Clar: Contrastive learning of auditory representations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.386271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:38cf355ffdcffb9bea99d89796010fe7a416d166f3a93f139b002689b153bdd5

Observation f331f772-8796-44e9-9838-cb4b07a675d5 · outbound

This paper cites Byol for audio: Exploring pre-trained general-purpose audio repre- 22 sentations,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Byol for audio: Exploring pre-trained general-purpose audio repre- 22 sentations,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.393585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:98896ec8304ee217262b513510ef27f730d12752fdfd2e255ce9c699e0c8cb77

Observation e4679cb1-8c95-49ee-a2be-dcbcd1130fe5 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Representation Learning with Contrastive Predictive Coding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.898991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:00e8078a24832f105c83d7563abcf8b95e1f0d558311bfb7e22865a7018c2d27

Observation 39855e98-1e4a-4296-814b-5edd9143696e · outbound

This paper cites Byol for audio: Self-supervised learning for general-purpose audio representation,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Byol for audio: Self-supervised learning for general-purpose audio representation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.447668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:dcfe5c59f31b984c4379055b70549fea8e93b74eb2ef6e519e492abb808943c1

Observation c28c2cd5-21e4-4a01-bcb2-a7bed13bdfa9 · outbound

This paper cites An Unsupervised Autoregressive Model for Speech Representation Learning.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning An Unsupervised Autoregressive Model for Speech Representation Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.910790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:3135cb069efb8972c7c02de20af1db230e03dc3c35640c57874f93eb52d42079

Observation 4526c40a-3467-4301-b3a9-b326e4f8b846 · outbound

This paper cites Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.383858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:fc7ced060c54f832b07e09923269c4093ada55fd16dbf458435723ed232e93b0

Observation 8ab2afd9-9ca3-4e50-acd4-93ffb9e8f98d · outbound

This paper cites Audio albert: A lite bert for self-supervised learning of audio representation,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audio albert: A lite bert for self-supervised learning of audio representation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.409863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:9e5eb03c252fdcfdb995d18ef87270d6cb46f92ff274d303fa3597fbcae5d2c6

Observation 5e014b4c-f80e-49d0-b73c-8c65e9531226 · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.417838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:7aae5e9b95c6bf9c39e5a23e4996e48bcd91131f758f3dcbcba15af509335696

Observation 418d0736-613b-4ee4-98d4-f3e77d558404 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.856520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:ba942e2b551c439650105860571fc65defc36cb3c997f0b9dd54bb841663c521

Observation 184b1b0b-25c5-4db5-84e2-06b255d85cd3 · outbound

This paper cites w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.457041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:72e16dd7f00f7dc6649e6cb81d4db0f860b10f86b481a8ec5b0ed2934cc13de4

Observation 8671d9b6-11af-48d1-9654-13eb951fece5 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.366138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:837874748747d894df4a50ab1719702cc71b994ce3a131d641e611d378b91a81

Observation 7a747007-214b-4571-a80e-cf14747836e4 · outbound

This paper cites Discrete audio tokens: More than a survey!.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Discrete audio tokens: More than a survey!

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.932308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:91caeaf3fd1e6c0173b3189345a4b4e7941d981356a7053ce6a83fcba6d83726

Observation c58dfbec-4b83-49b6-89fb-1b284891cad0 · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Data2vec: A general framework for self-supervised learning in speech, vision and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.373133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:d2acc87cca571cd55ffc9cd00f9237ecb896baf0560744d938058a45f6b686cd

Observation afd95574-7c8f-41cb-a244-7bbfc574dbc3 · outbound

This paper cites EAT: Self-Supervised Pre-Training with Efficient Audio Transformer.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.916640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:6c8a41d99e9eeff220534a44766d4c13fc6fbbfad29e48864739aa568df07150

Observation 9d8dd83f-1272-4b98-b238-0291ef7e9ef2 · outbound

This paper cites Look, listen and learn.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Look, listen and learn

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.363660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:d157a851ea30b9a3a1d9d4f412a42adaf24a3fc35d037bd6acea384bb4566bbf

Observation b0f5883d-5eb2-45b7-9102-60fa7b250c75 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Soundnet: Learning sound representations from unlabeled video,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.376544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:b8dba16812e27b5988b9bb3363715f19aa4f1171393811e6a4de5131ac8128a0

Observation e4c916b3-53db-4971-a3ad-ddf3decbcb74 · outbound

This paper cites Robust audio-visual in- stance discrimination,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Robust audio-visual in- stance discrimination,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.350626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:2c23d862a85785b602121c061380b05bf73b65807e02cc5c324d1046ecdc0c92

Observation 4b9c210a-1092-4f56-b8dc-72802efc662c · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audioclip: Extending clip to image, text and audio

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.217968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:cfc3ce8ea557129f13ec6a1eefe8eadcfb0cfd774926eca2e86211318982be58

Observation ab04c334-86a2-4bfd-ba50-c9208a35e265 · outbound

This paper cites Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.348438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:8e2aa06c537abfb5783f4d5d4025c746e127da0e45316973362ce8e0e8b2e6f4

Observation 7d8c8085-3369-4f67-a2a9-ff65d861caf5 · outbound

This paper cites Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.798062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:3686f0c9cfb6150941182c5211c617c4240f5824d78c9c512f1b1a8b709f6914

Observation f73f66d7-732d-4d73-a7eb-0f6c06f869d7 · outbound

This paper cites SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.815085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:ac67ee0596747f332603d33adca0ac01eeccdb3e7cdabd0e9378a54e7ff24fe5

Observation f5ab6965-481c-4113-8987-115d97b9ebb1 · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Mamba in speech: Towards an alternative to self-attention,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.352668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:85b8d6c6cd9da4ba7651d576a49fb96816312e1ca812f0be3634f52ec2425839

Observation 5c283c3e-87ff-46fc-9488-c38df3c319e1 · outbound

This paper cites xlstm: Extended long short-term memory,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning xlstm: Extended long short-term memory,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.198951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:8080b6c610abe9ae356594bd070d0e57441c45042546708243c2e5ea84cbf3e7

Observation 8ba750ba-0085-46ab-a873-33fd755affa7 · outbound

This paper cites AxLSTMs: learning self-supervised audio representations with xLSTMs.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning AxLSTMs: learning self-supervised audio representations with xLSTMs

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.837872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:bd05bb8de4ec816bf11f139f3ec9d6de5a661ec79ffc2ff365c4e6ae45ce4794

Observation 03da306f-1914-42f7-991a-538d5156499a · outbound

This paper cites Tera: Self-supervised learning of transformer encoder representation for speech.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Tera: Self-supervised learning of transformer encoder representation for speech

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.412584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:41546b2c571db318d528ec7000908d9d751b3c3a0dcc8cd7847e224ab926a463

Observation 8376e35f-5aeb-4bfc-aea8-66cf16962f44 · outbound

This paper cites AST: Audio Spectrogram Transformer.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning AST: Audio Spectrogram Transformer

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.820388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:55e0b5db8c8b0e251bbd355559de09d515149af843488483158124687f1c92d9

Observation 8d77ddb9-f659-4438-ba77-924dbc38be1f · outbound

This paper cites AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:56:39.906436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:091cddf74549514958bda6ca404c8a388bd1fff0ae8743e3df95bf7ad537a034

Observation 4ca4d881-c727-4479-a123-a09b332b838b · outbound

This paper cites Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.434677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:4e56d400ee5a75ceff52757de62b6298eb12e7547d700c77a1e1a9865118c808

Observation 880f7438-026d-4164-9e5e-bd9f6409e7a1 · outbound

This paper cites HyperConformer: Multi-head HyperMixer for Efficient Speech Recognition.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning HyperConformer: Multi-head HyperMixer for Efficient Speech Recognition

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.812037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:84bfc6e5ff9b6e729bbe9788dcb56804cbdefb5b3f3b4fe2dd85a4d03f20a267

Observation 33e22351-0061-4fd2-9fa8-035764780e79 · outbound

This paper cites Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.343513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:4145e99908b42c172946cf70902b007a700a3586c840f7962f3d67b48a373593

Observation 55f5ba7c-e139-448d-aa7a-3c92c36ddc1b · outbound

This paper cites E-branchformer: Branchformer with enhanced merging for speech recognition,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning E-branchformer: Branchformer with enhanced merging for speech recognition,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.336962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e0f9b544058d58341046ba4d179462190fce9e807bdead8745c712a0c954c49d

Observation b2610e2d-264d-4edb-aa02-8b362b035575 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Zipformer: A faster and better encoder for automatic speech recognition

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.902887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:3d0f3db51fee4b33acc93bb4137eba46bd8b003e697d071f678e1824a6eae45c

Observation 370cd1fc-f671-4e2f-a54f-28e9d9802468 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.267939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c99922c24b4384b8ce38fe064b626f5983b50aa14640c04dbad1930ac7e93831

Observation 1231c773-ab75-4841-b562-f728bf34f504 · outbound

This paper cites Joint semantic knowl- edge distillation and masked acoustic modeling for full-band speech restoration with improved intelligibility,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Joint semantic knowl- edge distillation and masked acoustic modeling for full-band speech restoration with improved intelligibility,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.331578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:3f678436e369550537b8daf144ff3434feccdf2bc632b76398afe41182b05c9c

Observation 50d28ddc-3d53-4abd-90f2-5ecee7f34da7 · outbound

This paper cites Distilhubert: Speech represen- tation learning by layer-wise distillation of hidden-unit bert,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Distilhubert: Speech represen- tation learning by layer-wise distillation of hidden-unit bert,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.333547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a0e8a353df0a9a428c9e2bb80fb74a65a1efa18e9910e616764503377502951a

Observation d88e755e-0abe-4657-ab26-18f59f406fc8 · outbound

This paper cites Skill: Similarity-aware knowledge distillation for speech self-supervised learning,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Skill: Similarity-aware knowledge distillation for speech self-supervised learning,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.323664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:d5af5f79eae4a0972ad2106dccab259855c8e1444e60696f2e5519a96aed8fd1

Observation de188346-7ff0-4192-880b-24806483e9be · outbound

This paper cites Semanticodec: An ultra low bitrate semantic audio codec for general sound,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Semanticodec: An ultra low bitrate semantic audio codec for general sound,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.461040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:ad65ca4401afa3437407f78bee01121c6d453c7d47c245b4b5a3c7aa3d2ed320

Observation 738c50fb-c616-4bae-9ec5-c8addfe11d27 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.919713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a5d767ea675a96bb71bce3b42a532148f9c29bcab003185ee7d5c9738162ff46

Observation 5fbd5ea8-fb63-4e40-b732-416492efbf74 · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.890564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:5474856b047fb4c02c8d2f813d09e00abaccbc9850df37300473747f5f05e122

Observation 729e9ef5-03f7-422c-862e-e34ec139aba4 · outbound

This paper cites Codecfake-omni: A large-scale codec- based deepfake speech dataset,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Codecfake-omni: A large-scale codec- based deepfake speech dataset,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.422988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:84f0753856e8313b4b24824aa03bbe74f921373cfe705e7423421d06450ddc62

Observation 9bf6d11c-0b5f-4a07-9949-777d777b2bfd · outbound

This paper cites TS3-Codec: Transformer-Based Simple Streaming Single Codec.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning TS3-Codec: Transformer-Based Simple Streaming Single Codec

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.854827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:9aa178ede4625260be5cc0883c2b6f18fbd228a4ea8d03d10b189987b222a523

Observation 707f8a58-c63c-46b7-b60e-1c704da53897 · outbound

This paper cites Fast- hubert: an efficient training framework for self-supervised speech representation learning,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Fast- hubert: an efficient training framework for self-supervised speech representation learning,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.321801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:66280e647291c532cc42793a91ef66f331176f84251cb3cee95a2bad18abc556

Observation 81a34224-e9bc-45ea-9f3a-77d0da107583 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Cnn architectures for large-scale audio classification,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.261383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c898131ac181154ec9922522918c91c8c0735f06a0a6aa8b9f3deb509db7a474

Observation d8255aa3-9af7-4e4f-aaca-05ad8ab16334 · outbound

This paper cites Long short-term memory.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Long short-term memory

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.320048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:78fb287ad38a59a38d6ff97021d1724d2ab8401184d6a9c493a8e80a403e2391

Observation 9af83738-121c-4722-92c4-444b91b4a13a · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Mamba: Linear-time sequence modeling with selective state spaces

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.327389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:af2a5cdc4974c5fb13112787de3ecaf4260c839be7156b4b6ef8435b37b36b59

Observation 759a09c5-31ee-4f15-9663-88093c3b3e71 · outbound

This paper cites Attention is all you need,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Attention is all you need,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.340711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:f32e30eb4d572590f6a4475075749e084743dddbfc9e8cd351bf2202d5060ac8

Observation 3c7bc2a3-01ec-425f-a7c0-b6088e83b0ec · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.297771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:9c1e2a336977b37d0c885805f9874ee9013b7f5e62ae7086309db48bdb40a587

Observation 9268124b-a2f4-4f8d-aea6-3087f0b1c290 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.897175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e2b9ecc94832b7afbcd7e59a7dbce5fc032c7a5f4a636d9a7f3ef56a7de90793

Observation abbf0d63-cf32-495b-a9cb-570350d25ac4 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.846458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e2b0cf8882f5bc697a67276eae4287d47a3a60edfc22424920172302032b44f6

Observation 379becd0-0e5e-4527-8962-460a79b8042a · outbound

This paper cites Investigating self-supervised learning-based front-end for multi-channel replay attack detection,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Investigating self-supervised learning-based front-end for multi-channel replay attack detection,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.313603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:6ab26d17cb052ffbf73a3ef610e6295dea0ef997641993a1cdc9844cbc70b145

Observation 2fc140d0-23bd-4548-9c8a-8b924dd0d98c · outbound

This paper cites SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.911902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:023a4f2e80cb3e97d31bdd9d85ca249790064a855eab5d86f0323a7a29c20fdf

Observation 833c88ef-6e09-4824-93f5-67ef7b9c71eb · outbound

This paper cites AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.830382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c5e76cee4005f96dbbf212ce265a3f3105426a905ed1d614aec63868d5f7c9e6

Observation 4fadf425-000d-4db8-af5b-f1363027a95c · outbound

This paper cites Dp- mae: A dual-path masked autoencoder based self-supervised learning method for anomalous sound detection,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Dp- mae: A dual-path masked autoencoder based self-supervised learning method for anomalous sound detection,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.318215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:1c54f0962d073df878872f46bfbd412bad4ecdb47fb4edc3d40a3bda6f76ab34

Observation 60b87f4c-1ab5-45d9-afd8-2cf0790c4242 · outbound

This paper cites Muq: Self-supervised music representation learning with mel residual vector quantization,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Muq: Self-supervised music representation learning with mel residual vector quantization,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.263379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:6bda01e06a6a024d93814711717959b4e21f285a478af58bd0289b529bb387b8

Observation d102e475-e100-4d7d-a099-c6fffd12f62b · outbound

This paper cites A foundation model for music informatics,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning A foundation model for music informatics,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.265159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:ea6f912faa0800d8dafa2e8781a2b25779d1eb01c0037bc552e517cb68ed0abd

Observation 4beb6b9c-57a5-4b25-b726-5ce93408a548 · outbound

This paper cites Unsupervised sound separation using mixture invariant training,.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Unsupervised sound separation using mixture invariant training,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T06:52:04.259505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a0931814da0c2c10b9f8694b4ef5db0dfdc14e63c733c884cdc2673399877b47

Pith citing papers

No inbound Pith citation observations are available.