Pith. sign in

Paper Citation Record · LEDGER

USAD: Universal Speech and Audio Representation via Distillation

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2506.18843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18843 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:26.977695Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:27:05.187538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:41:06.244426Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9803bd2d-a80e-4467-8fd1-1bf7fce1c619 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

USAD: Universal Speech and Audio Representation via Distillation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:20.950820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:20.950820Z digest=sha256:12bc09e36445ad7d8e990b9d01b7750281fefd9fbc83d169e93ff226c0eafa56

Observation 2857be56-1e71-4949-81eb-aaefadcbb708 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

USAD: Universal Speech and Audio Representation via Distillation Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:35.149503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.019181Z digest=sha256:b1683089ed63aad193432614f4f6a8eebe0e22098f761d7c133846635ef7fd5c

Observation 9580200e-b490-45ad-a885-d27f022ad6f8 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

USAD: Universal Speech and Audio Representation via Distillation Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.969084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.120019Z digest=sha256:06508b7846681ed76f56c0e4746d249cb05c2b45ca207b9d922a2da3ac6dc728

Observation bd83d77f-d534-4744-998f-41adca1bfe1c · outbound

This paper cites Ssast: Self-supervised audio spectrogram transformer,.

USAD: Universal Speech and Audio Representation via Distillation Ssast: Self-supervised audio spectrogram transformer,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.783760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.326189Z digest=sha256:e231354e122e43359fd053b26b484716fe23d29eeac0665e753d065ce4f27e4a

Observation 96f93f5c-1c38-446e-9399-342e15204231 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

USAD: Universal Speech and Audio Representation via Distillation Beats: Audio pre-training with acoustic tokenizers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.540726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.493407Z digest=sha256:a6dc361281c5b86191e35c37687a3864ffa7cd86683078ca6c8410c8fd09dedd

Observation 4aa95dd6-9ff3-465d-bed9-f5714ff419ff · outbound

This paper cites Mert: Acoustic music understanding model with large-scale self-supervised training,.

USAD: Universal Speech and Audio Representation via Distillation Mert: Acoustic music understanding model with large-scale self-supervised training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.366351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.664948Z digest=sha256:d82a09d28b3322db2bb17c44eae845279b776f4419e55b061107879a9301a680

Observation 0ac23d43-ed8b-4c66-b5d7-b980a7db7d3a · outbound

This paper cites Listen, think, and understand,.

USAD: Universal Speech and Audio Representation via Distillation Listen, think, and understand,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.178286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.782308Z digest=sha256:f873a3b5aed38138b036df42a0233ce115fd0ff991d861973db9a7a618beefb7

Observation 2797e9a3-3d69-4259-a68d-31139cf72acc · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

USAD: Universal Speech and Audio Representation via Distillation SALMONN: Towards generic hearing abilities for large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.985657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.848096Z digest=sha256:cd8e7a4a19fb3807646c7cfc5759ae221a5e02cd86eec9f6fe75f8748482e8d7

Observation adf77828-8d75-4f97-ab06-0a5b90908caa · outbound

This paper cites Qwen2-audio technical report,.

USAD: Universal Speech and Audio Representation via Distillation Qwen2-audio technical report,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.868863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:21.989905Z digest=sha256:49348e90aac0689608d2b7a8da62d262ea31de43a124edf8fd907761f3d41915

Observation 830cc216-8902-46d8-85d0-feafef981276 · outbound

This paper cites Gama: A large audio- language model with advanced audio understanding and complex rea- soning abilities,.

USAD: Universal Speech and Audio Representation via Distillation Gama: A large audio- language model with advanced audio understanding and complex rea- soning abilities,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.707482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.121921Z digest=sha256:4039de3cf1030dcbea66b608aebd9c007cdf2ceebc3236ecc15d27e283d2f82a

Observation ed7d1c45-9811-4b44-876c-b7527602b483 · outbound

This paper cites Google usm: Scaling automatic speech recognition beyond 100 languages,.

USAD: Universal Speech and Audio Representation via Distillation Google usm: Scaling automatic speech recognition beyond 100 languages,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.583245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.232038Z digest=sha256:8b46568c14b9f8fc71275ab272b07226532dbb0e4b0538f60100e82f23e0514e

Observation 9b99dd78-a658-4556-8360-1871eac56571 · outbound

This paper cites Speechtokenizer: Unified speech tokenizer for speech language models,.

USAD: Universal Speech and Audio Representation via Distillation Speechtokenizer: Unified speech tokenizer for speech language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.479959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.385294Z digest=sha256:bee0940c9ca4d83b5569fb0b5250402197dd8a28aa24d85b0be8cd767cec3a96

Observation 1aeebdb9-51bb-4a6d-9d2e-f833a35391ac · outbound

This paper cites Soundstorm: Efficient parallel audio generation,.

USAD: Universal Speech and Audio Representation via Distillation Soundstorm: Efficient parallel audio generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.361483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.542181Z digest=sha256:4e3736391c703cc7cc2f4e260028b58032bfd4b2ed991d4464806f8db58ed5fa

Observation f1cb00c9-0acf-4673-8a29-054a36ec84dd · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

USAD: Universal Speech and Audio Representation via Distillation Moshi: a speech-text foundation model for real-time dialogue,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.210165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.709397Z digest=sha256:d898e1ae94d51981972f38257a98e8c499d1f7d2c33b4f9bf36972a0e7dcdfc4

Observation d9e96910-e0ac-4769-896e-4ecfa949cbbd · outbound

This paper cites Dc-spin: A speaker-invariant speech tokenizer for spoken language models,.

USAD: Universal Speech and Audio Representation via Distillation Dc-spin: A speaker-invariant speech tokenizer for spoken language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.081468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.906742Z digest=sha256:f74daed34ca30800aea601293dd9a922eb422116afe22e232ed716e21e615e98

Observation a0c42ab5-3310-4048-9006-6d2db6be8ec1 · outbound

This paper cites Joint audio and speech understanding,.

USAD: Universal Speech and Audio Representation via Distillation Joint audio and speech understanding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.962859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:22.997389Z digest=sha256:ff2ba76c97fc1aefbc23579297707338e0e90954a8c85cf3c18e5c98a7efe0f8

Observation af58277e-7e2c-4572-be75-91a401135ca4 · outbound

This paper cites U-sam: An audio language model for unified speech, audio, and music understanding,.

USAD: Universal Speech and Audio Representation via Distillation U-sam: An audio language model for unified speech, audio, and music understanding,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.826324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.094069Z digest=sha256:a0c6836827d5bd377dd660bf4fec22f4af5f3637d6152a52b0e91856ab7eb18f

Observation c3385d0f-2302-49f3-b9cb-942180b51746 · outbound

This paper cites CoLLD: Contrastive layer-to-layer distillation for compressing multi- lingual pre-trained speech encoders,.

USAD: Universal Speech and Audio Representation via Distillation CoLLD: Contrastive layer-to-layer distillation for compressing multi- lingual pre-trained speech encoders,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.702387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.166044Z digest=sha256:02b917b68b7584a5533f12d45bd06fcabd9072c4369cb796c33215a760010742

Observation b2340a75-f99a-4c4b-9cc9-130ac425ffb7 · outbound

This paper cites Mae-ast: Masked autoencoding audio spectrogram transformer,.

USAD: Universal Speech and Audio Representation via Distillation Mae-ast: Masked autoencoding audio spectrogram transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.622842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.277080Z digest=sha256:3acde41efcd4f98b811aa50e88a419b8ce8030625820006c76f4057cc06ee8f4

Observation 21f37735-0ccd-47d3-82be-a112d6de1950 · outbound

This paper cites Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,.

USAD: Universal Speech and Audio Representation via Distillation Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.506691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.390695Z digest=sha256:ad905c4ec1049c5bf2c44f492402ea5c0b6bd117953479ef129d0d6731f7a153

Observation f0bd5528-3a08-4a0c-8b9b-690bcee749fe · outbound

This paper cites Masked autoencoders that listen,.

USAD: Universal Speech and Audio Representation via Distillation Masked autoencoders that listen,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:23.535235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:23.535235Z digest=sha256:3380a1c59e77c4d217ea296fa64ac60227492ac8de09f97a473fde1af9489a00

Observation b12c6f02-375e-4ffc-9089-6363a8b22cf3 · outbound

This paper cites data2vec: A general framework for self-supervised learning in speech, vision and language,.

USAD: Universal Speech and Audio Representation via Distillation data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.295570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.677938Z digest=sha256:6ddb7788e2d1a2891e69e1ade7805c31ad9f5a8b0a64abad91630cdce3cd90ae

Observation e8a8dcea-19d2-47e1-980f-e119fa9246b5 · outbound

This paper cites Efficient self-supervised learning with contextualized target representations for vision, speech and language,.

USAD: Universal Speech and Audio Representation via Distillation Efficient self-supervised learning with contextualized target representations for vision, speech and language,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.234671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.826605Z digest=sha256:47e31b1cf156d9ec5fd270198950824a4745775cf0c6c68eee7760132ec329f3

Observation cb9532be-b1a9-42b2-bd44-e3c451245d5b · outbound

This paper cites Dinosr: Self-distillation and online clustering for self-supervised speech repre- sentation learning,.

USAD: Universal Speech and Audio Representation via Distillation Dinosr: Self-distillation and online clustering for self-supervised speech repre- sentation learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.101137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:23.957169Z digest=sha256:d0dd6c875ad01842d1b817ef9dcc11e1c49375fc460c7578fa6942864662789c

Observation 2681e63e-d579-4757-b997-a9e21f2a0df1 · outbound

This paper cites Eat: Self-supervised pre-training with efficient audio transformer,.

USAD: Universal Speech and Audio Representation via Distillation Eat: Self-supervised pre-training with efficient audio transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.015670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.031734Z digest=sha256:baa0b7f1a42530669cef892bb32498ef49ab3c642a9c923b7d5c724838361c1a

Observation 59f5c611-2488-44ba-9586-f5601e530526 · outbound

This paper cites Sslam: Enhancing self-supervised models with audio mixtures for polyphonic soundscapes,.

USAD: Universal Speech and Audio Representation via Distillation Sslam: Enhancing self-supervised models with audio mixtures for polyphonic soundscapes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.851676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.185064Z digest=sha256:56f9409a5b1fa46af2b8beda3c3fa909b0e030ee6b8300aba977b411c34b2363

Observation 5afdf0c6-b789-492e-9204-ea3c020b92c7 · outbound

This paper cites Byol for audio: Self-supervised learning for general-purpose audio representation,.

USAD: Universal Speech and Audio Representation via Distillation Byol for audio: Self-supervised learning for general-purpose audio representation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.704745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.281964Z digest=sha256:4e70eccf383b0c955d7ae20a3b606cf6c814a89c42dee4daa918f48f969538e4

Observation 0f8249ce-4a93-4bc8-a80e-58a6b4792f51 · outbound

This paper cites Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,.

USAD: Universal Speech and Audio Representation via Distillation Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.535281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.398795Z digest=sha256:80f53b8f64f30cb562cae806fc9634597bfd8cc11780d138ecd76c19c9483822

Observation 4c8d1068-292d-4fbe-b655-ae6e422b240e · outbound

This paper cites Masked modeling duo: Learning representations by encouraging both networks to model the input,.

USAD: Universal Speech and Audio Representation via Distillation Masked modeling duo: Learning representations by encouraging both networks to model the input,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.423539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.493746Z digest=sha256:bffc65f5383de5ab6c1f0908c56494c639a96db92e4c43ab05e890aed401f464

Observation 349d3f1f-58f9-4cf5-b2b9-7b3a0324ee33 · outbound

This paper cites DistilHuBERT: Speech rep- resentation learning by layer-wise distillation of hidden-unit bert,.

USAD: Universal Speech and Audio Representation via Distillation DistilHuBERT: Speech rep- resentation learning by layer-wise distillation of hidden-unit bert,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.310082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.560928Z digest=sha256:36feddbf881ca0925d35ae0f11c9e98228c34c758e8d9a4548e3a3a3f5008cd4

Observation 13ca0c5e-f154-42ec-af26-02292cf4c903 · outbound

This paper cites Dphubert: Joint dis- tillation and pruning of self-supervised speech models,.

USAD: Universal Speech and Audio Representation via Distillation Dphubert: Joint dis- tillation and pruning of self-supervised speech models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.214849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.625963Z digest=sha256:23535bb03af6bba6c92165671260c0a549f221c2286042339cc094446a264a9e

Observation 8ce43848-a607-4024-8e75-8d684c6286a3 · outbound

This paper cites Dass: Distilled audio state space models are stronger and more duration- scalable learners,.

USAD: Universal Speech and Audio Representation via Distillation Dass: Distilled audio state space models are stronger and more duration- scalable learners,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:24.692213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:24.692213Z digest=sha256:db131ccc1259817890618574ab88d1dc51a541e34057f168cfb71587aee3b85c

Observation 03653dc8-1f29-4cfd-b528-f408fdcac4a2 · outbound

This paper cites Ensemble knowledge distillation of self- supervised speech models,.

USAD: Universal Speech and Audio Representation via Distillation Ensemble knowledge distillation of self- supervised speech models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.064742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.808046Z digest=sha256:e9f30ab564f1e4e3dd11463d1c02b672801082662e3af4ed02cb7d2daa6c9f62

Observation 71239d75-3b5b-4d83-bd53-4739d01a2d80 · outbound

This paper cites Distilling a speech and music encoder with task arithmetic,.

USAD: Universal Speech and Audio Representation via Distillation Distilling a speech and music encoder with task arithmetic,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.875657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.906549Z digest=sha256:35189f7bbff37bd3ffb9df7703030bb658d47df3143db918daf86927ce0130bc

Observation d0d6e2be-b8f1-4a92-a866-ce4130a97cdd · outbound

This paper cites Attention is all you need,.

USAD: Universal Speech and Audio Representation via Distillation Attention is all you need,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.686734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:24.997035Z digest=sha256:d238b15461adab963d9f1a1633675aff09a31d27dda92cc24e29cade9b005917

Observation 72a0f1f2-8997-4bd1-9b7f-1fa8c5e77264 · outbound

This paper cites Layer-wise analysis of a self- supervised speech representation model,.

USAD: Universal Speech and Audio Representation via Distillation Layer-wise analysis of a self- supervised speech representation model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.422581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.069275Z digest=sha256:98de49fa8946f98e9026a54a2b3379676d5c502e3aeb78550eb8295ab149f033

Observation 29cf1ab6-5aec-43ce-ae48-0e3f82afb655 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

USAD: Universal Speech and Audio Representation via Distillation Robust speech recognition via large-scale weak super- vision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:25.144389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:25.144389Z digest=sha256:d3a3c3be403174db409c73a1c63633a2d97fa4bbe4b03fbb906b7c00b7576c8e

Observation aa963bdb-c1a6-4c9b-b94b-0cc331f22410 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

USAD: Universal Speech and Audio Representation via Distillation Librispeech: An ASR corpus based on public domain audio books,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.306828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.242305Z digest=sha256:2859e2dceaaacf93451e9d0529b8c5d3dd4f943fc6bcff7c69fd7f4ac323e5c9

Observation 5b88cdb2-f84d-4909-a1e6-5f15722888ad · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

USAD: Universal Speech and Audio Representation via Distillation Libri-light: A benchmark for asr with limited or no supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.157667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.364102Z digest=sha256:6f1ca8f022a683fa81382a4570bc4ffcdc3e39883681c17512a8f3278e1c6202

Observation c24b73c5-4638-4bc6-8cc4-095b73dbadd2 · outbound

This paper cites Mls: A large-scale multilingual dataset for speech research,.

USAD: Universal Speech and Audio Representation via Distillation Mls: A large-scale multilingual dataset for speech research,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.984380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.477593Z digest=sha256:0b057173b1f24ec9e8b635565398a9b49b2cfc6ea46ac7f155928d545debca2d

Observation 70e7d4af-fe2e-449f-ab09-05f17da6fb30 · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

USAD: Universal Speech and Audio Representation via Distillation V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.730479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.592409Z digest=sha256:9341cfeb2d2ad7fb2db82e86971b69baebac326f44f581e354c879a17e9bd5ac

Observation d3bd4328-5d1c-4bef-9692-8674e551f07c · outbound

This paper cites Gigaspeech: An evolving, multi- domain asr corpus with 10,000 hours of transcribed audio,.

USAD: Universal Speech and Audio Representation via Distillation Gigaspeech: An evolving, multi- domain asr corpus with 10,000 hours of transcribed audio,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.469585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.710394Z digest=sha256:387627ba9c3a4287c506398868855101c6d9ab164dbf173f12ce70b9981644bc

Observation 7a3a4c6c-d745-42b7-8974-4864d027eae2 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

USAD: Universal Speech and Audio Representation via Distillation Common voice: A massively-multilingual speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.309280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.805425Z digest=sha256:fe0562aefe4a5e8741186b935a2df7ef0b6b9561c22646c584b27e9fdb203757

Observation 3c27e679-e562-4d1b-a244-f98e5a118880 · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

USAD: Universal Speech and Audio Representation via Distillation The fisher corpus: A resource for the next generations of speech-to-text

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.137544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:25.937214Z digest=sha256:5e3f9140004dc42cd9762543aa0d8e0f1fee13f0cb6b6b81127a3037776755e3

Observation 287d4539-4ada-4f48-945e-e56defe31875 · outbound

This paper cites V oxlingua107: a dataset for spoken language recognition,.

USAD: Universal Speech and Audio Representation via Distillation V oxlingua107: a dataset for spoken language recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.955667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.051603Z digest=sha256:c259dec24e1101a804cc7dd37a471cec8671654c8943d40dee679d1a289b0117

Observation ec99ff25-39c1-479f-af6c-f5297ae458f4 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

USAD: Universal Speech and Audio Representation via Distillation Audio set: An ontology and human-labeled dataset for audio events,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:26.141172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:26.141172Z digest=sha256:a65e44a1133a1539aedee6d143c311576c459d81a914f6fb66c76a6e5d20f21e

Observation 5bfdbea0-dfaa-491d-b550-6c3049143255 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video,.

USAD: Universal Speech and Audio Representation via Distillation Soundnet: Learning sound representations from unlabeled video,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.794265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.187046Z digest=sha256:e6ac1138b5deee2cc12a2933cefbe0276475c950d3d323c37ab77ff3460755e0

Observation cde9ede7-e5e0-4ccd-87d9-c4566b585fc0 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

USAD: Universal Speech and Audio Representation via Distillation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:26.234805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:26.234805Z digest=sha256:a536363418fa20cfdd0510c9d6bea6f94bce60af7713e05d051c5afc5270aaac

Observation 10effc48-0b51-4139-aa79-29293a9bee66 · outbound

This paper cites Music4all: A new music database and its applications,.

USAD: Universal Speech and Audio Representation via Distillation Music4all: A new music database and its applications,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.652831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.272726Z digest=sha256:0dd9598183a6de026339c5a0e1f0bb99a5cfc1e1481c217f85ca066dc3caf0a6

Observation 38c9d75c-da90-4391-b0db-e86f3fa19f13 · outbound

This paper cites fairseq: A fast, extensible toolkit for sequence modeling,.

USAD: Universal Speech and Audio Representation via Distillation fairseq: A fast, extensible toolkit for sequence modeling,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.543437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.352239Z digest=sha256:f02167b698deaa782944e4266adb34107acd392f8816d4ac385d1bcb727c8684

Observation f5a81b72-a1be-4027-a593-5ac2b613813c · outbound

This paper cites Layer normalization,.

USAD: Universal Speech and Audio Representation via Distillation Layer normalization,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.411742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.409954Z digest=sha256:01503faa275960bbe002a1c1a506dc353510fb8c73c2c87bf8f8ed25a1c87868

Observation b6dc93b6-4b9d-40c9-afe1-92e3732586c1 · outbound

This paper cites Self-attention with relative position representations,.

USAD: Universal Speech and Audio Representation via Distillation Self-attention with relative position representations,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.264759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.471051Z digest=sha256:f1edee50b4a1f809981da586b85429d25f7e4dab3a67572afb948433860edf44

Observation 8f1808d9-5252-4f93-bbdd-6ed5c959396b · outbound

This paper cites Speech commands: A dataset for limited-vocabulary speech recognition,.

USAD: Universal Speech and Audio Representation via Distillation Speech commands: A dataset for limited-vocabulary speech recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.140213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.529521Z digest=sha256:02f1767a3cea9d9746c32a9859b04a5e92dded9e23feacdabe2a839c97757c6b

Observation b16ec7cc-281d-45e9-b90e-a5bac8a52c51 · outbound

This paper cites SUPERB: Speech processing universal performance benchmark,.

USAD: Universal Speech and Audio Representation via Distillation SUPERB: Speech processing universal performance benchmark,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.959384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.600823Z digest=sha256:3e9f168ed114a3e72c1f6913d1eea223d5b5c3b525bce5467010577b4957c7b6

Observation 4926ca56-4405-4134-acab-f5b787454b07 · outbound

This paper cites SUPERB-SG: Enhanced speech processing universal PERformance benchmark for semantic and generative capabilities,.

USAD: Universal Speech and Audio Representation via Distillation SUPERB-SG: Enhanced speech processing universal PERformance benchmark for semantic and generative capabilities,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.723306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.658300Z digest=sha256:4567d95731b3fc210ba845ccec4bd75f46a16303b5d9b83476dbd7cf09c8be68

Observation eb20a4d8-dfbf-478c-834c-61c39d3e8eea · outbound

This paper cites A large-scale evaluation of speech foundation models,.

USAD: Universal Speech and Audio Representation via Distillation A large-scale evaluation of speech foundation models,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.603002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.722253Z digest=sha256:4580a143b0f3c181d0d02d91a2eccc9d87442751cfe269f0a9766353d7f95c22

Observation 3ef690ab-d0ea-4524-b326-34161c0a3c03 · outbound

This paper cites Hear: Holistic evaluation of audio representations,.

USAD: Universal Speech and Audio Representation via Distillation Hear: Holistic evaluation of audio representations,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.479424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.790110Z digest=sha256:514386855ede2878c2daec4a44586c8dd10e6d848ea98b4b1591b8250bd48cbd

Observation 20dcce66-1464-49a9-a821-5d5ab6a1bc5f · outbound

This paper cites ESC: Dataset for Environmental Sound Classification,.

USAD: Universal Speech and Audio Representation via Distillation ESC: Dataset for Environmental Sound Classification,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.325179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.882696Z digest=sha256:b11c5afc507d630dcb91f18bf67c5ecd69645a18bd08d369951d68f0280ab712

Observation 1c975ef0-3fe5-43a6-bd22-1e4d3b6f7699 · outbound

This paper cites Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,.

USAD: Universal Speech and Audio Representation via Distillation Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.152485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:20:26.977695Z digest=sha256:d88c9e3fbca69f5ef23701af63c331a75ed2f55ffafd5acb36a7ff795ec2b21e

Pith citing papers

Observation 509a8141-b875-4dfd-a562-9bf5547bf3f3 · inbound

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations cites this paper.

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations USAD: Universal Speech and Audio Representation via Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T07:27:05.187538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:27:05.187538Z digest=sha256:2ecb7f2666642884cfdafce36699e2327072e6bb04a144a865327060df4d6ce3

Observation b3c39e7d-ebbd-417a-bfa5-7ab57dcca0a4 · inbound

Alethia: A Foundational Encoder for Voice Deepfakes cites this paper.

Alethia: A Foundational Encoder for Voice Deepfakes USAD: Universal Speech and Audio Representation via Distillation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:33.721376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T19:27:59.124425Z digest=sha256:877f1d06c0009de79db0e26f82b16c046a728a77833c32dc950cf7ef853d71e1

Observation adc33ca8-be64-4795-a30f-20f8f644cdb5 · inbound

Stage-adaptive audio diffusion modeling cites this paper.

Stage-adaptive audio diffusion modeling USAD: Universal Speech and Audio Representation via Distillation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:06.248596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:21:34.140699Z digest=sha256:b557289c218cdb9996da757cf6bee1c0f206a8e452a6aa0db411f390037c26f1