Pith. sign in

Paper Citation Record · LEDGER

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2106.07447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.07447 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:27:22.494754Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dc48ff5f-8d49-4f63-8941-bf4181efd1bc · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.494754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.494754Z digest=sha256:28bca761e709ec953727dd8d44b755ed1af54cb35aa5c38caf9c0909aecae211

Observation 3efc365d-704a-4714-8d27-54e70b726133 · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.191969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.191969Z digest=sha256:128752642b91401d8d450f5fb0fa236740bf04b7d812ffba54e2b1af8c2cebdb

Observation c86a60e8-64fa-429d-b203-491345f93955 · inbound

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages cites this paper.

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:04.453086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:04.453086Z digest=sha256:fde0a4435112f7ae8643727e4a8282e701dd80696519d6b7913e579ee4ab3804

Observation b907b8c8-5dd0-4009-a434-2028a84c0a45 · inbound

Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech cites this paper.

Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3460

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:36.292242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:36.292242Z digest=sha256:d59a8559fede37f77b1e62a4e2ce5985e2e3af0fce1fbff54586814294eec57f

Observation 2757860d-2e1a-4049-800c-f969eeb485a8 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.569844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.569844Z digest=sha256:d706aaa913dccc0a75938203af4020f60611454bd1a8f2bf1607b9cf6628ff9e

Observation 762248e0-df10-4ea3-ac8f-a499ded69049 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.040172Z digest=sha256:a844eaf5c23295f8bb9fc4c2ff99ba6e02380ad9a6429d35d5b5d9554e8991dd

Observation ca1dcac3-5898-4a56-8ae9-a51f4ae5a008 · inbound

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance cites this paper.

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:11.289042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:11.289042Z digest=sha256:3e073bdb9d87546968daff92293daf2cc1ff590f4cbd2d55c5e97bad60ee8493

Observation 2f1a916e-e772-4f3d-8282-a493595b354f · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.224028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.224028Z digest=sha256:95d100a868d003a474513f6a8961cfd8bf3e601f91aaf27c115fbf65c819003d

Observation ff90399d-1651-42e3-97fb-5e6f8bcc7037 · inbound

Self-supervised learning of speech representations with Dutch archival data cites this paper.

Self-supervised learning of speech representations with Dutch archival data HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.355340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.355340Z digest=sha256:e11f67c73987454d69ce5bef3e735bd7f0f439f94ff0df6253ab0b51dd2629b0

Observation e26faa88-32ee-45c7-8e30-07a804b5bb7d · inbound

Leveraging Context for Multimodal Fallacy Classification in Political Debates cites this paper.

Leveraging Context for Multimodal Fallacy Classification in Political Debates HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:18.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:31:18.775855Z digest=sha256:37e3ae270ad474026f63c877ff489f8ec772f72a9b13c6274e70e886913b03e1

Observation 986ec86d-5bf2-489e-a7fd-de4c8516c67d · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:59:51.132301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:8bbd98f171b00a9c0121732ab85e1e23b4379a04f873516e4af53b5ce7d4e43a

Observation b7d3a70e-6d86-4d32-b58b-f91ce0aac5bf · inbound

A Concept-based approach to Voice Disorder Detection cites this paper.

A Concept-based approach to Voice Disorder Detection HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:52.253608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:49:52.253608Z digest=sha256:adeb8ccfb985c474c75bdf5ae8effc64a8e837479a1e9e6d349e55d82bbeea6d

Observation d20f6562-f71c-4c16-8d5f-bdf6236aaf4f · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.485256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.485256Z digest=sha256:d1f8c60d5ee42c8bd0499df0cc60b24f18b2eff5a6c9253c6cc485d43ea975ab

Observation c9954816-dbc8-43e3-9707-98f3ef9225ec · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:04.730722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:04.730722Z digest=sha256:b5a72c61cc50325b96ffed01d3a23d55992317604ee8e62916c082e016cd191b

Observation a0f2b216-8a76-4676-9c08-b3ca6222aa23 · inbound

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models cites this paper.

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:22.832912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:22.832912Z digest=sha256:c79a8057cabc92bcef3131ba78ca13a1069e7f5185fdddcd5f291db3694fae1c

Observation 14a60961-3961-41d0-a5f4-b6cea4c81c02 · inbound

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition cites this paper.

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:39:59.378823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T11:35:38.330061Z digest=sha256:a22ef08d2da79259a9a29ed06280b0e9ba455496556fc8e2404f2f8738e1a57e

Observation baacf565-a834-462f-bf1b-b9e2c8474f3a · inbound

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding cites this paper.

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T17:37:12.659819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:37:12.659819Z digest=sha256:9d7c8c3baa0667ed0e4a2b52eb82742189945315ffa419e968b9b06efe9951a3

Observation f7728530-8560-4a6b-9fcb-4bc70517722d · inbound

Neural networks for Text-to-Speech evaluation cites this paper.

Neural networks for Text-to-Speech evaluation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:49:54.655165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T09:46:26.884551Z digest=sha256:8c8cd0990519d654d456678c5035bbb264481346e7aa8757b509067ce2caf074

Observation b1c58a5a-15e7-47fa-849e-a00281d74a84 · inbound

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology cites this paper.

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:43.777162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:02:23.847187Z digest=sha256:7dba2a969a02c88a4b7014185dd81c880b898d1398bdf1ec368a8a24bfe6cfe5

Observation 590cc051-77a2-40c2-9753-b2bdf1d71bc4 · inbound

Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG cites this paper.

Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T06:53:06.014414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T06:48:52.078640Z digest=sha256:bd8e9ec8d677b56fda2761bc3d0470a0567afeb808e86a45013f3c7503a6cd04

Observation 27588fde-9bd5-4255-a614-0f6e73963047 · inbound

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation cites this paper.

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:47:06.020988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T23:36:27.369551Z digest=sha256:3dec6456a202a8b83f52326129e3c8f3a35b9e01b5a6c89d31fd43e2b5e1ec90

Observation fcef0e10-8117-47e2-959c-c7bf9d3907e5 · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:07:56.083448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:51f8abccb753c075556c413dbcdce3f08c3142c4caab8f90dc7a8a648b318e67

Observation 46238d6e-7817-465c-bc14-d6371a8c32b9 · inbound

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal cites this paper.

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.223989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T17:37:25.043607Z digest=sha256:ad977db90718a3cd21cd17b4308bd4777c0b01425bcc060c0f0e11412437b52e

Observation 4cba52db-200a-4e08-9c0a-e31d1b391a8c · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:42.891310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:81950927062b1335bb2afde56808e9e28f3b49fb90eaba852a10184c4a63cdf3

Observation 12d732ad-24e4-49be-9020-97b89026ede6 · inbound

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users cites this paper.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:38.221570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T13:30:12.101045Z digest=sha256:d1a6cd1c2ec727986d3bce9cff34610ae5c4163ba0b14d91fd080621bf6c31cf

Observation 05e6240c-cfc1-42e9-b23f-5ac9f84ef1b8 · inbound

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty cites this paper.

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T04:38:58.619776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T04:38:01.183423Z digest=sha256:c11b6c848f9a6b00c94ba2d933f04ed8ad2f31845f01c975a47702a1d42aa562

Observation 28de5576-2bb0-49ca-9446-7276fb253d47 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:09a6e3413495aeed000dd58b01c13ea35eb6f0210c7492b89a022a8520a939d6

Observation 2e838a15-a1a9-403d-9d33-4fc1b0085c1f · inbound

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment cites this paper.

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:46.343211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:46.343211Z digest=sha256:749df826306da0d6fb894e236b2ff74471ecc0d76d853c26c7bd5719be5d54db

Observation 3df35c21-9c57-4186-97af-02c2123f6fe0 · inbound

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild cites this paper.

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T23:56:38.448201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T23:51:55.422232Z digest=sha256:07bf50ddca68b6a0e9269021e7c02e1de96a0ad26f8cead3daa62682f4d15b8b

Observation 0c6ac7ef-ce87-48c6-acb2-7e2bfde25cf9 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.830211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.830211Z digest=sha256:fd7ea8c755e969b7f811f9a7ad7dcc05b8d3ff9035be6b7cb2e193d5042b9be2