Pith. sign in

Paper Citation Record · LEDGER

W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2108.06209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.06209 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:40.662723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T01:44:22.720612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bfc93eb1-b0b3-4c43-b320-f767692700d2 · inbound

Different Speech Translation Models Encode and Translate Speaker Gender Differently cites this paper.

Different Speech Translation Models Encode and Translate Speaker Gender Differently W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:40.662723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:40.662723Z digest=sha256:124f8337ccbae1b51a6680ac184df161ecd8c54d3a57813ace54ac1f7a9383f8

Observation 9516fb69-478a-4798-bfa4-95e9d6cda6bf · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.599394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.599394Z digest=sha256:85b85b91e4ec98d937ca973c7956a655a52cfd097dee3a40e86e01a28dd3576a

Observation 0cfb0907-fbf6-4a81-bc92-e6632aba70a2 · inbound

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion cites this paper.

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:31.283961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:44:31.283961Z digest=sha256:d8ba9cb722d8e464283a2a327f6d1ee99aedd258d9525fb5e50b1d135820e883

Observation 4a60ab31-8328-404a-8836-00cfa8260a5a · inbound

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model cites this paper.

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:44:22.724239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T01:44:20.282896Z digest=sha256:d969bb2fe671fde9ef89e1dda4b1d268a7777e221b7903b08302f55fa5a4cb02

Observation a6c55c28-7e24-44f6-b799-e29e7ae88445 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.822866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.822866Z digest=sha256:d004778c2154337e77a331b84e5a3173e3942b416766181d033658e6fcbd027e