Pith. sign in

Paper Citation Record · LEDGER

wav2vec: Unsupervised Pre-training for Speech Recognition

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 63 inbound Pith citation observations for arXiv:1904.05862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1904.05862 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:02:44.973306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:30:08.546246Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ce798540-d1b0-4cbd-b4b1-1e8c5ba8d8f6 · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.250772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:0a10dd05631437f8e533979512c93ae6e43e62dd54dbd888422df07a1943b32a

Observation 1eeae4cf-6622-43e5-90c0-7fcf9e1b1aad · inbound

Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models cites this paper.

Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:31:53.243382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:31:53.243382Z digest=sha256:e26e1bcbd66e7a5f0ad0391ffa766cd52558ef74cae112f2275928b7fd4227fb

Observation c3b75942-03a6-4b9c-bd97-87e0ddbf016f · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 185

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.978189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.978189Z digest=sha256:a6b8e4d4046ccf42a19a5acfb50927571e392a6bdd590ce4c0873ce1f321302c

Observation 4a5bae22-7b35-430a-8a8a-6e90caf29058 · inbound

Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network cites this paper.

Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:49.936349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:49.936349Z digest=sha256:fbf3975457ec58347b708adef9572ec5b5c61f8286a77fca7c19b8ca05e5e6dc

Observation a7dd5953-4b0f-491f-82d1-bc933439413b · inbound

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation cites this paper.

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:15.328188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T17:19:51.411937Z digest=sha256:db742305432d686067cbb19597f740b0592e109a98cc15c1912f8fb633d35be8

Observation 3bf2902f-54f8-42b6-861d-8660cf0724d5 · inbound

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer cites this paper.

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:06:43.857045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:06:43.857045Z digest=sha256:d42ac68df3922f6d7d0c5c00c49d227b7279a7eafc21434dfaad9ffcf5f56d3e

Observation ae92db97-c9e1-43cd-a818-8e7bd89ca952 · inbound

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction cites this paper.

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:49:45.797865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:49:45.797865Z digest=sha256:36c0e299bff9cbdcde0096998b2a86ea9447c068e4fb381ffae05d5295324569

Observation 6b6dc154-21c1-4a94-a3e5-ac196ba4ef52 · inbound

Real-time One-Step Diffusion-based Expressive Portrait Videos Generation cites this paper.

Real-time One-Step Diffusion-based Expressive Portrait Videos Generation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:59.657839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:59.657839Z digest=sha256:61406679eb7a3cb0bf2875fbec37b18d1fca7a27fd99b70a960d42495808fbd9

Observation 9c221058-7543-4c6a-859a-f85ec72bcbe4 · inbound

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation cites this paper.

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:23:50.654728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:23:50.654728Z digest=sha256:20fa47351aca8c1eee80aff31a66c02f0d3f6c878c2acbeedc810e5323320319

Observation dcc3f86b-bba6-49b1-bb64-2770c4ced7f2 · inbound

Optimizing Speech Multi-View Feature Fusion through Conditional Computation cites this paper.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.967080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.967080Z digest=sha256:148cbb94276eee7588f778857f66548145c84c9a4b6378fc66659df17b37352a

Observation 5300cfb3-f06b-46d5-a3b3-1c0f4050d361 · inbound

Tessellated Linear Model for Age Prediction from Voice cites this paper.

Tessellated Linear Model for Age Prediction from Voice wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:19.839014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:19.839014Z digest=sha256:590976cbe126d5a25bd7063127e66b7e0dcf95b198c16e78828216d53f5f638b

Observation a2db48f8-ec86-40b7-834a-c5a81a49ed46 · inbound

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition cites this paper.

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:48:06.562996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:48:06.562996Z digest=sha256:f9c09f05dc3f824bd21914f43e36b55656b3e557295578455b30aa8fbfe1d3da

Observation e069ab83-3241-4b8a-bd64-766445ddf959 · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.620790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.620790Z digest=sha256:f5c64b14282f524805b7b09fb3425dc46f150bd05bf02004e0a718d2303fdc0c

Observation cb4deb4a-4ded-4ddd-91a1-20d164e9dcc0 · inbound

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models cites this paper.

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T16:47:13.015378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:47:13.015378Z digest=sha256:c59e002f3a93ad2eea4110fe6361b17759486d8d647512546b50a4935dc79bd8

Observation e04c8143-d829-4d61-995b-dd344e649c4c · inbound

Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends cites this paper.

Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:26:13.931617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:26:13.931617Z digest=sha256:387c8ffeb200665c5c42436f59e9abaf206aeeb00ab8d880db31963dad85fb95

Observation 3e4e6b38-1910-4a57-a1ed-bb0b929aa685 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:07.994420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:07.994420Z digest=sha256:af7f6362a308fa05de4190a6f91fa4153324410e0d7343ebcb14ac8edfe6ef23

Observation 1a6f2899-cca5-46c4-9a03-c33486018c95 · inbound

Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning cites this paper.

Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:02:44.973306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:02:44.973306Z digest=sha256:8d0675049a2a0a21ca5c009766de78401a988eeccec267232282a3d8e3c80769

Observation 3fbe7391-59f9-4613-81fb-6e3e9b2623df · inbound

Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning cites this paper.

Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:06:50.295833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:06:50.295833Z digest=sha256:066816fb90ba003d41429c1909fdf9a0481b2cb01d0ae0191d6bf9897ccdba99

Observation 780d4055-048f-4c61-9f8a-fdb0d592921b · inbound

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play cites this paper.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.666078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.666078Z digest=sha256:ebc2b1ea38083adba95573d8c11f584aaac41542c37affdb8aaab047f141b202

Observation ec376700-f940-4ad6-bb8d-14b1b91cd02e · inbound

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation cites this paper.

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T23:53:10.585147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:53:10.585147Z digest=sha256:9efacacd0b6cecf377f7ff104489c7ebf7b0c493abfe2870238f639e53a5133e

Observation 6c47dc4f-ccf0-487f-9318-a04e7d3f7167 · inbound

Model as Loss: A Self-Consistent Training Paradigm cites this paper.

Model as Loss: A Self-Consistent Training Paradigm wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:40.142844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:40.142844Z digest=sha256:50da5bb99cfcf3dc57313fe0eabc4e22197a7d20b9fdb0ad13e9a3ce4434fd71

Observation d3a5ef45-374b-4484-983e-4c3fa9ffc8f5 · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:31.128176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:31.128176Z digest=sha256:a3d362d5a02e5e2b8ab96828d7897a1677bc64577475350854a9d1e6c28946b5

Observation f6496d82-3799-4a4e-b019-ed3142d7a88e · inbound

Automatic classification of stop realisation with wav2vec2.0 cites this paper.

Automatic classification of stop realisation with wav2vec2.0 wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.747657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.747657Z digest=sha256:247c78a64dc4bcd505541966ac6be7dcb6dda5cd98c96bbf49f1dc28e4752c27

Observation e7f091b9-a0e6-4a24-993f-f5a1d4d7890b · inbound

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec cites this paper.

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:39.796888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:39.796888Z digest=sha256:762732dd3126f90a114f4eb3e12ba2f1bcd73203eca84d7da8cbcb0b4271c583

Observation 02691a0c-79b2-460b-9721-f69bf7e95692 · inbound

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction cites this paper.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.618021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.618021Z digest=sha256:e84a291587485df5ded1bfd19ace6c96a61be16adbaa849fdcc5ab3a8453d5dd

Observation 0cd0cc7e-cfb5-4cac-8002-73ea19f1c9e0 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.988059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.988059Z digest=sha256:6291cb41488d4169eb4669415d8094ed9fd5db4ec63c711f8da9df7a765a69e2

Observation bbba3595-c4d5-4e38-9e09-6c455e5d1628 · inbound

Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning cites this paper.

Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.946181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:04.946181Z digest=sha256:d01aaa2be2664d857b4770d98adf54a73857caa95d06fd264ea2a25453b6690e

Observation eded0f9a-9741-4aa3-b944-1b04c6b765ce · inbound

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques cites this paper.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.191276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.191276Z digest=sha256:3ffbadb18164355c932becf34341a8a2a7fc0c32685ef9196f468ea10975064e

Observation fbf8f96b-d270-49c2-93b7-92681762c5ff · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.115489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.115489Z digest=sha256:99da684540e31d76936bf54c237793aba3b5a1f59de29d99d92fbfcd2adae399

Observation ad2f536e-d363-4745-ab5f-ffb2d3fc48e7 · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:14.039893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:14.039893Z digest=sha256:a2d5f3c44b29ba1c80b23a5e2d32fdb74a2b606b4bd98685e877fc3f102985cc

Observation b41b945c-92e7-4dce-9c77-ab82425a6e3f · inbound

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis cites this paper.

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.003341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:09.003341Z digest=sha256:0a8238f9cca2e4958e2207f148b1e1944bffe59649dd2dda27a2bfcf8638b8bc

Observation 4311960b-b606-41e4-a1f4-1bd4c7ea2419 · inbound

Manipulated Regions Localization For Partially Deepfake Audio: A Survey cites this paper.

Manipulated Regions Localization For Partially Deepfake Audio: A Survey wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:56:32.139392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:56:32.139392Z digest=sha256:1497b565d0633cb71386e9dea0d13af6211dc1a94548d70b4de205abed1b3fce

Observation 0d44c9b3-5ea3-4053-be17-e2763b175e0f · inbound

A Dataset for Automatic Assessment of TTS Quality in Spanish cites this paper.

A Dataset for Automatic Assessment of TTS Quality in Spanish wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:26.540877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:26.540877Z digest=sha256:406a65c4d3d9f5f0d34a884b0ea14dcfbe3237efe7a772ad7ceecfb20904779b

Observation 568f8819-1c0c-4d3c-8df3-6da741a4fd4e · inbound

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation cites this paper.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:42.919123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:42.919123Z digest=sha256:a6689b8fb2882b4e2a208f5c491a4c3f7047094054b22b225df54ef67e3a8e2b

Observation f33d1aa8-4a70-49d7-b9a9-a28f852fa65f · inbound

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder cites this paper.

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:13:17.511657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:13:17.511657Z digest=sha256:afb5c5c9470f1224c490ba6754fd382b816f0a925685fded1f66a2e6b1ea392f

Observation 827e119e-ffb4-4984-a5d8-e2dd6d862668 · inbound

Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems cites this paper.

Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:05.733703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:05.733703Z digest=sha256:d4f1b39d49cd418802ae10708a2d62bb040500c54da730381be97c8a78754e43

Observation df222587-a596-4f16-9829-02dffc8e538c · inbound

Scaling and Distilling Transformer Models for sEMG cites this paper.

Scaling and Distilling Transformer Models for sEMG wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:28:56.354336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:28:56.354336Z digest=sha256:ad83425b4b7a2933fdcd80300b174842ebdccef326287123d15834ff02c25cd7

Observation c13b9691-63b7-4200-b309-67be32579770 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.266125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.266125Z digest=sha256:142a69106500628e88430b4cbf5a72fcd886835f15a084e9666c394e3e84652e

Observation 7f0dfb18-be61-48e4-a07e-b129e4107d22 · inbound

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition cites this paper.

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:01:23.671306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:01:23.671306Z digest=sha256:67721609ce23e3a56b528b7761108df8d83b402f19554099e707575d7036c227

Observation a928eae6-f83b-420b-8800-168703fb3db6 · inbound

DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches cites this paper.

DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T18:31:36.742828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:31:36.742828Z digest=sha256:9966e91c066cb656d8681d0f839b0cb946e824a720dbeeba13ed93314e9a76cb

Observation 58a00d04-42ad-4dd7-a6ad-02046ed1ef14 · inbound

Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition cites this paper.

Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:22.800840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:22.800840Z digest=sha256:8d6b214d78975410ced59a4e2a58b8f9da535cb39180495378654425cfe70602

Observation 9fd4856e-f8aa-4d72-be37-0229cb13e30a · inbound

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning cites this paper.

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:22:20.571864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:22:20.571864Z digest=sha256:a0f1c2c8e32471347788f8617aba9845d6010ce96754d2b3f04451607ed5dd0c

Observation 2df71453-5bde-4389-8e4f-9818f88521cb · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:05.908306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:05.908306Z digest=sha256:d5a7656b7427ed8020be4cb2a44edb81ff169e700f1a6558714188b6deaca371

Observation c737ac1c-8927-4983-8369-8b7f16955cdc · inbound

Contextualized Token Discrimination for Speech Search Query Correction cites this paper.

Contextualized Token Discrimination for Speech Search Query Correction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:12.074420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:16:12.074420Z digest=sha256:56fb0038bf13e577a0fbb33d471400ed161daed779a151b2e3596c57f92911ee

Observation 8696e58f-7177-4cd2-8f24-63c0c5ab6eed · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.850388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.850388Z digest=sha256:5b1a9b914d0d6ca30cfc54589b24a0971a6f82cf1925faa73658cf74e7ca87b1

Observation c3ada0bf-9a1b-410e-9062-db6b3e576a50 · inbound

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent cites this paper.

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:35.635172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:46:15.275441Z digest=sha256:cd816c853d5be85c3479c8e2a94e61a9c52dac62fd8428a5636d839e2eec7136

Observation accda2fa-7009-4344-9ada-51a61be75f66 · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.174342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:5b1b6b676fde363e9037a1ec62c77e1cb188babae7e84d806354c7383e4fab48

Observation 10f075cd-643d-4c24-960f-f04cbcc4213c · inbound

The Indra Representation Hypothesis for Multimodal Alignment cites this paper.

The Indra Representation Hypothesis for Multimodal Alignment wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:48.751340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T20:06:03.531145Z digest=sha256:5e5ef8eedc2e38d8b067b6a869b114f018222b4a79adc4b57fc877a9b9d3f1d8

Observation 130decd9-62cf-4431-958e-298031d78e5f · inbound

Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection cites this paper.

Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:51:00.710066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:46:58.566821Z digest=sha256:e8e0c23795b8b1fac3d773c02a50e41bf30b5b8b7294c710b7d81dde37c3ec0a

Observation f2f67d08-f403-4fa5-a139-9ed8de2bebdd · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.993740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:fe23cb24058c2522205165123545cf8f5619fd80097863ffd0cdfad572826dc2

Observation 6eb36e49-f82f-416e-a613-f46fef0c9cf7 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.677029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:e8e4547a0c62fc02b448aa506ca757ac5d3adf7349d70ddcf3b07a91e1a7b012

Observation 8c0ac0c5-0803-4aec-a953-4a25a3e1dc4e · inbound

EGI: A Multimodal Emotional AI Framework for Enhancing Scrum Master Real-time Self-Awareness cites this paper.

EGI: A Multimodal Emotional AI Framework for Enhancing Scrum Master Real-time Self-Awareness wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.686763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:04:50.082469Z digest=sha256:156eeeda66817424c903fbbb571d5548e58832fd87f5c28460aa9617b0dae53a

Observation ea22beaf-3b1f-4edf-89d9-0d39535ec478 · inbound

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition cites this paper.

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:39:35.091185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T04:37:41.980535Z digest=sha256:02fe4c4f1a8342aeb2b9a04de300a7b7fcbb5e6dc82e4ac0fc22b7a831e2f8bc

Observation 12eb3e77-30c5-4fac-bb4d-63cbaf818c0f · inbound

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing cites this paper.

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.569440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T14:45:45.359616Z digest=sha256:7d8265e4b36adec0a858325724cb0fb2d746a1baaf16d3dc2d4973ea9d228051

Observation a866f1c8-24d9-475d-a7a7-6a4db8a68998 · inbound

Extracting Governing Equations from Latent Dynamics via Multi-View Contrastive Learning cites this paper.

Extracting Governing Equations from Latent Dynamics via Multi-View Contrastive Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:48:21.212460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T07:30:17.092770Z digest=sha256:c2e2c0997c72cd4ad19bf3c7204611ccc1c26e027f98caad2912e863059434e5

Observation 24e94bc1-9898-4567-940a-96cebad7cf17 · inbound

Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming cites this paper.

Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:08.548256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-25T20:00:25.206417Z digest=sha256:a5e343b797a50467132f0bde4bef9a0a50975d01f55c11bfb2411fe5630ff87a

Observation 4e0ca598-ac30-4646-9f0a-4b12dd9ea6cb · inbound

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel cites this paper.

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:49.300052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T05:13:56.626647Z digest=sha256:c607daa7f17f87483929beda82b43ef3d23cba0e0dc8b525d98e232f8f5c099f

Observation 93435d25-db93-479d-8e4f-1ab7fdc1df1a · inbound

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel cites this paper.

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.213982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T10:06:51.628147Z digest=sha256:e9dc8e7ba57c0e0276b5ded02a50baf16b54688829bb87d05e69324302c04ef4

Observation 3bf9f104-4fb6-4c13-928f-d12585603dc3 · inbound

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2 cites this paper.

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2 wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:04:32.922022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T08:55:16.695293Z digest=sha256:55b361853b1d789bdfa8c1ddd5cab11a18cfe65bace47ffd16a229064d891130

Observation 8cf086c1-9165-45f6-8adb-5028be3f62b0 · inbound

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System cites this paper.

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:44:27.383983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T08:41:24.068876Z digest=sha256:4d33513f06f06efa95e667694d96e30a4a25eeb85e0d311c42dc1bdac5a4359c

Observation d26ce9db-c44b-433a-8ffc-a124754bf397 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.839281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:988910b99b27fa19e9dfea0160469a9b7d0566533de48298607441c3522da4a5

Observation ce08860b-31a8-4bb5-8558-5a1c77aefc8b · inbound

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis? cites this paper.

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis? wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:37:21.698795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:37:21.698795Z digest=sha256:69dd6cff87e2ba14e65e987ab0a5c80f99a4117a52e1fc15d6aa06a771f02a88

Observation 3ceb5d73-c69e-4672-9245-0a033c51e0f7 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:10.875081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:10.875081Z digest=sha256:335cb0cffcf55eaa3aeddfbb628fb1ba8b71c53ab359a6879641218d05629a9a