Pith. sign in

Paper Citation Record · LEDGER

Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2104.03502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.03502 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:07.451790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:34:43.114766Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73c8657a-c5c1-4469-9563-4549b887a009 · inbound

Once More, With Feeling: Measuring Emotion of Acting Performances in Contemporary American Film cites this paper.

Once More, With Feeling: Measuring Emotion of Acting Performances in Contemporary American Film Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:08:42.895117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:08:42.895117Z digest=sha256:d399087b857e17a1b0d3f319bd240a68a3b315590333b462baeef67cffdbdea7

Observation 14e7e7f4-71bc-43d2-b84b-6a49421b81c4 · inbound

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer? cites this paper.

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer? Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:25.777825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:43:25.777825Z digest=sha256:039721abe5ea9b11a5a11b9479fe0ca728474c15d84f725dcc5825f9e69260bd

Observation fa02346d-452d-47ac-91a8-54bdf61fe6c4 · inbound

Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition cites this paper.

Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:46.274112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:03:46.274112Z digest=sha256:d7873dcc598ef6ad32c1a2e388a85468b5893ac2af54b011ee0cdfb33160ddcb

Observation 42e07762-53d0-40bd-9304-00a22760f0bc · inbound

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities cites this paper.

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:42.651816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:42.651816Z digest=sha256:55b8b938ec6213b9cf6f8be6f56917e9823c44c69e1b79a778bdc0fa3b669595

Observation 6840b528-b5f6-4b71-a05d-fc00a2c0dff3 · inbound

Synthetic Audio Helps for Cognitive State Tasks cites this paper.

Synthetic Audio Helps for Cognitive State Tasks Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:41:51.174050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:41:51.174050Z digest=sha256:28ee2c405db399673e73847b0a1e97f8b66a00f5caf3a8b0a97b94b4b513ed60

Observation a71d5443-d719-4183-a034-43b3a7a0effe · inbound

Investigating the Impact of Word Informativeness on Speech Emotion Recognition cites this paper.

Investigating the Impact of Word Informativeness on Speech Emotion Recognition Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:15.791716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:15.791716Z digest=sha256:5a1cdb26f70bd6b8be13709fd92a44058f6a9a01a4c6f23a8b35b6707edcbf98

Observation f15749ac-0c24-496d-aae4-a5f7766a748f · inbound

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? cites this paper.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.459509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.459509Z digest=sha256:8816e164b8ea09fe431c9996cde3eb2625e0feead4975b8686a7a04e71f48349

Observation 87888b61-f0a7-4ea2-8609-c1f8381350d5 · inbound

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews cites this paper.

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:39.878332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:39.878332Z digest=sha256:ab81fc9349e5cf7ffafef79b41e0d651705a77978aa49753b9413c21948005ad

Observation 8f82d945-779b-4bf2-bb6a-5dbef9fe54ec · inbound

Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond cites this paper.

Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:23.587518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:23.587518Z digest=sha256:281e102d491b8982d701bcbe04b239dff03e19170967c68a8ec10366818606c6

Observation 07a9346d-c2d2-4a23-8987-7b6918b937c0 · inbound

Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection cites this paper.

Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:07.451790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:29:07.451790Z digest=sha256:b565969682c4896dfbdbf0a843db0e07baffa3575dd0410061698d527c77fb7d

Observation fedc5dc6-d724-47b7-b347-6f96fc92f8c8 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.317218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:1c5cca904bda1624d51ed0ce93245178c62a4bb3cd8859fdb785d525490f95c5

Observation 290e3121-0c63-41c9-93dc-8c059de5046f · inbound

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction cites this paper.

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.849407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T07:58:49.423628Z digest=sha256:124e661e09342e601ca520ab96f35f509ef3b231dd6c85afd5a2be10bce023ab

Observation 1f25fb35-91ca-4465-b66b-7067feb29066 · inbound

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures cites this paper.

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.082710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T11:14:20.892333Z digest=sha256:e4a0de7f63638fe046b95d5f9254272528a681236c679e38d2491285415a1d03

Observation 6c93ab97-6a5e-4c93-bfe7-0cc20f230153 · inbound

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition cites this paper.

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:57.691826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T04:44:59.279562Z digest=sha256:c5ed2e8919ae4187fa11a187d609142b9124c93c617ff71147c47117f2a4c046

Observation 1c122ce8-1fa7-490d-8cb0-98cddfc424c3 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:ff0b1acc9ab36eb832b54f4dbc9345f41fe6930d7163a70a0b852ef66b761351

Observation 3677867c-4038-4993-976c-4a15694b23b4 · inbound

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment cites this paper.

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:46.343211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:46.343211Z digest=sha256:afb8b1a4a009c2da4262b8a590034ae2d54ac3e2a29e769edc3f730a32ec28b4

Observation d94f2228-de5b-4331-8f5c-64da30fd43e3 · inbound

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective cites this paper.

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:34:43.117158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T07:33:22.898615Z digest=sha256:45d7f9110ee01d373b2ada0375a9b48daaf29450a99e26e5ee40060eefab6f12