Pith. sign in

Paper Citation Record · LEDGER

Learning Emotion-Invariant Speaker Representations for Speaker Verification

As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.18498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18498 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:14.007487Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:13.912562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:33:14.053965Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d272e83d-db89-464e-a96d-f5c90b0553ec · outbound

This paper cites Currently, SV systems that use low- dimensional speaker representations extracted from deep learning- based speaker encoders have become the dominant approach in this field.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Currently, SV systems that use low- dimensional speaker representations extracted from deep learning- based speaker encoders have become the dominant approach in this field

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.358571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.908650Z digest=sha256:22eefae20e6b7d3770ee35892ce6caafd71dfb9953d52ab141e787a39ebe30a7

Observation 1c53738c-ad4d-4cb8-b7b5-c8af21ec7dfb · outbound

This paper cites Learning Emotion-Invariant Speaker Representations for Speaker Verification.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Learning Emotion-Invariant Speaker Representations for Speaker Verification

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:33:14.058991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.912562Z digest=sha256:90b6f2f2eac48e32199c4c5a17626f9b449d84c104dc5afcf6a3f860168265eb

Observation ef6aa4ac-7b3b-43de-9116-967227f13384 · outbound

This paper cites Datasets Our models are first pre-trained on the V oxCeleb [22] and then fine- tuned on the Dusha [18] dataset to evaluate the performance of SV.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Datasets Our models are first pre-trained on the V oxCeleb [22] and then fine- tuned on the Dusha [18] dataset to evaluate the performance of SV

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.346234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.916593Z digest=sha256:61a4730bd16b26478893802d19c6dc15db1b5dc884e9f9f5bc43993c2a86f1fa

Observation e4b5f3c6-29e9-4647-b9dd-79d75c8cae9a · outbound

This paper cites Performance of the baseline system The performance of the pre-trained model on V oxCeleb1 is presented in Table 3, with the results of the ECAPA model referenced from [3].

Learning Emotion-Invariant Speaker Representations for Speaker Verification Performance of the baseline system The performance of the pre-trained model on V oxCeleb1 is presented in Table 3, with the results of the ECAPA model referenced from [3]

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:33:14.335678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.920331Z digest=sha256:0c66c02a691d4a51459fd7fd11e4e28f1710162ad8c1c81a52131962c9578747

Observation 51d46830-4f89-40c2-9dd7-eb89b07fa94e · outbound

This paper cites We first verified that emotional utterances degrade SV performance, with cross-emotion test trails performed worse than same-emotion test trails.

Learning Emotion-Invariant Speaker Representations for Speaker Verification We first verified that emotional utterances degrade SV performance, with cross-emotion test trails performed worse than same-emotion test trails

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.323742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.923606Z digest=sha256:39ccfabd5cdb2731c25fdcceb9528008a664cd017dd20f089bd326b34cf13db6

Observation e6b4887c-d265-4784-9cd3-313a6928f3c6 · outbound

This paper cites X-vectors: Robust dnn em- beddings for speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification X-vectors: Robust dnn em- beddings for speaker recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.312011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.926783Z digest=sha256:971ce0b65a6cfaf2c4a41014f0bf19d152f5567daa03e18d4732ac80326da9db

Observation 4c2a3cbe-4731-48f2-836e-011dc408225f · outbound

This paper cites BUT System Description to VoxCeleb Speaker Recognition Challenge 2019.

Learning Emotion-Invariant Speaker Representations for Speaker Verification BUT System Description to VoxCeleb Speaker Recognition Challenge 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.929697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.929697Z digest=sha256:80116ac822a56debb30e64b54ae08a0f71a027e1d8b1d3dcbb15cb74e74cef18

Observation acb2c15c-c308-4464-85c9-2b6052663fdb · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.300839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.933411Z digest=sha256:66a5d08bd3381475636c95d3a957e5077d8ba59a07311fdccb048df5ddf4e7f2

Observation 451a8d6f-545d-4d5f-966b-bf4bb031e276 · outbound

This paper cites MFA-Conformer: Multi-scale Feature Aggregation Con- former for Automatic Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification MFA-Conformer: Multi-scale Feature Aggregation Con- former for Automatic Speaker Verification,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.289827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.937296Z digest=sha256:2a1b1eb2e32f1d2a3a57065c3d5e80cb150ef9d495037917fabc71371c05eea0

Observation 662e5a3b-e0d7-424e-bc6b-256c079af3e4 · outbound

This paper cites Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.279108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.940222Z digest=sha256:cca28794e45105ec35d851ad0c2548c808aa9f2c5f085b753326daa486e4dbf6

Observation 447260bd-c7cd-4507-916b-ec3ed4f55f6a · outbound

This paper cites In Defence of Metric Learn- ing for Speaker Recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification In Defence of Metric Learn- ing for Speaker Recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.267645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.944520Z digest=sha256:ecb36ea2e552ace9cb786e2d9081adfa763f0a2b46635f57fdaccaad35224a58

Observation 49ed5cc9-3b10-4dea-a8fb-32e56cc92aeb · outbound

This paper cites Multi-query multi-head attention pooling and inter-topk penalty for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Multi-query multi-head attention pooling and inter-topk penalty for speaker verification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.256580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.948506Z digest=sha256:0d3187a0eb1a1c2b40d9a636fd1c6a00228f6ce2805b5a4d83c08dd25fd47499

Observation ca578001-9df4-4d34-95a1-ec41e6d25e45 · outbound

This paper cites Explor- ing binary classification loss for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Explor- ing binary classification loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.242783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.952200Z digest=sha256:45bcf3979ad689e966fd050cffab2534a5d2ae3949b9cbf924e1881d858f57db

Observation 6cf5c2e2-72ec-481d-a0c4-de75ad522a61 · outbound

This paper cites Nplda: A deep neural plda model for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Nplda: A deep neural plda model for speaker verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.231624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.956178Z digest=sha256:3394d159905c554e657900e2e00bcd76fd3ab207ff95b60f53d1ccb3a1789f09

Observation ca1df279-d672-42a6-8fad-01f644036535 · outbound

This paper cites Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.219413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.959639Z digest=sha256:af01c108f2d484d1ab15a580db0ef1aa72e7a33ceb8c5fa08edb0db03f57e0d5

Observation cc0dba39-3f10-4b23-a105-ad9b91c58656 · outbound

This paper cites Attention back-end for automatic speaker verification with multiple enrollment utterances,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Attention back-end for automatic speaker verification with multiple enrollment utterances,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.207711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.962936Z digest=sha256:5f0db7d7749f05e4be4f555cec2e7ca0ff64c53c983986bb31521cfeb6c7a685

Observation 8fb1c116-b121-4d24-ac15-c9e94d28bc45 · outbound

This paper cites Prob- abilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Prob- abilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.196784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.966314Z digest=sha256:4bd85a6aab507dd81e8721f12839f267b643b1738c5a47e07a37e700cb212f1f

Observation 7a8bbea9-f537-430f-9007-e5b796bfa35a · outbound

This paper cites A study of speaker verification performance with expressive speech,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification A study of speaker verification performance with expressive speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.185802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.968866Z digest=sha256:d870ebaa123044e8cf016dd2e5e9f3c57890c3cbee4009dde28762be10d1fae8

Observation 975d0073-5454-4f20-b597-0ea549de3dd3 · outbound

This paper cites x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.175812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.971447Z digest=sha256:058b02f50842357ab85d434450e5a59248a4250e4d23f1feb73a5efacb9145e1

Observation 355036ba-0493-4a31-888b-0724812840b3 · outbound

This paper cites Emotion attribute projection for speaker recognition on emo- tional speech,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Emotion attribute projection for speaker recognition on emo- tional speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.163436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.974634Z digest=sha256:d36ad6a06fcdf1117e8f1e67474ff868719b51d588c9425e1cd0df16b8da895b

Observation cf971125-39da-41ee-9c76-acbc9fee0ae7 · outbound

This paper cites Segment- Level Effects of Gender, Nationality and Emotion Informa- tion on Text-Independent Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Segment- Level Effects of Gender, Nationality and Emotion Informa- tion on Text-Independent Speaker Verification,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.151537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.977992Z digest=sha256:1cc0e74dc60ca7ecad98beaf4ab999c1116e761eb7e4a8747be1f42a32beb4da

Observation 0a3ed5c0-f612-479b-98ef-9108cdf3f6e5 · outbound

This paper cites Instance-based Temporal Normalization for Speaker Verifi- cation,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Instance-based Temporal Normalization for Speaker Verifi- cation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.139636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.981024Z digest=sha256:22f5c5ecd4b64a231f5261a7fdee81346460321546addcc34727b9a020d56d67

Observation 37549ff7-4e44-4184-a0a5-a54e562cd1f2 · outbound

This paper cites Hy- brid Dataset for Speech Emotion Recognition in Russian Lan- guage,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Hy- brid Dataset for Speech Emotion Recognition in Russian Lan- guage,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.128075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.984661Z digest=sha256:ef7492a9d888221ca7cafec75505d28a5d99bf5d4a362250a5774cf4a0c05b19

Observation b7d2bb8d-0695-40f2-81fd-39aa33242bf7 · outbound

This paper cites Copypaste: An augmentation method for speech emotion recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Copypaste: An augmentation method for speech emotion recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.116391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.988353Z digest=sha256:1bb8c55e561d82ef5b7b06eed1b7e46c2675daee41882fd01a5d0ef94bd20d62

Observation 992077d4-9643-4452-af41-65fa8401c116 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recogni- tion,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Arcface: Additive angular margin loss for deep face recogni- tion,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.991313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.991313Z digest=sha256:bb572ff34bd7e471bc288007e30916b4e855906d7384ea5b4c00ebbbb3e8c85e

Observation 32601c19-731e-4846-9047-ba2e1f0537ab · outbound

This paper cites Sur- vey on speech emotion recognition: Features, classification schemes, and databases,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Sur- vey on speech emotion recognition: Features, classification schemes, and databases,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.100109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.993871Z digest=sha256:50ab37741826ab2adb37bbaf875b150aab4c720de9ff3292868ffde45321a9e6

Observation e8787bbb-9e0f-4918-a1b3-ed414dc1f771 · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification V oxceleb: Large-scale speaker verification in the wild,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.090707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.996952Z digest=sha256:87cd99daa304cb8a109f985b008c10248eb4661d5ae852d67d3489e9d623dfd7

Observation 9fe6b536-cc38-4501-b8de-3d7e303c9a97 · outbound

This paper cites Re- visiting the statistics pooling layer in deep speaker embedding learning,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Re- visiting the statistics pooling layer in deep speaker embedding learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.080892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.999565Z digest=sha256:f68c148ebed494c8748e2d6c8641aa4dcbea926eb2d6c0bbbdba3a71af3087e0

Observation da5662a4-e6e1-4ce4-a827-e6df1c00a14b · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

Learning Emotion-Invariant Speaker Representations for Speaker Verification MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:14.003406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:14.003406Z digest=sha256:1740ce79075b20365be7a44f53ce830ca954396120e2a438e8697bbe4f8fdda8

Observation 647e1a2e-06b3-48ff-ac2e-916d31a40bdc · outbound

This paper cites A study on data augmen- tation of reverberant speech for robust speech recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification A study on data augmen- tation of reverberant speech for robust speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.070342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:14.007487Z digest=sha256:c1548610403833d34adbe667ec4096888ac47bd5f9a7af83cb904e2b9220acf0

Pith citing papers

Observation 1c53738c-ad4d-4cb8-b7b5-c8af21ec7dfb · inbound

Learning Emotion-Invariant Speaker Representations for Speaker Verification cites this paper.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Learning Emotion-Invariant Speaker Representations for Speaker Verification

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:33:14.058991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:33:13.912562Z digest=sha256:90b6f2f2eac48e32199c4c5a17626f9b449d84c104dc5afcf6a3f860168265eb