Pith. sign in

Paper Citation Record · LEDGER

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

As of 17 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.19356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19356 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:58:27.925157Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65a1ec94-ccd5-4f1a-9375-d501dab9dd1e · outbound

This paper cites Integrating emotion recognition with speech recognition and speaker diarisation for conversations,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Integrating emotion recognition with speech recognition and speaker diarisation for conversations,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.544060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.543398Z digest=sha256:e5cb4add380e7d5188fb0608400be5deccdab60c2388a36aca4e9b23deb8b74d

Observation 4ae37139-1c2c-42de-b31d-67004035000d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.626269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.626269Z digest=sha256:0eaf6d8b28afb23afd4f5d69d985d10d18972fa161de951c03e95759d6c3e16b

Observation 8c0807a3-afe2-4d28-ba97-8f7d905e066b · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.637936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.637936Z digest=sha256:b1644757d9e23752b61bc192689f3b6672d915d75666188b3f1c5f8dbca43bb5

Observation a00965dc-6ad0-49e7-9e78-a44e24b5a903 · outbound

This paper cites Cross-Modality Gated Attention Fusion for Multimodal Sentiment Analysis.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Cross-Modality Gated Attention Fusion for Multimodal Sentiment Analysis

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:58:28.089107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.642608Z digest=sha256:afd48fbfb2e76f889e2edb3d0849761b259858b4ee427f9bfdb144d317062a62

Observation 8abb5c52-f9d6-4b3e-9112-00a3efa89468 · outbound

This paper cites IEMOCAP: In- teractive emotional dyadic motion capture database,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization IEMOCAP: In- teractive emotional dyadic motion capture database,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.424998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.648490Z digest=sha256:0b480e4e36fa964538ac79051896083dc2b3613d6c88e303e23bb45dead69df6

Observation 6c6fd2b4-234e-4d8e-b18e-e88208b4279a · outbound

This paper cites WhisperX: Time accurate speech transcription of long form audio,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization WhisperX: Time accurate speech transcription of long form audio,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.653087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.653087Z digest=sha256:6548199824a96e3ebccd52adcccc2590a82a02d8153076dad655a03d7c0892b6

Observation c4bae0ed-fb30-40e4-a72a-f4fa25a3730a · outbound

This paper cites Speech emotion recognition with ASR transcripts: A comprehensive study on word error rate and fusion techniques,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion recognition with ASR transcripts: A comprehensive study on word error rate and fusion techniques,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.658492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.658492Z digest=sha256:db06e2afc5a2b017909b385b2315dd1895c3fbb0caee7f76ff3ee737b2693036

Observation 66813e96-3786-467d-9511-4e03379b63fa · outbound

This paper cites emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:58:28.010121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.662205Z digest=sha256:693faaf546b6a283ff50f177348f223efe4113f7ce8fafc3cbe4110cef7b4448

Observation ea78dee3-a181-4a97-94fb-15f08dc95cfe · outbound

This paper cites Speech sentiment analysis via pre- trained features from end-to-end ASR models,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech sentiment analysis via pre- trained features from end-to-end ASR models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.412577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.666530Z digest=sha256:0afe6a44da20eaa525d5a09ef4e7899f7b0f0dbf0e51b4c70d1b1a043b156579

Observation 4094e105-2344-4330-86bd-c97793158f5b · outbound

This paper cites learning discriminative features from spectrograms using center loss for speech emotion recognition.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization learning discriminative features from spectrograms using center loss for speech emotion recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.670639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.670639Z digest=sha256:d05595db3ded5ffd79c6b2d6ad942bf642403f862b8274c88481f5e4b107ec71

Observation 45de3c0f-aed2-4757-806d-e376c58a24bf · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Multimodal multi-loss fusion network for sentiment analysis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.399819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.674919Z digest=sha256:3e40082c5a4642ca072e1e000ebffb1a5eedf609acd7671a16e399f3e2d3b280

Observation 31112366-bbc2-44f5-be01-9b8ad3113a6c · outbound

This paper cites Speech emotion recognition in dyadic dialogues with attentive interaction modeling,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion recognition in dyadic dialogues with attentive interaction modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.385822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.679943Z digest=sha256:df0c58ecf4c3d8da98f588a29b336db22138c5ca84ef7b6cf8afcd35134fde7b

Observation 9dded3b1-4def-439f-8428-588568e78806 · outbound

This paper cites Temporal context in speech emotion recognition,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Temporal context in speech emotion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.332793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.779799Z digest=sha256:e9d7dd12b94c8ee2ab1860393428f4efb5b22095cd804687f945d759533f6a64

Observation 8ab3af26-85b1-4931-92d5-6e9de366e262 · outbound

This paper cites Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed ASR,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed ASR,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.299558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.824966Z digest=sha256:c01e270e3ad1d96c9cc27788a047f6140315d35dfba03d5c8e740a8a03531b69

Observation 0d1ed5e0-e93d-41d3-86b1-1d85510c30f6 · outbound

This paper cites Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.863490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.863490Z digest=sha256:4d07cad028485a9cebe3af86ff132478191ce207a2bab66f6352d1349d5ac974

Observation 083e3fea-e673-4231-a515-c77139045883 · outbound

This paper cites Speech emotion diarization: Which emotion appears when?,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion diarization: Which emotion appears when?,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.287116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.908678Z digest=sha256:dfc8de939c8d4fd30362965c04e658944b6a75c8456f640b57ed4ffa5584b57d

Observation f3862308-facf-4b79-8a18-660760e1d4c9 · outbound

This paper cites Meeting recognition with continuous speech separation and transcription-supported diarization,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Meeting recognition with continuous speech separation and transcription-supported diarization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.273428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.912835Z digest=sha256:38854e43cf07bcbecea0e982b3e9b1202b52076dbe9efc615fce8491d61fb018

Observation 58eb620d-d29f-4bf9-8404-c0571fff8df3 · outbound

This paper cites pyannote.audio: neural building blocks for speaker diarization,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization pyannote.audio: neural building blocks for speaker diarization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.259375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.916716Z digest=sha256:b39eb4ec09202012c013b774b68167a2ca6a811d8316784987520ad933b18fc0

Observation c950f8c1-2d87-46f9-8716-5ef2086e7a65 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.920852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.920852Z digest=sha256:ae7d0d0ffd845d9d8aa2630c604a93ae3501df5e6a1109471f74993cd54189d0

Observation df8c584f-eb5a-4005-b8a8-38828189d440 · outbound

This paper cites 2019, Accepted at EMNLP-IJCNLP 2019].

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization 2019, Accepted at EMNLP-IJCNLP 2019]

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.199442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:58:27.925157Z digest=sha256:c591fe48a10d8bea1e54026c80711f423fa819bec505403e0eb59dd86b0bf615

Pith citing papers

No inbound Pith citation observations are available.