Pith. sign in

Paper Citation Record · LEDGER

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.19356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19356 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:58:27.925157Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65a1ec94-ccd5-4f1a-9375-d501dab9dd1e · outbound

This paper cites Integrating emotion recognition with speech recognition and speaker diarisation for conversations,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Integrating emotion recognition with speech recognition and speaker diarisation for conversations,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.544060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.543398Z digest=sha256:b2b9f2f6b7a7ef877db12bf1983ee3ee73deb502723ebd572e324ec44b610ec5

Observation 4ae37139-1c2c-42de-b31d-67004035000d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.626269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.626269Z digest=sha256:27f18d65ec4ae7471c300687c142fbc106d3d1c8782ed631e8365d32dc90b2c3

Observation 8c0807a3-afe2-4d28-ba97-8f7d905e066b · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.637936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.637936Z digest=sha256:825dda283ed44a118f664666add7658243ad245baa0b9514f1aabce5c5ae818e

Observation a00965dc-6ad0-49e7-9e78-a44e24b5a903 · outbound

This paper cites Cross-Modality Gated Attention Fusion for Multimodal Sentiment Analysis.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Cross-Modality Gated Attention Fusion for Multimodal Sentiment Analysis

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:58:28.089107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.642608Z digest=sha256:07a09f1da49d5467d813f0e7d81cd9800e1a709d1b81b3d81dc3af900549e7e2

Observation 8abb5c52-f9d6-4b3e-9112-00a3efa89468 · outbound

This paper cites IEMOCAP: In- teractive emotional dyadic motion capture database,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization IEMOCAP: In- teractive emotional dyadic motion capture database,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.424998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.648490Z digest=sha256:2165d1030cccbe22d3218600a7bd5f7cbfecb9ddcf833d034a7aef1d42f67963

Observation 6c6fd2b4-234e-4d8e-b18e-e88208b4279a · outbound

This paper cites WhisperX: Time accurate speech transcription of long form audio,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization WhisperX: Time accurate speech transcription of long form audio,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.653087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.653087Z digest=sha256:1010f4044c9c0f315d548b64bd2a9c0c51f5e6d916ba1273369e540b5b147974

Observation c4bae0ed-fb30-40e4-a72a-f4fa25a3730a · outbound

This paper cites Speech emotion recognition with ASR transcripts: A comprehensive study on word error rate and fusion techniques,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion recognition with ASR transcripts: A comprehensive study on word error rate and fusion techniques,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.658492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.658492Z digest=sha256:13dfa0ff57d4ab5ffea9b9157272c78f6cf934eb6f110f85d9107fc1e6d37fa7

Observation 66813e96-3786-467d-9511-4e03379b63fa · outbound

This paper cites emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:58:28.010121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.662205Z digest=sha256:d41448623177ffb81d3834e5c44df4a2e41d4e182380f73bf7bba44f5543517b

Observation ea78dee3-a181-4a97-94fb-15f08dc95cfe · outbound

This paper cites Speech sentiment analysis via pre- trained features from end-to-end ASR models,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech sentiment analysis via pre- trained features from end-to-end ASR models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.412577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.666530Z digest=sha256:14786a52f8056743235b95d41923ab9587359e6c401dd1ff6414134dfe6bc91e

Observation 4094e105-2344-4330-86bd-c97793158f5b · outbound

This paper cites learning discriminative features from spectrograms using center loss for speech emotion recognition.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization learning discriminative features from spectrograms using center loss for speech emotion recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.670639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.670639Z digest=sha256:84e8deddddc1e0361f277c8417530aa5ffff592645ee6b217525d6e1da20b1d8

Observation 45de3c0f-aed2-4757-806d-e376c58a24bf · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Multimodal multi-loss fusion network for sentiment analysis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.399819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.674919Z digest=sha256:a60d7387e6d334d8b899d6815f57e0f06a577d384f8bb302b17e4440f70db9f1

Observation 31112366-bbc2-44f5-be01-9b8ad3113a6c · outbound

This paper cites Speech emotion recognition in dyadic dialogues with attentive interaction modeling,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion recognition in dyadic dialogues with attentive interaction modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.385822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.679943Z digest=sha256:c01f1793b6722363cbb7368d04e6cc0f05bb3978885b9584f3db5398a8fd3945

Observation 9dded3b1-4def-439f-8428-588568e78806 · outbound

This paper cites Temporal context in speech emotion recognition,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Temporal context in speech emotion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.332793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.779799Z digest=sha256:1bcc6d436b279a4110c60bfc9fada548dea78b581ece26b38dcf02cf1db59175

Observation 8ab3af26-85b1-4931-92d5-6e9de366e262 · outbound

This paper cites Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed ASR,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed ASR,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.299558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.824966Z digest=sha256:846adc03779278c4e5f87b8a77a4f59390230b90e50c309736f6544475e62399

Observation 0d1ed5e0-e93d-41d3-86b1-1d85510c30f6 · outbound

This paper cites Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.863490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.863490Z digest=sha256:cc53a7e45f7caba6386cbebc2202d9e2fb429c64a2b63e6f8238ef4d801e413e

Observation 083e3fea-e673-4231-a515-c77139045883 · outbound

This paper cites Speech emotion diarization: Which emotion appears when?,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Speech emotion diarization: Which emotion appears when?,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.287116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.908678Z digest=sha256:aee5f196b42fe248767956f92d7144ba211f9b5f2a3d2f72d9882e7a3d5a91bc

Observation f3862308-facf-4b79-8a18-660760e1d4c9 · outbound

This paper cites Meeting recognition with continuous speech separation and transcription-supported diarization,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Meeting recognition with continuous speech separation and transcription-supported diarization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.273428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.912835Z digest=sha256:3bfa55923ac483aacb4f5ebfe52adcd00654c1cc23fe8b576c983f95d68e6124

Observation 58eb620d-d29f-4bf9-8404-c0571fff8df3 · outbound

This paper cites pyannote.audio: neural building blocks for speaker diarization,.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization pyannote.audio: neural building blocks for speaker diarization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.259375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.916716Z digest=sha256:4d83c79ae988daf4cbf7908be215433bff0746590ab9ea715c49204f9f14ffb5

Observation c950f8c1-2d87-46f9-8716-5ef2086e7a65 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:27.920852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:58:27.920852Z digest=sha256:3151479e1bed3403b3ffd92e7fb290d25fb11328d0a280b39f40473cda751ce6

Observation df8c584f-eb5a-4005-b8a8-38828189d440 · outbound

This paper cites 2019, Accepted at EMNLP-IJCNLP 2019].

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization 2019, Accepted at EMNLP-IJCNLP 2019]

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:28.199442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:58:27.925157Z digest=sha256:409e9d5e93e2fc77ade337d91543e154fef99097747c21f9d0f8c722d9a6b2a0

Pith citing papers

No inbound Pith citation observations are available.