Pith. sign in

Paper Citation Record · LEDGER

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.06611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06611 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T01:53:00.127636Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact7
  • verified fuzzy45
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f693040-bd69-46b8-ae57-62bf2b178b88 · outbound

This paper cites The evolution of sentiment analysis and conversational AI: Techniques applications and future research directions.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts The evolution of sentiment analysis and conversational AI: Techniques applications and future research directions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.942291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:abd1ebf3ca1581263ba381f2f6f5a05c60ff13bbfab167fc9dff29a38aed44c0

Observation c797137c-7de2-4edd-ae15-66f2ec2735aa · outbound

This paper cites Sentiment analysis and emotion recognition from speech using universal speech representations.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Sentiment analysis and emotion recognition from speech using universal speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.881121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:a5f3a247dbc0a281ba02fd4847e3140d41e2355be65ef684a7b3ef51f95f7b8c

Observation f07b8f8e-8391-4d93-8498-6176ad613e87 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Proceedings of NeurIPS, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Proceedings of NeurIPS, pp

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.703203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:663cf5e39ba4f8b5ef33f417cccfc45c50f1d509174d3c82dbf78f8394d6d04b

Observation d21193c9-b6e7-431c-bbec-bdcb507027ae · outbound

This paper cites TweetEval: Unified benchmark and comparative evaluation for tweet classification, in: Findings of EMNLP, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts TweetEval: Unified benchmark and comparative evaluation for tweet classification, in: Findings of EMNLP, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.341409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:df88076942a23f55b09f0a72d442d7f1be01fd37e52deb3bd9301af6de1c869a

Observation 1e7a7188-cd28-45ff-9ddc-32c0a3f60164 · outbound

This paper cites Sentiment Analysis of Customer Feedback and Reviews in E-Commerce Systems, in: Proceedings of ICTCS, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Sentiment Analysis of Customer Feedback and Reviews in E-Commerce Systems, in: Proceedings of ICTCS, pp

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.623963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:aab64273e14a5ffa4d73c9f070c24a81c105700b9b6b38f818eee0dd01b76fcc

Observation 9f5f687a-49a2-4fcc-beff-cb3be8a79c5b · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts IEMOCAP: Interactive emotional dyadic motion capture database

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.008905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:e92cc8936d0cb8c4d14b46ae7b896728e9b50cee702e642796b09b7527f4fc68

Observation 6bf36bf0-23ef-4bd0-a977-ea0adbef5d99 · outbound

This paper cites The MSP-Podcast Corpus.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts The MSP-Podcast Corpus

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.125343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:f63a5f72473a2c7ff4cf332eec0ea2b1f6084c69f0719fee023dd6656a8c6237

Observation 5581cd06-56bc-48e7-a1d8-cc58f3a6080f · outbound

This paper cites German’s next language model, in: Proceedings of COLING, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts German’s next language model, in: Proceedings of COLING, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.561625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:f5148c30824476893f1eac741a7e8266c762aa686be6aaf0faa7886f5365e2bf

Observation 639c29ce-3311-40e5-ac9e-dc5278c8b5b3 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.538198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:7ae2ef034674d3a1eb1e728111fd24e1de7868d4f47ba6fee33536021a64044d

Observation b7b2eaf5-d8ad-4360-9fbb-d7fe89dfd425 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of NAACT-HLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts BERT: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of NAACT-HLT, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.571023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:a6d6f70f9a53d79dc8838ec9cc6599ea609cacf78ea385afb8af9cbc307c2bdc

Observation 0f7eca62-d1b6-4a2c-ab23-17fd0297fd19 · outbound

This paper cites Optimized Sentiment Analysis in Tagalog Speech Using PCA and BRNN on Prosodic Suprasegmental and MFCC Features, in: Proceedings of ICTC, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Optimized Sentiment Analysis in Tagalog Speech Using PCA and BRNN on Prosodic Suprasegmental and MFCC Features, in: Proceedings of ICTC, pp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.601183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:2cc5a83cd700d3cb085cb85505a2760adae0f4ee9d24d9db6ad1ec0a08dcff34

Observation 0f0b3717-bb94-4ffc-aecc-b9e4d329bfec · outbound

This paper cites A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.404930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:9ddd4d1154172258e45117e28c148cbb0e6f9fec1d7d040dbb1affff4aaba5ac

Observation cdeb7a1d-0ee1-46fa-9d83-315bb24b73f7 · outbound

This paper cites Teacher-Student Training and Triplet Loss for Facial Expression Recognition under Occlusion, in: Proceedings of ICPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Teacher-Student Training and Triplet Loss for Facial Expression Recognition under Occlusion, in: Proceedings of ICPR, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.726874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:55d6b7e5d325b66d0d9464d76471dbfef40dda1d39c80c29289a2ce2ad21d062

Observation 7e0b00e8-8c0e-40c6-acd1-5426bd551698 · outbound

This paper cites AST: Audio Spectrogram Transformer, in: Proceedings of INTERSPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts AST: Audio Spectrogram Transformer, in: Proceedings of INTERSPEECH, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.944100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:04abe9cd75065521f0b6470cfb5939113718889a5388e492dc372ff65b921921

Observation 173827fc-beb8-4ce7-b67f-01a05ee034c6 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Distilling the Knowledge in a Neural Network

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.151963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:38ef631a385af12546e2d1967beb3afeab4a82d99c1739e55eb2646fd1bcfe46

Observation fc45812c-5ef5-406c-bc89-7e90096498fc · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.352399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:2b95e3b60195de7e7cf9a0a5a1892233267da1929210cb3695007ca249549bff

Observation c9ae9372-df68-4ce1-ac83-bdbdeda3aba4 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models, in: Proceedings of ICLR.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts LoRA: Low-Rank Adaptation of Large Language Models, in: Proceedings of ICLR

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.383973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:513cf618d44eefd753c97e0580bc48092301a95cfed52dcfd26d2cfded5c2283

Observation 44ae0fd6-1df5-4da5-99f2-d10fdf2153b6 · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, in: Pro- ceedings of ICML, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, in: Pro- ceedings of ICML, pp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.474514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:a2193470794db519708ebab09c32e5e922d03b6eca67b364bd4aba068e197d46

Observation 2daecc1d-f41f-4b82-973e-41c02b5f2660 · outbound

This paper cites faster-whisper: Faster Whisper transcription with CTranslate2.https://github.com/SYSTRAN/faster-whisper.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts faster-whisper: Faster Whisper transcription with CTranslate2.https://github.com/SYSTRAN/faster-whisper

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.381658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:fff6f32318519568442cbe4c0c319d54815ae30f6ac9088fa286a91294112737

Observation b98508a2-01cb-49c5-bd1f-3374212a27c6 · outbound

This paper cites Incorporating end-to-end speech recognition models for sentiment analysis, in: Proceedings of ICRA, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Incorporating end-to-end speech recognition models for sentiment analysis, in: Proceedings of ICRA, pp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.562121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:3b28075bc8a6dc91ca32dcc7454fe1b8233b95c4b3de713f68ce58ac28d0fff5

Observation 7a028ca4-f932-4d96-92e0-96691cde939e · outbound

This paper cites Unimodal-driven distillation in multimodal emotion recognition with dynamic fusion, in: Proceedings of ICME, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unimodal-driven distillation in multimodal emotion recognition with dynamic fusion, in: Proceedings of ICME, pp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.410591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:0cb3f03eb2743fed91d6f2c47112f90f168ab38dc3628d97a82aeb1ba95bc58d

Observation 5bd5a4c2-13a7-4ecd-ac3f-a003d5b00c58 · outbound

This paper cites Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion Techniques, in: Proceedings of SLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion Techniques, in: Proceedings of SLT, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.261230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:3e893675caf5bcd3a818e7cd6f5add88840825452b08ed7a967ce01cf9048624

Observation 8ac18c4f-f752-4748-87f7-d42d163a878e · outbound

This paper cites Decoupled multimodal distilling for emotion recognition, in: Proceedings of CVPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Decoupled multimodal distilling for emotion recognition, in: Proceedings of CVPR, pp

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.322770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:464b1fa6b32c3d5d381d2b363713fe8cfd565cf0a2d5986b74cb76e29e87d9a0

Observation 3aea66e4-3850-4ebb-9bcd-e5852175771b · outbound

This paper cites A survey of deep learning-based multimodal emotion recognition: Speech, text, and face.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A survey of deep learning-based multimodal emotion recognition: Speech, text, and face

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.507364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:df0d8fb7a18b0f1de5a178b149daa4e2397e590a93dd91edd631798903e524e0

Observation 2ebb9d24-e0ba-4c28-9175-b8ad767de9a0 · outbound

This paper cites Development of interactive English e-learning video entertainment teaching environment based on virtual reality and game teaching emotion analysis.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Development of interactive English e-learning video entertainment teaching environment based on virtual reality and game teaching emotion analysis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.435195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:31e39ced783b14a5990d259c65336b5bf0296576acfd3998142dca7d8cf9e076

Observation 470f6a7e-e1b9-4267-a186-e558dc4a4df9 · outbound

This paper cites Unifying distillation and privileged information, in: Proceedings of ICLR.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unifying distillation and privileged information, in: Proceedings of ICLR

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.727337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:f547748d771ade897e0ff7a42dce4454a59f96ad370c7ba9e43a11e05ec23f43

Observation ffdce282-5df0-4fcb-bfb4-1c1f1d883b29 · outbound

This paper cites ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.113290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:ac98ed629d55da137ab0a8ae550fa13814def1a1734a55b79f6a066e0ccf8507

Observation becdda28-330f-4f03-83d9-c406ffb7824f · outbound

This paper cites Audio sentiment analysis by heterogeneous signal features learned from utterance-based parallel neural network, in: Proceedings of AffCon@AAAI, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Audio sentiment analysis by heterogeneous signal features learned from utterance-based parallel neural network, in: Proceedings of AffCon@AAAI, pp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.096140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:3695fbd3c03e0da25ad41d9a2621387b3dba31d929b4804aabdc7c0b70be32cf

Observation ffef5c22-058d-4bc6-b934-a8361559a09d · outbound

This paper cites CamemBERT: a tasty French language model, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts CamemBERT: a tasty French language model, in: Proceedings of ACL, pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.356234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:7a3cd22a578ad7b279eeace588e0d1eb3dab26a8d75f023bf0ccbf9dfa3cf56d

Observation b74604ae-6a73-4ab8-b296-85bad7aa1bed · outbound

This paper cites Learning using generated privileged information by text-to-image diffusion models, in: Proceedings of ICPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Learning using generated privileged information by text-to-image diffusion models, in: Proceedings of ICPR, pp

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.328815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:e2246a14ceb222690106bde503d5f2c2fc582ad5482651b07d53bec64d9e86a6

Observation 9d77120e-2752-4f43-a185-3bdf66a764c7 · outbound

This paper cites Verbal sentiment analysis and detection using recurrent neural network, in: Advanced Data Mining Tools and Methods for Social Computing.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Verbal sentiment analysis and detection using recurrent neural network, in: Advanced Data Mining Tools and Methods for Social Computing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.882398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:66a6421f4a49b6df60e3018ab77eec2817af79c91772868fe9aed32304841da5

Observation 5c34b6f5-2a3b-4f4d-af1c-0af8fd4e45ba · outbound

This paper cites Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.031561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:d79e40cac8526a31d48978e0a4bc97ccdb3607a175141e8df043d2f5f2bc79dc

Observation bfaf5578-4fc6-4d6f-8a12-8647fb491780 · outbound

This paper cites an unresolved cited work.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-07-11T02:07:49.174874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:59a6d513fc63a6243583f0e99307a62880e682e1e5d634c963edce771d7b8225

Observation 7fcf8860-9b95-48d0-8dcb-d715a2158b6e · outbound

This paper cites 4668–4672.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts 4668–4672

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.433904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:d24544093fb9083ccc8c532794f08c38ed19ad55b7bbf95170a6849297f842eb

Observation 8adb56cc-3895-4df3-8298-bf67eef50be9 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.159386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:16db5a2b7cb231644913b564e6e1733de49d7e85a40d61c08fe9ff57ada9dd68

Observation e3810431-9f63-4532-bbbb-7601d50ccc77 · outbound

This paper cites RoBERTuito: a pre-trained language model for social media text in Spanish, in: Proceedings of LREC, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts RoBERTuito: a pre-trained language model for social media text in Spanish, in: Proceedings of LREC, pp

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.034169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:73a234b0bfc2529a56096042fdb294ae376c40fdfa5ea51f4e3c4fb816e6b335

Observation f4e197aa-3d68-4436-b2cc-caa9b2cea59c · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, in: Proceedings of ACL, pp

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.133028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:ef24c824365d65cfb69ae70987f02b6815c075be67db34eef10c0c35b3909e81

Observation 31401229-8535-41aa-abe5-79ab315dbe2a · outbound

This paper cites Scaling speech technology to 1,000+languages.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Scaling speech technology to 1,000+languages

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.201643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:712d843b7a6ecf8385c567d2864ea0eda4932dc83f548a4ecdfc694db5106b5e

Observation 2b27edf8-5f60-414a-839d-285db828943c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision, in: Proceedings of ICML, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Robust Speech Recognition via Large-Scale Weak Supervision, in: Proceedings of ICML, pp

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.443273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:cae4a1bef9d0c2b33a66a1a3b492ece12c228c48cbab1f6234de8675926e2c7f

Observation 7687e83f-2f1a-478f-8084-fbcba3c8d3b2 · outbound

This paper cites Cascaded cross-modal transformer for audio-textual classification.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Cascaded cross-modal transformer for audio-textual classification

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.696578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:ffd4f5cd9f7580a85997e021a8a85ab56bf2f19881b6ffca199f83ae54a1ca92

Observation 17fe5ee7-83eb-411a-814d-bb49c472a88e · outbound

This paper cites Cascaded cross-modal transformer for request and complaint detection, in: Proceedings of ACMMM, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Cascaded cross-modal transformer for request and complaint detection, in: Proceedings of ACMMM, pp

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.792233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:69a263712246967aed3f69eda27e5a9f7d6b7440835a24ab6a06a004bdda0ff3

Observation f3cdb596-e0bf-4499-b398-13ab84f207f4 · outbound

This paper cites SepTr: Separable Transformer for Audio Spectrogram Processing, in: Proceedings of INTER- SPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts SepTr: Separable Transformer for Audio Spectrogram Processing, in: Proceedings of INTER- SPEECH, pp

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.758610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:7516c7620a454815e679d65bd304068b5ae07cc0ad4f0fcc0248bb92282dd805

Observation 050246b2-a0f1-436a-8777-1513cb80f3a5 · outbound

This paper cites An integrated approach for mental health assessment using emotion analysis and scales.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts An integrated approach for mental health assessment using emotion analysis and scales

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.665945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:c50a31ffc05a81d7da7f7bbcce699464313fd98ff43582bc6ea3a0bb7c9fd728

Observation 13152bb5-235f-490d-a9b3-8ff72df03477 · outbound

This paper cites Leveraging pre-trained language model for speech sentiment analysis, in: Proceedings of INTERSPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Leveraging pre-trained language model for speech sentiment analysis, in: Proceedings of INTERSPEECH, pp

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.276886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:65322af6a210684e8461c06300b2f557a4efffe4a4b6d61ec4d7b566ed15541b

Observation 9a154783-b248-45ab-9fcc-c64df45744a8 · outbound

This paper cites A comparative study on Bengali speech sentiment analysis based on audio data, in: Proceedings of BigComp, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A comparative study on Bengali speech sentiment analysis based on audio data, in: Proceedings of BigComp, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.142487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:90912942afc47bcd0954f8f3eae0b2b84959a4fae61c36a74edc734cd04e9ead

Observation 3ad7cd62-36fd-4bd7-8197-18dbb4cbf17d · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Multimodal transformer for unaligned multimodal language sequences, in: Proceedings of ACL, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.106555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:bcc8ed09d14274e0a2e5eebf7d59ae4709bf30176e7d0b417cfc59027470f6dc

Observation d4331dbd-6328-40ab-8fa6-2c0ad0ab1bc1 · outbound

This paper cites Hierarchical cross-modal attention and dual audio pathways for enhanced multimodal sentiment analysis.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Hierarchical cross-modal attention and dual audio pathways for enhanced multimodal sentiment analysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.912555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:552200546cc86ef8d0143e0f140ff2ee231c5f40d5d05dfa14e7fe2bfefbcbe5

Observation d41049fb-4d5a-4183-8be6-3675fee23ca4 · outbound

This paper cites Learning using privileged information: Similarity control and knowledge transfer.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Learning using privileged information: Similarity control and knowledge transfer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.245218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:35a3d053cd105767eb460aeb83471f21cbfbc5b050dd2eae701dca94ef591d82

Observation 2b030a15-2b87-45b2-8590-d35b6fbd9979 · outbound

This paper cites A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.061776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:32029e124a4b1d58ce2533dd9704d69f81ee795c94a8ae2e1461660e0df22f04

Observation 55ee1f27-4db3-4afe-ac07-226c743311bf · outbound

This paper cites Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content, in: Proceedings of AAAI, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content, in: Proceedings of AAAI, pp

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.664775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:47606b01dce792ec5b48dc1209978d2e7517433e792fe64afe80b1524c731ade

Observation 53ca2bb2-41f4-4adb-9b3f-f91edeec52d4 · outbound

This paper cites A Self-Adjusting Fusion Representation Learning Model for Unaligned Text-Audio Sequences.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A Self-Adjusting Fusion Representation Learning Model for Unaligned Text-Audio Sequences

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.092558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:d9f62b957899d148b2e91a7ae7bbd832f42deddd6f3126f7415d6e06a2078cd2

Observation d1bbf595-ef6c-43fa-a12b-4ec23e26d35b · outbound

This paper cites Multimodal speech emotion recognition using audio and text, in: Proceedings of SLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Multimodal speech emotion recognition using audio and text, in: Proceedings of SLT, pp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.302756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:ab4bffdc12ff7e2f92ade85192dfe0e43b71cce869a413db353bdd475ac21d68

Observation 7c0dee6f-1b7b-4c3b-971c-47b38a7c9102 · outbound

This paper cites Personality-aware multimodal driver emotion recognition towards intelligent connected vehicles.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Personality-aware multimodal driver emotion recognition towards intelligent connected vehicles

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.632327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:3e39ab7e8f397a89c5de7fac5b7b55ef0f90581b6a6d71feae5a30ee48008b5d

Pith citing papers

No inbound Pith citation observations are available.