Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:43.208707Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.00903.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:43.208707Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 913ffe1a-f211-4c30-af5c-e927cb5bed4b · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition A survey of state-of-the-art approaches for emotion recognition in text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04dee490-1e60-42c1-a13c-2b93d050c45a · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb2f650d-4705-42df-82f5-acd630d937c1 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal machine learning: A survey and tax- onomy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 725b42f8-3194-498b-a8be-df7e1ff62ba0 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Openface: An open source facial behavior analy- sis toolkit
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a8c76f5-a6b2-46f3-8f2c-5699eb9df6b1 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Narayanan
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0d19d05-500e-45c6-ad2b-aa5b037fce91 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 588d6738-779a-43ee-b67c-c0eb68d2bfc2 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Covarep: A collaborative voice analysis repository for speech technologies
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b89e5de-fde1-4d06-8145-4ca440a0d685 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Data determines distributional robustness in contrastive language image pre-training (CLIP)
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c09c4d1d-0654-41a6-befb-da3b952077e9 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Emoclip: A vision-language method for zero-shot video facial expression recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6dc821e6-76b2-4026-9523-8119e830f26b · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Bermano, Gal Chechik, and Daniel Cohen-Or
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e6b04ec-510e-470a-b56f-d88659f333f8 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1df064c-6823-4790-912c-e0642933214f · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Audioclip: Extending clip to image, text and au- dio
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84caed99-495d-4b09-bcd0-730004d2f085 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Misa: Modality-invariant and -specific representations for multimodal sentiment analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5fc61c7-fa32-4458-8aae-b34888d590ab · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Deep multi-task learn- ing to recognise subtle facial expressions of mental states
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dbea8fd7-ba72-46b7-a93c-a1a526f2a0cb · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8b24d84-ef50-4c66-882f-6552ae2e0a09 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd1d26e9-3cf5-4d7c-847e-a407c6ee3032 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Decoupled weight de- cay regularization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44bdd621-0451-497c-b7f2-6dfafc317213 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Progressive modality reinforcement for hu- man multimodal emotion recognition from unaligned multi- modal sequences
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8aa31651-cea6-402e-86ed-01a570660d34 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition The Stanford CoreNLP natural language processing toolkit
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 857fd497-e843-41bf-a4db-a024a7f02497 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition M3er: Multiplicative multi- modal emotion recognition using facial, textual, and speech cues
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef198261-108d-436b-8010-f7645ee4e41e · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Ma- hoor
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a1aa5eb6-223d-478b-9755-8b8da34bb7ea · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ccb7c58-ee45-4233-9188-b514ebe4a1fb · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Glove: Global vectors for word representation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 678e6850-0043-4be8-aed8-1df42c66272e · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Found in translation: learn- ing robust joint representations by cyclic translations be- tween modalities
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a253ac0e-e3c9-4743-9234-9ad32eacc4c9 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition MELD: A multimodal multi-party dataset for emotion recognition in conversations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6de451b0-7f0e-4ee8-a5f2-bbc8c2ead762 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning transferable visual models from natural language supervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dbd69e5-5131-447a-b23a-44a33b9bee9d · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Fine-tuned clip models are efficient video learners
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 85e58ac6-61c4-4912-bcc2-38ce52928e72 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Accommodating audio modality in clip for multimodal processing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d193c81f-e17c-4225-9cc9-47453ffa297a · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02fa32c0-2af0-4ba3-b118-eafcd6f16406 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Sheikh, Rupayan Chakraborty, and Sunil Kumar Kopparapu
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e28770d1-9d25-4e3d-a5e7-78a91c58250d · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning Factorized Multimodal Representations
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3658bf-d11c-48c9-90d6-c4241b1a6637 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 878e9b75-0fdc-44e0-b7d8-64a1d62c6019 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Suppressing uncertainties for large-scale facial expres- sion recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 384ab5e6-038a-4eca-866e-fabbf20f0b97 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition A novel end-to-end speech emotion recognition network with stacked transformer lay- ers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 663d340d-2ba6-4d8b-b022-586b131c6b3a · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Words can shift: Dy- namically adjusting word representations using nonverbal behaviors
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9b77623-a8ab-450a-9d25-cdb6f242da6e · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Wav2clip: Learning robust audio repre- sentations from clip
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb306a66-21c3-4ca9-a31a-1794d9dd62a7 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi- view multi-label learning with view-specific information ex- traction
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abd59139-02e1-435d-a24c-f2ca5b648dc0 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Disentangled representation learning for multimodal emotion recognition
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1129c12d-7252-48e6-aaf6-2ba57958ff04 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Tensor fusion network for mul- timodal sentiment analysis
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce7f53f3-42dc-45dd-b3ec-9a27b9195793 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Memory fusion network for multi-view sequential learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b3618fd-9d69-44ce-858b-4ac775089ed0 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc2ab590-c864-42cc-9ecb-8436f887145a · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91d30615-0785-4e2c-8a48-97bf67bb004d · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi-modal multi- label emotion recognition with heterogeneous hierarchical message passing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 909c6c5d-dfff-445c-8cf2-6a801e897d6a · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition A review on multi-label learning algorithms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 29082bb6-f489-41d7-9edd-fb6defc99e03 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Tailor versatile multi-modal learning for multi-label emotion recognition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba244c3c-8c56-433b-9be7-460865dbfa41 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Manning, and Curtis P
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ea844c1-c774-4158-8c72-3a5fb3995ca2 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Prompting visual- language models for dynamic facial expression recognition
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation beeab9aa-fa5f-480c-ba3f-6039327ce2a8 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition To demonstrate the diverse range of words associated with emotions, we conduct experiments to compare the performance of syn- onyms for ‘Emotion’ and ‘Sentiment’
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98a67b5e-ecc8-4b7c-a153-e5365ab6d229 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 427fb53b-9d71-4285-baa8-990e14555f45 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 162639de-ddba-4c62-8f94-63a655e1d9a6 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work
Reference 536
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dab4dd87-d934-4ba9-9187-0ae552ea2eb9 · outbound
Leveraging CLIP Encoder for Multimodal Emotion Recognition 2, 4, 6, 7
Reference 6569
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.