Pith. sign in

Paper Citation Record · LEDGER

Leveraging CLIP Encoder for Multimodal Emotion Recognition

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.00903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00903 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:43.208707Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 913ffe1a-f211-4c30-af5c-e927cb5bed4b · outbound

This paper cites A survey of state-of-the-art approaches for emotion recognition in text.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A survey of state-of-the-art approaches for emotion recognition in text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.579814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:36.992164Z digest=sha256:0356b50325ddbfcbcea5dcce1bf4884e76453f13a116bfd0a03dc32c35ecc6ad

Observation 04dee490-1e60-42c1-a13c-2b93d050c45a · outbound

This paper cites Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.339800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.079396Z digest=sha256:33dde9d6d6c035cd309591e38ee169e478d7c98ee882032ab755125066647128

Observation cb2f650d-4705-42df-82f5-acd630d937c1 · outbound

This paper cites Multimodal machine learning: A survey and tax- onomy.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal machine learning: A survey and tax- onomy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.111628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.209014Z digest=sha256:6d88be51ebfb90a82e9441aeca13f01b0dc0654a8869754f241b922e2e068fcb

Observation 725b42f8-3194-498b-a8be-df7e1ff62ba0 · outbound

This paper cites Openface: An open source facial behavior analy- sis toolkit.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Openface: An open source facial behavior analy- sis toolkit

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.867373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.350776Z digest=sha256:25f5af1f3db942756cf471640632872fa661546e4264a657f0cedad46cb67754

Observation 2a8c76f5-a6b2-46f3-8f2c-5699eb9df6b1 · outbound

This paper cites Narayanan.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Narayanan

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.620722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.485921Z digest=sha256:7c8d7c67327bfd1695bea340bd388b82ac08099e335455a459ee745575abd895

Observation e0d19d05-500e-45c6-ad2b-aa5b037fce91 · outbound

This paper cites Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.456005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.573830Z digest=sha256:cb063708cbbda6a4633c6c72ebbd4d77cc59ce12a709daf2b9093af20788a1da

Observation 588d6738-779a-43ee-b67c-c0eb68d2bfc2 · outbound

This paper cites Covarep: A collaborative voice analysis repository for speech technologies.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Covarep: A collaborative voice analysis repository for speech technologies

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.196225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.677902Z digest=sha256:ebfee98e6a85243f55fbbe848bfeac64f450bc889d4ce63bb2ebaaf1af10bf9e

Observation 5b89e5de-fde1-4d06-8145-4ca440a0d685 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (CLIP).

Leveraging CLIP Encoder for Multimodal Emotion Recognition Data determines distributional robustness in contrastive language image pre-training (CLIP)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.920451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.803797Z digest=sha256:f0e35a34722386fc4d0099a2d0bf9f85820ddb788a39ad44855c5166bb0b0bd0

Observation c09c4d1d-0654-41a6-befb-da3b952077e9 · outbound

This paper cites Emoclip: A vision-language method for zero-shot video facial expression recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Emoclip: A vision-language method for zero-shot video facial expression recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.685758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.900242Z digest=sha256:b173f9c4961df71a78ed2e62cc0c1b2d86bbceb75720bc327311d48362cb8f9b

Observation 6dc821e6-76b2-4026-9523-8119e830f26b · outbound

This paper cites Bermano, Gal Chechik, and Daniel Cohen-Or.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Bermano, Gal Chechik, and Daniel Cohen-Or

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.518184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:37.970173Z digest=sha256:ad8abaf39781ee56d86d1c435dcf93ad4fb2e1c6d89cd1a4886b4ed9f625dad5

Observation 9e6b04ec-510e-470a-b56f-d88659f333f8 · outbound

This paper cites Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.247113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.152192Z digest=sha256:7ebd14ba154fa8a04ca6c4eb7aff38ec8e68045524a2c82b268c4963a8dc1029

Observation d1df064c-6823-4790-912c-e0642933214f · outbound

This paper cites Audioclip: Extending clip to image, text and au- dio.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Audioclip: Extending clip to image, text and au- dio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.908148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.269976Z digest=sha256:e388c15442dc81986a92c23191ffa4f9cd16262663aa22895c6e1f73db7dc76e

Observation 84caed99-495d-4b09-bcd0-730004d2f085 · outbound

This paper cites Misa: Modality-invariant and -specific representations for multimodal sentiment analysis.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Misa: Modality-invariant and -specific representations for multimodal sentiment analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.684058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.402857Z digest=sha256:410e37d1e14a0e943138bc582fdf5824db0237a2015e8c1c5676fc5881cccd29

Observation c5fc61c7-fa32-4458-8aae-b34888d590ab · outbound

This paper cites Deep multi-task learn- ing to recognise subtle facial expressions of mental states.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Deep multi-task learn- ing to recognise subtle facial expressions of mental states

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.470990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.536098Z digest=sha256:7ed78e3d3ae2dbbdaa2547b138e3c5cf7e3c70ec64350d04b59ddd7f7673cb45

Observation dbea8fd7-ba72-46b7-a93c-a1a526f2a0cb · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:52.286100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.664993Z digest=sha256:4fac6125d09b53c0694cd141427bb0296e8a4514b2bdb5b03cc69d0f439ad226

Observation c8b24d84-ef50-4c66-882f-6552ae2e0a09 · outbound

This paper cites Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.056692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:38.777293Z digest=sha256:da1dbb36b5965facf6a814372c52f1b5db3b8a2e0dcdb034bd89690d78ed4285

Observation cd1d26e9-3cf5-4d7c-847e-a407c6ee3032 · outbound

This paper cites Decoupled weight de- cay regularization.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Decoupled weight de- cay regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.895254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:38.895254Z digest=sha256:7226cc0fca0635b06141300fb88645ba98d255dc733e91bfaca815e65d8af01b

Observation 44bdd621-0451-497c-b7f2-6dfafc317213 · outbound

This paper cites Progressive modality reinforcement for hu- man multimodal emotion recognition from unaligned multi- modal sequences.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Progressive modality reinforcement for hu- man multimodal emotion recognition from unaligned multi- modal sequences

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.838240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.037880Z digest=sha256:5d7f31c56cece03209d6eb0052aec7293aed1552e7862da68914cc169fbe0415

Observation 8aa31651-cea6-402e-86ed-01a570660d34 · outbound

This paper cites The Stanford CoreNLP natural language processing toolkit.

Leveraging CLIP Encoder for Multimodal Emotion Recognition The Stanford CoreNLP natural language processing toolkit

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.582091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.145592Z digest=sha256:779c588e0b331753d9d38b07e6c4bef7a54d048e843274f4b9f939085fffb10f

Observation 857fd497-e843-41bf-a4db-a024a7f02497 · outbound

This paper cites M3er: Multiplicative multi- modal emotion recognition using facial, textual, and speech cues.

Leveraging CLIP Encoder for Multimodal Emotion Recognition M3er: Multiplicative multi- modal emotion recognition using facial, textual, and speech cues

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.277342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.261470Z digest=sha256:ba027b719c48c036f7638ec45c2873516d08e5b5541b5fa9f7b0718951759bbc

Observation ef198261-108d-436b-8010-f7645ee4e41e · outbound

This paper cites Ma- hoor.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Ma- hoor

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.041927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.377863Z digest=sha256:ffdcc23acaa9b33e16baded6edd4caf545a5692c65d193d06559a9a9cf42238e

Observation a1aa5eb6-223d-478b-9755-8b8da34bb7ea · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:50.749488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.513973Z digest=sha256:ff345fbbd1de239676168c5db5b6ae80720cff74890032116a4fb22112ca8c83

Observation 7ccb7c58-ee45-4233-9188-b514ebe4a1fb · outbound

This paper cites Glove: Global vectors for word representation.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Glove: Global vectors for word representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.550289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.602591Z digest=sha256:ce0a4dca0e10fb10ed515802ee4217938af40b7dd68676a13de1ea6faaf63d1b

Observation 678e6850-0043-4be8-aed8-1df42c66272e · outbound

This paper cites Found in translation: learn- ing robust joint representations by cyclic translations be- tween modalities.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Found in translation: learn- ing robust joint representations by cyclic translations be- tween modalities

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.354952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.698272Z digest=sha256:e6bd28c65e48dca6b425e72a25bdb7f81007cab2e96179613cb94e9d4671dd58

Observation a253ac0e-e3c9-4743-9234-9ad32eacc4c9 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations.

Leveraging CLIP Encoder for Multimodal Emotion Recognition MELD: A multimodal multi-party dataset for emotion recognition in conversations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.857845Z digest=sha256:8fe191a53c28430a2f1444a2daafc1ffffc47e2716a7448df224aef451478317

Observation 6de451b0-7f0e-4ee8-a5f2-bbc8c2ead762 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning transferable visual models from natural language supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.580704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.099817Z digest=sha256:54808c3f48cf4f271ced4dcc9a649b155bde0b3b658baa391946a0b457e9e069

Observation 3dbd69e5-5131-447a-b23a-44a33b9bee9d · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Fine-tuned clip models are efficient video learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.302570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.216660Z digest=sha256:eb67da72be73831b6dfb81f6f3ded331d83f0e7ae95dfda5b658ceea3ee1e6cd

Observation 85e58ac6-61c4-4912-bcc2-38ce52928e72 · outbound

This paper cites Accommodating audio modality in clip for multimodal processing.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Accommodating audio modality in clip for multimodal processing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.054639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.305682Z digest=sha256:bb2604f96d9fabf7772a383ed934c85b568b00669e15d1886e59ffde9ab8fd48

Observation d193c81f-e17c-4225-9cc9-47453ffa297a · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Leveraging CLIP Encoder for Multimodal Emotion Recognition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.406313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.406313Z digest=sha256:843bd810996997d426be47b7a4bb6336c3ffe625d583ad75ead4397ae915aca3

Observation 02fa32c0-2af0-4ba3-b118-eafcd6f16406 · outbound

This paper cites Sheikh, Rupayan Chakraborty, and Sunil Kumar Kopparapu.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Sheikh, Rupayan Chakraborty, and Sunil Kumar Kopparapu

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.809591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.519096Z digest=sha256:2d84833c8601fee11fbe423b5bf8b559e42376c93a270880c235635a8ea4ee0e

Observation e28770d1-9d25-4e3d-a5e7-78a91c58250d · outbound

This paper cites Learning Factorized Multimodal Representations.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning Factorized Multimodal Representations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.613619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.613619Z digest=sha256:773ad7af47c9139ea554575638fcc09974b9b8c7aaa4b44f16a4a2495e8e3556

Observation ae3658bf-d11c-48c9-90d6-c4241b1a6637 · outbound

This paper cites Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.556394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.740192Z digest=sha256:5363ea9b6a400658fb445365963c0ff96dcfd142968234a2616e33945ea66d33

Observation 878e9b75-0fdc-44e0-b7d8-64a1d62c6019 · outbound

This paper cites Suppressing uncertainties for large-scale facial expres- sion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Suppressing uncertainties for large-scale facial expres- sion recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.107389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.964911Z digest=sha256:645a93bf1c49ce22aed8574678c5569674347158d066384f792462a90a843567

Observation 384ab5e6-038a-4eca-866e-fabbf20f0b97 · outbound

This paper cites A novel end-to-end speech emotion recognition network with stacked transformer lay- ers.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A novel end-to-end speech emotion recognition network with stacked transformer lay- ers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.859062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.093397Z digest=sha256:b1ddf8d84d23407110a9ec0249ec0f919e5f26816667944442e613948a338746

Observation 663d340d-2ba6-4d8b-b022-586b131c6b3a · outbound

This paper cites Words can shift: Dy- namically adjusting word representations using nonverbal behaviors.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Words can shift: Dy- namically adjusting word representations using nonverbal behaviors

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.622625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.200740Z digest=sha256:23e7af95d77787f25c65b737b5a4c24fb14fdcedb7b4a39c1bbcf02409e9d3a7

Observation e9b77623-a8ab-450a-9d25-cdb6f242da6e · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Wav2clip: Learning robust audio repre- sentations from clip

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.380397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.345797Z digest=sha256:272215563a2583c66f93852d1de77832be73da5f09015d8aafc6645f29667142

Observation cb306a66-21c3-4ca9-a31a-1794d9dd62a7 · outbound

This paper cites Multi- view multi-label learning with view-specific information ex- traction.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi- view multi-label learning with view-specific information ex- traction

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.133005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.460363Z digest=sha256:fc6c3b3132519aa589890b71bf34b9bf4171dbe5a8f2b5cb409b155640e555e1

Observation abd59139-02e1-435d-a24c-f2ca5b648dc0 · outbound

This paper cites Disentangled representation learning for multimodal emotion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Disentangled representation learning for multimodal emotion recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.907855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.608214Z digest=sha256:9558c7495b1a62c4bad80c5ad6e24fa47d14b338f3764fdd1f87a0db58205b67

Observation 1129c12d-7252-48e6-aaf6-2ba57958ff04 · outbound

This paper cites Tensor fusion network for mul- timodal sentiment analysis.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Tensor fusion network for mul- timodal sentiment analysis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.670853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.742164Z digest=sha256:8aa8a8c2635efe2ef207b2031cd040dfffe361cc244e98b9eef135144f81635e

Observation ce7f53f3-42dc-45dd-b3ec-9a27b9195793 · outbound

This paper cites Memory fusion network for multi-view sequential learning.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Memory fusion network for multi-view sequential learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.412953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:41.843834Z digest=sha256:e15fa8534b5dc67b360b4a0aa2aa7665c0501d64a07604a16280406ca206d1d1

Observation 6b3618fd-9d69-44ce-858b-4ac775089ed0 · outbound

This paper cites Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.049658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.001440Z digest=sha256:4f4298bb1a59db134e0176a042d9f3d765f81a1413269835f9877de9f2a12f37

Observation cc2ab590-c864-42cc-9ecb-8436f887145a · outbound

This paper cites Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.782901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.155413Z digest=sha256:041afc458a9cf74c144d67659ad55e3badc91ee625af44010c8a9ba9c84aa0fe

Observation 91d30615-0785-4e2c-8a48-97bf67bb004d · outbound

This paper cites Multi-modal multi- label emotion recognition with heterogeneous hierarchical message passing.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi-modal multi- label emotion recognition with heterogeneous hierarchical message passing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.506580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.300181Z digest=sha256:2c1eb12f0c769806d9abf34c53c017bdc2d42b59ccfc234e299f59b2faabbc19

Observation 909c6c5d-dfff-445c-8cf2-6a801e897d6a · outbound

This paper cites A review on multi-label learning algorithms.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A review on multi-label learning algorithms

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.246829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.411096Z digest=sha256:100b87bda97c9d3ccc26577515ec53fe44490053db307116e54091fdd2ee4f13

Observation 29082bb6-f489-41d7-9edd-fb6defc99e03 · outbound

This paper cites Tailor versatile multi-modal learning for multi-label emotion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Tailor versatile multi-modal learning for multi-label emotion recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.965250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.527630Z digest=sha256:b67a4871938335e9f3d68d0d4427f4b6318ae17a4b91c798f7a095d38d6dacb5

Observation ba244c3c-8c56-433b-9be7-460865dbfa41 · outbound

This paper cites Manning, and Curtis P.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Manning, and Curtis P

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.703770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.649215Z digest=sha256:bd1647181d80dd6403190edebb1c90208c5b02c125fa2c3ee6b60f4ee07e6c3d

Observation 0ea844c1-c774-4158-8c72-3a5fb3995ca2 · outbound

This paper cites Prompting visual- language models for dynamic facial expression recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Prompting visual- language models for dynamic facial expression recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.459062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.811706Z digest=sha256:6d0594c53595bc3bf50e3595538fd48f947a183e3a24f52f46a4a66faee2f6d2

Observation beeab9aa-fa5f-480c-ba3f-6039327ce2a8 · outbound

This paper cites To demonstrate the diverse range of words associated with emotions, we conduct experiments to compare the performance of syn- onyms for ‘Emotion’ and ‘Sentiment’.

Leveraging CLIP Encoder for Multimodal Emotion Recognition To demonstrate the diverse range of words associated with emotions, we conduct experiments to compare the performance of syn- onyms for ‘Emotion’ and ‘Sentiment’

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.141826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:42.950644Z digest=sha256:41c45daa9ba5106e40d22b82a9e36c0c456fbb0509337669d49afe34e3351dd1

Observation 98a67b5e-ecc8-4b7c-a153-e5365ab6d229 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:43.830165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:43.066089Z digest=sha256:ae1214ef831922953e292c6f1eff2f33f28753818a05c1fd3006db301d944185

Observation 427fb53b-9d71-4285-baa8-990e14555f45 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:43.512945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:43.208707Z digest=sha256:599f521bc02dd96c243b94fe08be90330e863b5e68f7f341a99718cd89271406

Observation 162639de-ddba-4c62-8f94-63a655e1d9a6 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 536

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:49.864258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:39.978857Z digest=sha256:a9db12f798493659125e2850e9fa81ab604d83f8a28aff1586efc2f64ed9b9bb

Observation dab4dd87-d934-4ba9-9187-0ae552ea2eb9 · outbound

This paper cites 2, 4, 6, 7.

Leveraging CLIP Encoder for Multimodal Emotion Recognition 2, 4, 6, 7

Reference 6569

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.319086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:58:40.871715Z digest=sha256:21c7ba1dc357f5ba4b7c8c4648076549fac8fefe904d6bae2ddb28e218e7d55a

Pith citing papers

No inbound Pith citation observations are available.