Pith. sign in

Paper Citation Record · LEDGER

Leveraging CLIP Encoder for Multimodal Emotion Recognition

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.00903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00903 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:43.208707Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 913ffe1a-f211-4c30-af5c-e927cb5bed4b · outbound

This paper cites A survey of state-of-the-art approaches for emotion recognition in text.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A survey of state-of-the-art approaches for emotion recognition in text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.579814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:36.992164Z digest=sha256:01f8aaea9d78b157cdb04e50f8fd52e9e8776f0a44494ff4d3f144fb64398951

Observation 04dee490-1e60-42c1-a13c-2b93d050c45a · outbound

This paper cites Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.339800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.079396Z digest=sha256:9a01ff4364434b989eba46366ee4268f43bc63dd67d7664becc82787ef73bffe

Observation cb2f650d-4705-42df-82f5-acd630d937c1 · outbound

This paper cites Multimodal machine learning: A survey and tax- onomy.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal machine learning: A survey and tax- onomy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.111628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.209014Z digest=sha256:36aeebc085119c8309bb6c24c6c0c341e50173c18834aa6f380b2f70cd052d36

Observation 725b42f8-3194-498b-a8be-df7e1ff62ba0 · outbound

This paper cites Openface: An open source facial behavior analy- sis toolkit.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Openface: An open source facial behavior analy- sis toolkit

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.867373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.350776Z digest=sha256:674165a67e2f9673c09ec1911cf315e3f4bd762e867fdf9003785667b2a6eddf

Observation 2a8c76f5-a6b2-46f3-8f2c-5699eb9df6b1 · outbound

This paper cites Narayanan.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Narayanan

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.620722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.485921Z digest=sha256:e9fb3e6c627c8ce5854d1a889e2f4afc107f84bea788d70789e9023aadc4de31

Observation e0d19d05-500e-45c6-ad2b-aa5b037fce91 · outbound

This paper cites Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.456005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.573830Z digest=sha256:21cd74f5fff58853f0450a8f1b775f6b9b0cc80c588131d539de1f5994ca373c

Observation 588d6738-779a-43ee-b67c-c0eb68d2bfc2 · outbound

This paper cites Covarep: A collaborative voice analysis repository for speech technologies.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Covarep: A collaborative voice analysis repository for speech technologies

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.196225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.677902Z digest=sha256:f62967b7b276d0b0bdc62a68d3c61f8755377c682873d3f11d6ae49e85043802

Observation 5b89e5de-fde1-4d06-8145-4ca440a0d685 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (CLIP).

Leveraging CLIP Encoder for Multimodal Emotion Recognition Data determines distributional robustness in contrastive language image pre-training (CLIP)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.920451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.803797Z digest=sha256:d9db5b0c32b4ce2ea3ad1b256448e159e514f9ba121178e9bacd6b9d83bfd65e

Observation c09c4d1d-0654-41a6-befb-da3b952077e9 · outbound

This paper cites Emoclip: A vision-language method for zero-shot video facial expression recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Emoclip: A vision-language method for zero-shot video facial expression recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.685758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.900242Z digest=sha256:ac34106bf45ec7f65f051e6d96af6ea1352ebad948094bec701d9e0f43dc7ab8

Observation 6dc821e6-76b2-4026-9523-8119e830f26b · outbound

This paper cites Bermano, Gal Chechik, and Daniel Cohen-Or.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Bermano, Gal Chechik, and Daniel Cohen-Or

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.518184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:37.970173Z digest=sha256:54fb102bf82a5fcb4be79e83e15bb6dd4493fb9a5c2a47d08ed91a979e8cebdb

Observation 9e6b04ec-510e-470a-b56f-d88659f333f8 · outbound

This paper cites Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.247113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.152192Z digest=sha256:8d9680c993e3344033544a2d074eb7d0058c80c580dda7af38e1f4e53ab22349

Observation d1df064c-6823-4790-912c-e0642933214f · outbound

This paper cites Audioclip: Extending clip to image, text and au- dio.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Audioclip: Extending clip to image, text and au- dio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.908148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.269976Z digest=sha256:edff0f6ffb77e118cacbbe769e6e6101b7f2120e2d756bdba0bd68b17f75d36b

Observation 84caed99-495d-4b09-bcd0-730004d2f085 · outbound

This paper cites Misa: Modality-invariant and -specific representations for multimodal sentiment analysis.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Misa: Modality-invariant and -specific representations for multimodal sentiment analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.684058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.402857Z digest=sha256:263b85ef42cf95acf2f9084f8f9182634970a1fd5c42f44fd2a9f5292d73af04

Observation c5fc61c7-fa32-4458-8aae-b34888d590ab · outbound

This paper cites Deep multi-task learn- ing to recognise subtle facial expressions of mental states.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Deep multi-task learn- ing to recognise subtle facial expressions of mental states

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.470990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.536098Z digest=sha256:d63277a9abc5185e83bf6ef609391dcc3414da5e465fe7adb8471a50db31bf83

Observation dbea8fd7-ba72-46b7-a93c-a1a526f2a0cb · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:52.286100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.664993Z digest=sha256:0049176baba41de8d321261ee8d4834828ba5a7f126dabd9dbc63eed7f604543

Observation c8b24d84-ef50-4c66-882f-6552ae2e0a09 · outbound

This paper cites Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.056692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:38.777293Z digest=sha256:40ab88d8ad27de63e98f9aa7cfc05c29b9567ad60b1c45af931b06d8d7e000b0

Observation cd1d26e9-3cf5-4d7c-847e-a407c6ee3032 · outbound

This paper cites Decoupled weight de- cay regularization.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Decoupled weight de- cay regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.895254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:38.895254Z digest=sha256:af596c01b112bf11ada15853070fd2f6d7486a352e77676e99deb085bcf56d6c

Observation 44bdd621-0451-497c-b7f2-6dfafc317213 · outbound

This paper cites Progressive modality reinforcement for hu- man multimodal emotion recognition from unaligned multi- modal sequences.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Progressive modality reinforcement for hu- man multimodal emotion recognition from unaligned multi- modal sequences

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.838240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.037880Z digest=sha256:44882f9808bb363dc7a5c0f2c161c76998f7386c57c3c9f0808bc872f9e5e938

Observation 8aa31651-cea6-402e-86ed-01a570660d34 · outbound

This paper cites The Stanford CoreNLP natural language processing toolkit.

Leveraging CLIP Encoder for Multimodal Emotion Recognition The Stanford CoreNLP natural language processing toolkit

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.582091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.145592Z digest=sha256:892722d5e835d417ee963b3e2c7c0efa69fc60ba7d06a37980b7dee726018d65

Observation 857fd497-e843-41bf-a4db-a024a7f02497 · outbound

This paper cites M3er: Multiplicative multi- modal emotion recognition using facial, textual, and speech cues.

Leveraging CLIP Encoder for Multimodal Emotion Recognition M3er: Multiplicative multi- modal emotion recognition using facial, textual, and speech cues

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.277342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.261470Z digest=sha256:d31c7ccb3fd10631faa2ab36263008613a524b1226fe2ac69431182fc5fcfdc5

Observation ef198261-108d-436b-8010-f7645ee4e41e · outbound

This paper cites Ma- hoor.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Ma- hoor

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.041927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.377863Z digest=sha256:cc314b4308c5065aa0dcf5f4256518ff143d1614d90294b8b48ffd9598367149

Observation a1aa5eb6-223d-478b-9755-8b8da34bb7ea · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:50.749488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.513973Z digest=sha256:fa2c50bbb80e1b4f561d5665237c70ac0e719d4067130c801c6cda8884da1701

Observation 7ccb7c58-ee45-4233-9188-b514ebe4a1fb · outbound

This paper cites Glove: Global vectors for word representation.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Glove: Global vectors for word representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.550289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.602591Z digest=sha256:136f2e3026a5974492313f91d067fddd3d0c1bfd30dacb8c3672ec9a6e78363b

Observation 678e6850-0043-4be8-aed8-1df42c66272e · outbound

This paper cites Found in translation: learn- ing robust joint representations by cyclic translations be- tween modalities.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Found in translation: learn- ing robust joint representations by cyclic translations be- tween modalities

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.354952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.698272Z digest=sha256:3282452e23f2a41f9beb68f9a67a95dde8cd39474be3112510c2539b99fdc9c5

Observation a253ac0e-e3c9-4743-9234-9ad32eacc4c9 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations.

Leveraging CLIP Encoder for Multimodal Emotion Recognition MELD: A multimodal multi-party dataset for emotion recognition in conversations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.857845Z digest=sha256:642fa3e9ac034c6f5b0e039549085d8b15dd8b8b9fbed2f55a424ea46b76b705

Observation 6de451b0-7f0e-4ee8-a5f2-bbc8c2ead762 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning transferable visual models from natural language supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.580704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.099817Z digest=sha256:463c930140aea2d6163d94920e3b9277b448c49f169799b0f0c1184f3a557883

Observation 3dbd69e5-5131-447a-b23a-44a33b9bee9d · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Fine-tuned clip models are efficient video learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.302570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.216660Z digest=sha256:77c080a314f0c82e08e2ae6b0a8a52d32b6881c4e48d49b2b79fb17ec843f187

Observation 85e58ac6-61c4-4912-bcc2-38ce52928e72 · outbound

This paper cites Accommodating audio modality in clip for multimodal processing.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Accommodating audio modality in clip for multimodal processing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.054639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.305682Z digest=sha256:3a5ac0b5ac29a9dca06915f416bf7f5e1e22907980b10b64a6029580e7399dc5

Observation d193c81f-e17c-4225-9cc9-47453ffa297a · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Leveraging CLIP Encoder for Multimodal Emotion Recognition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.406313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.406313Z digest=sha256:0e827690e91c30c601dd34696d61ac715405f1b13083b1445e5968d0d891baec

Observation 02fa32c0-2af0-4ba3-b118-eafcd6f16406 · outbound

This paper cites Sheikh, Rupayan Chakraborty, and Sunil Kumar Kopparapu.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Sheikh, Rupayan Chakraborty, and Sunil Kumar Kopparapu

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.809591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.519096Z digest=sha256:f268caeb820c40f655ed882077790210f3358772f23d099296ed0d8be335b0d4

Observation e28770d1-9d25-4e3d-a5e7-78a91c58250d · outbound

This paper cites Learning Factorized Multimodal Representations.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Learning Factorized Multimodal Representations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.613619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.613619Z digest=sha256:73765fb6c558c543e0c2cdd016d6b4290f5ffb2cc634fcb39711aab78ef0d89a

Observation ae3658bf-d11c-48c9-90d6-c4241b1a6637 · outbound

This paper cites Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.556394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.740192Z digest=sha256:cb197703d84e2eae980a0c2777832eaaa4a7f1be2876ec8d4714250e58e36051

Observation 878e9b75-0fdc-44e0-b7d8-64a1d62c6019 · outbound

This paper cites Suppressing uncertainties for large-scale facial expres- sion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Suppressing uncertainties for large-scale facial expres- sion recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.107389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.964911Z digest=sha256:f1d8f78518ada61e95d6948ea3cd257c681acd637a4c05eabb4363ec5064beb6

Observation 384ab5e6-038a-4eca-866e-fabbf20f0b97 · outbound

This paper cites A novel end-to-end speech emotion recognition network with stacked transformer lay- ers.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A novel end-to-end speech emotion recognition network with stacked transformer lay- ers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.859062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.093397Z digest=sha256:d2082e75685ac6204c4215ab20c0d811abb0f81dc5568fa9fdfe62bd10efb8f4

Observation 663d340d-2ba6-4d8b-b022-586b131c6b3a · outbound

This paper cites Words can shift: Dy- namically adjusting word representations using nonverbal behaviors.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Words can shift: Dy- namically adjusting word representations using nonverbal behaviors

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.622625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.200740Z digest=sha256:85b92f59a2e6012e2f311abc6b18a3ed4831976f4823d5701411eb6db33864ec

Observation e9b77623-a8ab-450a-9d25-cdb6f242da6e · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Wav2clip: Learning robust audio repre- sentations from clip

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.380397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.345797Z digest=sha256:c9c2622532f41658f62a9d7f442c93d19105d230a504a02dc64f20221b030bdf

Observation cb306a66-21c3-4ca9-a31a-1794d9dd62a7 · outbound

This paper cites Multi- view multi-label learning with view-specific information ex- traction.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi- view multi-label learning with view-specific information ex- traction

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.133005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.460363Z digest=sha256:93e5e14297d89ed758aaba702d7e3553fae13b72a5582fe8f028c213a44334f0

Observation abd59139-02e1-435d-a24c-f2ca5b648dc0 · outbound

This paper cites Disentangled representation learning for multimodal emotion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Disentangled representation learning for multimodal emotion recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.907855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.608214Z digest=sha256:b296c70895e84f625137d52086190a261726831cd72f5af571edca8b36bdbb88

Observation 1129c12d-7252-48e6-aaf6-2ba57958ff04 · outbound

This paper cites Tensor fusion network for mul- timodal sentiment analysis.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Tensor fusion network for mul- timodal sentiment analysis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.670853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.742164Z digest=sha256:9daab4a0cad3d3da0c6712e18608ca26cc558c2b532b0bbf9f92a3e5d258d8f8

Observation ce7f53f3-42dc-45dd-b3ec-9a27b9195793 · outbound

This paper cites Memory fusion network for multi-view sequential learning.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Memory fusion network for multi-view sequential learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.412953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:41.843834Z digest=sha256:1696ad97bf98f143932e8f5bd68335b406146c44d6e75dd094af35887157df7d

Observation 6b3618fd-9d69-44ce-858b-4ac775089ed0 · outbound

This paper cites Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.049658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.001440Z digest=sha256:06d52fa1aa03f3aca12183245ba75e3b8bedef15f8d8bba54c265ede668e7dd8

Observation cc2ab590-c864-42cc-9ecb-8436f887145a · outbound

This paper cites Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.782901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.155413Z digest=sha256:aec3cc317ef6309d9df9c3f7411d0c64c3fbf0d2686f9cbf9928fa998646a36a

Observation 91d30615-0785-4e2c-8a48-97bf67bb004d · outbound

This paper cites Multi-modal multi- label emotion recognition with heterogeneous hierarchical message passing.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Multi-modal multi- label emotion recognition with heterogeneous hierarchical message passing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.506580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.300181Z digest=sha256:1cce02e29353ebd86c0d8630188e4c983e1004667e7e9169a20dee11e4d1219c

Observation 909c6c5d-dfff-445c-8cf2-6a801e897d6a · outbound

This paper cites A review on multi-label learning algorithms.

Leveraging CLIP Encoder for Multimodal Emotion Recognition A review on multi-label learning algorithms

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.246829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.411096Z digest=sha256:c24c292bd9d11827fc4594eb2057fc8544ffb11455ec6cf922314e8133f815e7

Observation 29082bb6-f489-41d7-9edd-fb6defc99e03 · outbound

This paper cites Tailor versatile multi-modal learning for multi-label emotion recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Tailor versatile multi-modal learning for multi-label emotion recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.965250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.527630Z digest=sha256:16dc2885a6d6fca5f527c8f72979e0e253e7fd95366400ca159ca96675b3cea9

Observation ba244c3c-8c56-433b-9be7-460865dbfa41 · outbound

This paper cites Manning, and Curtis P.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Manning, and Curtis P

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.703770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.649215Z digest=sha256:3708a0d7f9c53f57be337030f64cee16025116adc1435cf8ba71d092a3b002da

Observation 0ea844c1-c774-4158-8c72-3a5fb3995ca2 · outbound

This paper cites Prompting visual- language models for dynamic facial expression recognition.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Prompting visual- language models for dynamic facial expression recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.459062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.811706Z digest=sha256:7ddbc57a736c2b8a830a575b5160fa5db1a4c2b654bb39ffe7187c3290398565

Observation beeab9aa-fa5f-480c-ba3f-6039327ce2a8 · outbound

This paper cites To demonstrate the diverse range of words associated with emotions, we conduct experiments to compare the performance of syn- onyms for ‘Emotion’ and ‘Sentiment’.

Leveraging CLIP Encoder for Multimodal Emotion Recognition To demonstrate the diverse range of words associated with emotions, we conduct experiments to compare the performance of syn- onyms for ‘Emotion’ and ‘Sentiment’

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.141826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:42.950644Z digest=sha256:5c9b23595dad5bb2fa84eac4365b9dc3cd883afa0567865923238d51b9412055

Observation 98a67b5e-ecc8-4b7c-a153-e5365ab6d229 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:43.830165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:43.066089Z digest=sha256:bbbe3c3269ae338870ba824f0a8fd62d3ce18ebeb9d82dd9c7d1ec9dc52238c7

Observation 427fb53b-9d71-4285-baa8-990e14555f45 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:43.512945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:43.208707Z digest=sha256:2030440fae7cc22ea77ecc84cc971c8c12f50ff5c0ef792ea2cf623bb2229d31

Observation 162639de-ddba-4c62-8f94-63a655e1d9a6 · outbound

This paper cites an unresolved cited work.

Leveraging CLIP Encoder for Multimodal Emotion Recognition Unresolved cited work

Reference 536

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:49.864258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:39.978857Z digest=sha256:7890e20926197e60e0af90dd08e2234d420c71d6f371fb6da85aec6b6214c6ed

Observation dab4dd87-d934-4ba9-9187-0ae552ea2eb9 · outbound

This paper cites 2, 4, 6, 7.

Leveraging CLIP Encoder for Multimodal Emotion Recognition 2, 4, 6, 7

Reference 6569

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.319086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:58:40.871715Z digest=sha256:fc0268d881a72c19e91b34ee6517da5d9b5e739850f3b865f2d42c8b2875448d

Pith citing papers

No inbound Pith citation observations are available.