Pith. sign in

Paper Citation Record · LEDGER

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2507.11892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11892 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:08:58.980743Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f80eff5-5e07-4e4f-bc11-824bde58144a · outbound

This paper cites The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:10.037724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.233116Z digest=sha256:6906429f9f4032a67d5e9610c34d6fc5e6c053d3e0329d927e97c7f612cab09f

Observation 8f33e461-9f0b-45e0-87f3-fc36ada3e9af · outbound

This paper cites Induced disgust, happiness and surprise: an addition to the mmi facial expression database,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Induced disgust, happiness and surprise: an addition to the mmi facial expression database,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.788431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.371685Z digest=sha256:79f3714b6d051353717a01c277e868e2a7ebc1b9124e38f72b568e91d8e8adff

Observation c7d2cc0f-ba65-45e7-82ba-d384efa2e394 · outbound

This paper cites Facial expression recognition from near-infrared videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Facial expression recognition from near-infrared videos,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.507381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.524779Z digest=sha256:4b75861588a5cddc7c855e5a5f025520845ef684b8fd9c79d89751513e6b227d

Observation 697b9dc7-4c81-4042-ba30-c4b59b5cdca8 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.232452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.678700Z digest=sha256:3e5cdf035bdd1e0c28ea54a15fc158d99ece1ac51ee6d81dba8a0f5efc88ab0e

Observation 7b370099-2a7c-449f-acc8-6c8f058cd5f8 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.923395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.819097Z digest=sha256:f3affe3b9ab64fafec242b5a13379dbe34e021731289fcf1b6f198c8e0344dba

Observation bbf98089-a486-4f67-9f52-4f89fe49eff0 · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for facial expres- sion recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ferv39k: A large-scale multi-scene dataset for facial expres- sion recognition in videos,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.691148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:49.985284Z digest=sha256:a5ccaff3f0b7a4f154adf8a46926c362b367b844c4a7e89db2a3c366ebee8edb

Observation 154984c9-e876-47c2-84ae-4d67398609e4 · outbound

This paper cites Dep-fer: Facial expression recognition in depressed patients based on voluntary facial expression mimicry,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dep-fer: Facial expression recognition in depressed patients based on voluntary facial expression mimicry,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.492490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:50.172639Z digest=sha256:6c5f24db1aca52534123b07b8c281efad9a4fb023e2156d31e8d2ebd740a18bd

Observation 770416af-30bd-44e0-ab48-7554f99f46f0 · outbound

This paper cites Efficient facial expression recognition with representation reinforcement network and transfer self-training for human–machine interaction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Efficient facial expression recognition with representation reinforcement network and transfer self-training for human–machine interaction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.367493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:50.384854Z digest=sha256:0aa303b208ef7750755186d0d19fc7bb8ea6a8f40e8be3e3e15a69382309e2db

Observation afccfec2-b81d-41ca-baa9-d4b7a9b331e4 · outbound

This paper cites Predicting personal- ized image emotion perceptions in social networks,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Predicting personal- ized image emotion perceptions in social networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.225797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:50.553598Z digest=sha256:67d6afb8b38dcbda74ee48474a5eb2b0481723bb2a482c84fe427c6871de8e10

Observation 8e3d8a74-1cb2-45fd-9683-5afaff8dc945 · outbound

This paper cites Spatio-temporal convolutional features with nested lstm for facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Spatio-temporal convolutional features with nested lstm for facial expression recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.056382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:50.726269Z digest=sha256:05d0d168371f9dfa7b3ae4e85a2e6ae5c9730c85ca90540376d09cde0bf227ef

Observation 5639bf19-85ab-4294-8870-a9829cdffe31 · outbound

This paper cites Saanet: Siamese action-units attention network for improving dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Saanet: Siamese action-units attention network for improving dynamic facial expression recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:50.885960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:50.885960Z digest=sha256:e272128bac1e5a6501fc7c1179f3ef6671d4bc3496c4b5017fbfa3d1f2e79100

Observation 9fe8fd94-4f81-4457-8460-8a451088566b · outbound

This paper cites Deep residual learning for image recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:51.446133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:51.446133Z digest=sha256:80e00e48097a0047b8ffec77ff5dedb0804a5a52e042cae864aeeca9059d6a60

Observation 76ef193a-40c2-48a4-8ada-3bf341c47643 · outbound

This paper cites Former-dfer: Dynamic facial expression recog- nition transformer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Former-dfer: Dynamic facial expression recog- nition transformer,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.867624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:51.579156Z digest=sha256:68db38fca095ba851cf2a1b7fa94ba36eb36f10d2a0d63d82fc2355ece9d30cf

Observation 80ec1c43-07f3-4a26-b909-4f2dc26d0eb7 · outbound

This paper cites Ex- pression snippet transformer for robust video-based facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ex- pression snippet transformer for robust video-based facial expression recognition,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:51.756681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:51.756681Z digest=sha256:01f4efe33734e53e42ccce9a1cf6034906ca01c4a45c62319d3460c8baca1cfe

Observation 0b439d16-a4e4-443e-8a0d-f2d8140ab652 · outbound

This paper cites Freq-hd: An interpretable frequency-based high-dynamics affective clip selection method for in-the-wild facial expression recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Freq-hd: An interpretable frequency-based high-dynamics affective clip selection method for in-the-wild facial expression recognition in videos,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.037086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.037086Z digest=sha256:14d61383f063399303407481d8bda2fafc393044824fba405c100e7a71b8c271

Observation 18c7a973-d1b5-4155-985c-bc9ab4855e05 · outbound

This paper cites Facial expression recognition with adaptive frame rate based on multiple testing correction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Facial expression recognition with adaptive frame rate based on multiple testing correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.651357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:52.159837Z digest=sha256:4212bac7ac2e171482dfaa5fbcef0767c911a010b0d4d50fa464639fd4f348ea

Observation a028bf9e-5a31-4acf-b5d9-cbfd4792e5d4 · outbound

This paper cites Video-based facial micro-expression analysis: A survey of datasets, features and algorithms,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Video-based facial micro-expression analysis: A survey of datasets, features and algorithms,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.455501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:52.268113Z digest=sha256:36ece1365357646f68ca5e7ab14c5ef523b4e13a41af28d6fe8c6ddbab70c775

Observation 7fd07951-24e8-4737-90d0-a71bd023e60d · outbound

This paper cites Dynamic facial expression recognition under partial occlusion with optical flow reconstruction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dynamic facial expression recognition under partial occlusion with optical flow reconstruction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.145228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:52.428689Z digest=sha256:01d55da932d9f9221d2d243c733a0a10b9b139b487870654a0dbb50cff581b0d

Observation c2d87dcf-a356-49e0-8c39-70c1a6bfed9d · outbound

This paper cites From static to dynamic: Adapting landmark-aware image models for facial expression recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition From static to dynamic: Adapting landmark-aware image models for facial expression recognition in videos,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.532422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:52.667666Z digest=sha256:cf4706da9dd653c734e01dd410c3ae1430caa03bcebbdd03ff64d9e6a48b0dd3

Observation e9feacc7-04c8-4699-b675-a01cc7560e19 · outbound

This paper cites A Survey on Facial Expression Recognition of Static and Dynamic Emotions.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition A Survey on Facial Expression Recognition of Static and Dynamic Emotions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.781534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.781534Z digest=sha256:806a9bd3590ad15e28fce08b3f01f72ca21b275f568ead6e4308ba13c4ffc5dd

Observation 0e0c819f-6ba6-426b-852d-26ab597a8c06 · outbound

This paper cites Prompting Visual-Language Models for Dynamic Facial Expression Recognition.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Prompting Visual-Language Models for Dynamic Facial Expression Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.937811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.937811Z digest=sha256:5ce1b851260f811a4babda46fcd541e056dbaeb0c8de8ad4ac739ae7b370de43

Observation 4e4d0b94-72eb-4c57-bbbd-9f175b5202df · outbound

This paper cites Emoclip: A vision-language method for zero-shot video facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Emoclip: A vision-language method for zero-shot video facial expression recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.252373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:53.030094Z digest=sha256:77110f17196db1a0ae4a1a6d1af0c5ef57eb9d73e948d4a896fdd656352289ab

Observation 964eea29-3d44-4c2b-b79f-40a7c4518bee · outbound

This paper cites Domain knowledge enhanced vision-language pretrained model for dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Domain knowledge enhanced vision-language pretrained model for dynamic facial expression recognition,

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T17:08:59.856215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:53.192045Z digest=sha256:93d6dd7f99bb69ac4b9108ecd0f7f379f6a9dff4fd6f71b02dfbae77c13d3194

Observation cfdaf705-c9c9-4083-b02b-5a5371ec2dfe · outbound

This paper cites Enhancing zero-shot facial expression recognition by llm knowledge transfer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Enhancing zero-shot facial expression recognition by llm knowledge transfer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.963193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:53.286611Z digest=sha256:d4e7ffc5f9d52e1f8d77546fe948a8fdba2c8c04b252a9d36de452a81aca5b19

Observation 65ef274e-e86c-4c71-a2ac-6d5b54a12094 · outbound

This paper cites Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.456059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.456059Z digest=sha256:6ec128a74386db78969726811cf1b2bbd27790a378b6f8be1c86171b29adddfd

Observation 1013a341-58d0-45a2-ba81-9eef0ddfdd18 · outbound

This paper cites Describe your facial expressions by linking image encoders and large language models.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Describe your facial expressions by linking image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.528965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:53.642988Z digest=sha256:6def0996f5892b193695cb4ab49d556e078d2fe97cbdf38cbc296e51d9d9b0fa

Observation 4efb4a3a-9c7f-49e8-8ad4-233405174190 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.771017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.771017Z digest=sha256:7b13c65c47176c06410d9b70c4f2a954153adf2e63f3c2f975a1a6f43128d668

Observation acbe33dc-ae8f-43e1-a5a4-b2d18cd43e08 · outbound

This paper cites Hierarchical Transformers for Multi-Document Summarization.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hierarchical Transformers for Multi-Document Summarization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.883130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.883130Z digest=sha256:d4118df684cca6a669ffdeeb233ee6514f7e57776584adc18fca8b5bc4fdd846

Observation 725a3fc7-35fd-49bc-ac14-a37a4d0e6b46 · outbound

This paper cites HDT: Hierarchical Document Transformer.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HDT: Hierarchical Document Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.989568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.989568Z digest=sha256:85240467b458a35ac8be04fb7a524e2b0b4cecb201e4dc8c7779baf5da06dc1e

Observation ce50ff7e-9644-4481-b55c-2d587c8e3144 · outbound

This paper cites Multi-task learning of hierarchical vision-language representation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multi-task learning of hierarchical vision-language representation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.349116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:54.100810Z digest=sha256:77fe40aa804f946acd6e991e4ecbdfaaff8f897933b0bcfeac1e7bf8d7f908a0

Observation 0ca66b65-6b30-4633-b844-1cc8f35be18f · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.214011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.214011Z digest=sha256:92469fdf5aa848dcd4dee17a34f50fc6869220a009f80507d096a8015556d5bc

Observation 242d3e97-996a-4f4f-86b9-67bcf9fb00d2 · outbound

This paper cites Hierarchical modular network for video captioning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hierarchical modular network for video captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.135439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:54.323924Z digest=sha256:2784d93a84cab06c06d333100927666f1163980a6922a7013c34f0355b51bd07

Observation cb83ea55-fd32-41b7-87bb-c41d4f594515 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Learning transferable visual models from natural language supervision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.486076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.486076Z digest=sha256:6a6fa55764c9d0bb34ca11f6aa77c9f3e360ebaf144a1ad913cd48f6941797b8

Observation 190d69ec-35eb-46a9-8f1f-3895e2b534c9 · outbound

This paper cites Ceprompt: Cross-modal emotion-aware prompting for facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ceprompt: Cross-modal emotion-aware prompting for facial expression recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.672039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:54.620430Z digest=sha256:4ae64543cda39b1bf4d56525e7b356bd8d399c84be4d465a198837f0a53fe2c0

Observation 06cef10d-69f3-47cb-a657-62d3cab51910 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.917377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:54.819539Z digest=sha256:78da25c83ad4e3ded907ad1b3f104e5aeb0c998d74a5942653ecb85ad845b796

Observation 5110dfd2-d14c-4c5c-9502-d00fe488b898 · outbound

This paper cites Recent advances in optimal transport for machine learning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Recent advances in optimal transport for machine learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.570793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:54.948095Z digest=sha256:a5b05d4f095c4f54ce63825cbffebacdc36578aa3d83cc8212764a00555b9064

Observation 7cb898f7-e789-4677-a77b-f6aee9e2dcfe · outbound

This paper cites Reliable weighted optimal transport for unsupervised domain adaptation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Reliable weighted optimal transport for unsupervised domain adaptation,

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-06T17:08:55.068589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:55.068589Z digest=sha256:39a8e9651cce1b43266a2ae56cc8a2e5993857a503b35a96282fa23e62980bef

Observation ff886e30-8ed9-46c3-8ff7-4223c9bcf03d · outbound

This paper cites Unsupervised learning of visual features by contrasting cluster as- signments,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Unsupervised learning of visual features by contrasting cluster as- signments,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.349557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:55.184879Z digest=sha256:51404c7cf0226c5c23d480806df38afb349b02cb5edc00b460bb6697bf78152e

Observation 8550292e-d670-4762-98d3-edd840d6d036 · outbound

This paper cites Optimal partial transport based sentence selection for long-form document matching,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Optimal partial transport based sentence selection for long-form document matching,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.049201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:55.298606Z digest=sha256:e93662d42e0155ef5e39964059d9436f28c1ddbd022b1af7de8b328d1cd14294

Observation ee3bd80c-b722-45b9-8bfb-73909537e983 · outbound

This paper cites Learning to align sequential actions in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Learning to align sequential actions in the wild,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.796463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:55.446654Z digest=sha256:f67d57ac20a3148677ac4f8132869fbbc5ef0ce8af8eb48b1ae8624a12b53985

Observation 55dcb424-4071-4d6f-abbd-814214ae3886 · outbound

This paper cites What when and where? self-supervised spatio-temporal grounding in untrimmed multi-action videos from narrated instructions,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition What when and where? self-supervised spatio-temporal grounding in untrimmed multi-action videos from narrated instructions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.611461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:55.583412Z digest=sha256:0798af3e225832be89ec696db2d865a84f05c79df2b587664d53592b3dab5610

Observation a0a922a5-baee-4e19-869d-6cc760f1c08f · outbound

This paper cites Multi-granularity Correspondence Learning from Long-term Noisy Videos.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multi-granularity Correspondence Learning from Long-term Noisy Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:55.704482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:55.704482Z digest=sha256:49ed2f71a3b0f946abfa6c9be036b8813135ef60c10452f98649ba0192e3d600

Observation 3acb2635-48d1-4bd5-8d6f-82a43a92f8d2 · outbound

This paper cites Spatial- temporal graphs plus transformers for geometry-guided facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Spatial- temporal graphs plus transformers for geometry-guided facial expression recognition,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.842125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:55.829988Z digest=sha256:db5b6d5f2b629f0f716eb002ac8703a463844eacc6cabaad094c80e55054f333

Observation 22258bd3-e252-4538-b727-1377306c1864 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multimodal transformer for unaligned multimodal language sequences,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.956146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:56.109359Z digest=sha256:38335e300b1f279cc0b77acb0b911d4a8b570e32453a6a7ab6c79a7c56e863fe

Observation e3b68206-aaa7-440c-8809-87610c2582a6 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.248682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.248682Z digest=sha256:a12cf8fe9c4c19c5d44fc15e5876d95a1ef3346aaec4b5a4d5d8a6d71bb39499

Observation adc6628f-859d-410f-b0bc-9ed5bdce20f6 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.385574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.385574Z digest=sha256:248d1f634455af2a9ed3a60447f6af1ddbc8a8e71362cb8e0fa9b9ff707e7c06

Observation 91adfa9d-5d95-413a-b0dc-543745217bdc · outbound

This paper cites Cliper: A unified vision-language framework for in-the-wild facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Cliper: A unified vision-language framework for in-the-wild facial expression recognition,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.695424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:56.546603Z digest=sha256:026c0c3a97665371608ecadaec8ad73afcf9d69ce07825906b1bd2c49b617cf3

Observation a771a906-1941-46cd-9384-d776e716758d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.717302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.717302Z digest=sha256:1bbd5550c3b1a2adff52c9e3479694dc4220dcc624295b6b9614ff2c0532cc4b

Observation 2ef3652f-b4ba-4a6a-9be8-6020b5bdcb3b · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.854748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.854748Z digest=sha256:1300e9efa5e950e9216a9f5d0bd891af7753977cd4cf9e8950a0842e0882b8cc

Observation 0ff0bbb6-eb8a-41d8-ba58-e1eb879b040d · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.989600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.989600Z digest=sha256:8e23d950d05c67bd1dbcb0ab1e3572279317cce0f7dfc393f1309af374b2f9e3

Observation d7fdddc8-f420-4089-a31d-40ca86630da1 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:57.095778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:57.095778Z digest=sha256:7e7330a856bca2dd6c8925f1e87e886d86ac777689ff557133ee774626860af2

Observation c0f13447-3841-4612-bf12-eab2e9546562 · outbound

This paper cites Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:57.212905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:57.212905Z digest=sha256:6cdab770ff21377bd3c500a6540b0e75ca14b998b52909d4dcf07fbcb7e149f5

Observation 40c53bc4-6176-455b-b879-1db2e02025b8 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- 14 tioning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- 14 tioning,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.372262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:57.346343Z digest=sha256:5ae40bdd9edc406aa3217a4e53cacc84b893523cfbccfb8178d2c35195d26fe6

Observation 0d0349cd-faf4-4122-9945-7563471fbdbb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.133934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:57.491789Z digest=sha256:76fa50a1988eb54379c8d1b6ea2fcf72afe1afd35f9db217a8429b4a4aff803b

Observation c3208862-0942-423d-9cbc-8b85e139500c · outbound

This paper cites Emotion recognition using imperfect speech recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Emotion recognition using imperfect speech recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.900774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:57.635608Z digest=sha256:7619d366a1a5c8dffe7fb803cd90e099171570f61660a564925704269d1a6538

Observation 0edd08c9-2483-4ae2-b727-2e19eecae091 · outbound

This paper cites Posterior calibration for multi- class paralinguistic classification,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Posterior calibration for multi- class paralinguistic classification,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.577572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:57.805680Z digest=sha256:b5b914c55f0d48fa1019954d1295e7b95388c96d0e19fc855351e53f0d455c77

Observation b06a9716-7b97-40fa-a894-d07f957a2a4c · outbound

This paper cites Rethinking the learning paradigm for dynamic facial expression recog- nition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Rethinking the learning paradigm for dynamic facial expression recog- nition,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.317518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:57.917571Z digest=sha256:a4127eec705e328ded64f95b0a51475c5408c94957ebb3e9729b08054218fc3b

Observation e9fc0632-63cf-4362-9066-95f90e68fbd1 · outbound

This paper cites A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.053197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.053197Z digest=sha256:bd0fdf0dbdfb4e2265cccbd08aeafbaed5be464b0bc04854c4c104a93c67c066

Observation bb674402-af34-407a-86c6-4a7b01908490 · outbound

This paper cites Clip-aware expressive feature learning for video-based facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Clip-aware expressive feature learning for video-based facial expression recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.028233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:58.167769Z digest=sha256:8e81c5fa1ad48a7cd340534709a8bdea296f24a60f06924233ed2ac1d46f23fa

Observation 5c53bea4-72b1-4926-8ec7-ed45a3ceb9b8 · outbound

This paper cites NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.322047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.322047Z digest=sha256:e512129d9ff0e5c17f4429d47e1640de6b3536ba9d2656135d680a9a1ef6f4bf

Observation c01291c4-fcdc-4459-8ccd-dd3fbabce61f · outbound

This paper cites Logo-former: Local-global spatio-temporal transformer for dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Logo-former: Local-global spatio-temporal transformer for dynamic facial expression recognition,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.457153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.457153Z digest=sha256:7b3b2cddf4384ec391aa278852aad7669f96408cb62cae296f91e8d438b7447a

Observation 025acfcc-723b-4024-9ea7-4130c4153e70 · outbound

This paper cites Intensity-aware loss for dynamic facial expression recognition in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Intensity-aware loss for dynamic facial expression recognition in the wild,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:58.592846Z digest=sha256:2fe9669151c40b94223c9a4241dbcf140cfc967c49d5179538e12a5fe8e51faa

Observation 44e6292c-dc8c-490b-8ef8-c2126613213e · outbound

This paper cites Transformer-based multimodal emotional perception for dynamic facial expression recogni- tion in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Transformer-based multimodal emotional perception for dynamic facial expression recogni- tion in the wild,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.714047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:58.705759Z digest=sha256:00f9adcc930d184b5d1e9ba66231ae5cec59782df8ed23f9a04f23800eb3fa13

Observation 6d6c81c3-f280-4671-9ef9-d9f219746c94 · outbound

This paper cites Svfap: Self-supervised video facial affect perceiver,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Svfap: Self-supervised video facial affect perceiver,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.483481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:58.844014Z digest=sha256:4178c90a44fae97cc034e2ce91a397f2e4e2dce69f370cd791554c36b7e596b2

Observation 56f4fb15-b13d-4326-8e4b-6bd0070b72bd · outbound

This paper cites Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recogni- tion,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recogni- tion,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.250849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:08:58.980743Z digest=sha256:684194a1182f6e19324927604ffe76279c7959708942ac144a45c4c02a554d24

Pith citing papers

No inbound Pith citation observations are available.