Pith. sign in

Paper Citation Record · LEDGER

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

As of 11 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2501.07978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07978 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:54.320561Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:24:36.541729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:02:37.644844Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc93ae32-cc76-4acc-ba60-cbec4923ba3e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.884278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.884278Z digest=sha256:1b5e8ac96cfb80724b1f5f74b5f53c1a9d7d822faae74c061c2b983d5b7ded9e

Observation 4abc8b02-edbf-4fe1-a5c3-b75fdb296c4a · outbound

This paper cites Emotion recognition in speech using cross- modal transfer in the wild, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Emotion recognition in speech using cross- modal transfer in the wild, 2018

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.889385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.889385Z digest=sha256:c090d7fc7c90fce152e764fd32cfc0525e46feaeff2d776aab43cae159f63539

Observation cdd82447-b527-47b6-8bcc-f1a8e953315f · outbound

This paper cites Claude-3.5, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Claude-3.5, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.893802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.893802Z digest=sha256:134d5276786608357a258c1c50b214cfa39443d8e6a0760b4c491bc7e8c4a082

Observation d6f6185c-f02d-4cf3-a693-a66927114450 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.903232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.903232Z digest=sha256:e03dc7724176fec4d7932448f119fccb929f01f2a72da281785203d5623ba226

Observation 8cbe17d3-6f5a-421c-ba14-585c11660183 · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Collecting highly paral- lel data for paraphrase evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.907800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.907800Z digest=sha256:fcf925298d0a07f9a80a2b1d7b1347227659b5bc8d73ea466c71feb5bc83cdb0

Observation 14417cbe-13db-43dc-8f44-377172cdc49b · outbound

This paper cites FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.913411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.913411Z digest=sha256:6320af4e793046ee0d11e24aac812ab147f5beb939f59db521774a0fcdc9ed82

Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.918678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.918678Z digest=sha256:930bea96e5ab951568372154b32b5a05b5d2a3af5ccd6b61a35e5aaa3993aa7a

Observation a70c08be-1453-477b-a074-daa4cfb9a060 · outbound

This paper cites Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.923207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.923207Z digest=sha256:37409349e3315727b5d834d38ed960e98489dc6fbe07c29241fa4410773b98ff

Observation 058f14fb-70b5-4ece-a08c-26afc70ea1c0 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.927345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.927345Z digest=sha256:59e09778b724397b9c60ac9345c63f0e8dba3813f7cb828e3d34c57e5b210ab2

Observation 8d653a64-b7f1-4921-926f-8b741118b8ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.930933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.930933Z digest=sha256:bf0aa00f67ce04ffae48be612468264cf02e9dd537ecea7fa7bbc78292fd008b

Observation 184a51db-02cc-49ff-a15b-11671cca3d01 · outbound

This paper cites Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.934690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.934690Z digest=sha256:716f72209ca380d7cf490b2229df0ba8f8870d12463ff8ddbb405f681102c0c4

Observation 70be547e-fdad-44e1-8463-20e460e9a722 · outbound

This paper cites Diffusionrig: Learning personal- ized priors for facial appearance editing.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Diffusionrig: Learning personal- ized priors for facial appearance editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.937932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.937932Z digest=sha256:9a5ed9af30a9d8772f88053d0f646ad42e87cda38873b788cada59d9bbdc4c4b

Observation 868dbd0f-f384-43a1-8d93-af5d185986ac · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.942983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.942983Z digest=sha256:dc2b1432ff8f97c11b10e4d3b1002fc05a8cd83ce5590d74cb5dfac137e57c70

Observation 6236c81b-5727-4fb4-8b60-8bc7774588a9 · outbound

This paper cites Strongsort: Make deep- sort great again.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Strongsort: Make deep- sort great again

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.947008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.947008Z digest=sha256:8f1cd04428ebac1c27e7e35ce5d5408c644de868c2347d967adc0f3935346b62

Observation 17472c72-cec0-40a2-99b2-84c691b507e7 · outbound

This paper cites EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.950534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.950534Z digest=sha256:631ed04507a74834e1ca743ecf1ff67ae268aff68acd4732ca8b374bba4be2df

Observation 8e20eb0f-a12a-465a-9847-5d95dbc2c3b4 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.954454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.954454Z digest=sha256:34496ea5a0f8359463c45a320e2d3ee8bf6b61ffa7afd9dbb200eee4984486ba

Observation 2ee3927a-96fb-42f8-b2fd-123ecbc2904c · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.958465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.958465Z digest=sha256:8c76eddac5fb69d4e0225ab7ca46f8050afac127985ec3a8afd2b8b392a1e77d

Observation 289d0ca9-4987-45a4-9795-01f08b933ada · outbound

This paper cites Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.963217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.963217Z digest=sha256:52ee2e061d69e18f16cbceb5bbf814f6a7b27e94e77d95ca82b61039584b949e

Observation b69d2990-0a6e-4675-a16e-24c27a8da406 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LoRA: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.967225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.967225Z digest=sha256:1be8e71d64b577adb41999354e6583b3308341ed0214c07643f3c87efc0bfa92

Observation b7602365-cb26-4dfd-ba34-4fa8ff056ba9 · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multimodal Pretraining for Dense Video Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.972371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.972371Z digest=sha256:f4451e24f871dbff9bcd195c4fe8d558fbfc73a48da6f76d15df4ad61b869700

Observation 4d501ff5-cfb1-45a9-8ef6-2707817343ae · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.977080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.977080Z digest=sha256:a452b6ca93097008e29e768a71321df04e867f8a3aa9e7ee5e077af800e0eb84

Observation 33d0f7e5-2d02-4c8f-9663-7bed3fa6c57e · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.982395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.982395Z digest=sha256:e0f9648196c13425bcf7fb8d2b415b29449713f18d4c2d92122ac30230a685ae

Observation 0d82b7f6-3584-472c-9619-29f1ddd9bf3a · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.986831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.986831Z digest=sha256:7458e0e98803a3ccd0150222e801f1357f485f534b0ce6e9ba2514352602cf03

Observation ac506810-02a2-41f5-af59-04498bbaecb6 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.991893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.991893Z digest=sha256:c3bd165a526ec983989f10dd546157265cb89598ca73ffc632b0269efa49c90b

Observation b19a64f9-dbfe-445d-8e6b-174955e70bb1 · outbound

This paper cites Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.996900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.996900Z digest=sha256:b9512946674ea7ebb4f4722434347f29994d320d6e7bb2e4cac4728c12e90484

Observation 49941daa-11fa-4191-86d9-ffe12cf973be · outbound

This paper cites Afew-va database for valence and arousal estimation in-the-wild.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Afew-va database for valence and arousal estimation in-the-wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.000691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.000691Z digest=sha256:e88c3b9fdceec2275f583834b4d1c669956c31f7677341e558c28eeb07fc0947

Observation 61377989-453c-4780-bcc3-f74d63bdc9c6 · outbound

This paper cites Dense-captioning events in videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dense-captioning events in videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.005726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.005726Z digest=sha256:ff06bfde6de64df8a154ebc2f93d468a2ca70b6a2dbd3aae55816f64188a4559

Observation b64a0f55-59a7-4ea3-b690-8dfe4fe2cbb8 · outbound

This paper cites Context-aware emotion recognition net- works.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Context-aware emotion recognition net- works

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.430420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.016846Z digest=sha256:ab65b2ca8a9a2550e3c91e8fdffdc49e6d8aad2e64fcd23bdfad3d372f8c0765

Observation c2632afb-5b1a-4883-b728-12af4c6b4f4e · outbound

This paper cites Llava-next: What else influences visual instruction tun- ing beyond data?, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava-next: What else influences visual instruction tun- ing beyond data?, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.416512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.021962Z digest=sha256:3cf2cfac9415a171ef96e2d109f20eac2644b39178ff1f0600970962d153310e

Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.026558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.026558Z digest=sha256:0ceca0198dfda741f4cce57a20bc61493475590abeeee16b07aef489dd19705c

Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.031175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.031175Z digest=sha256:c9a8b7c90ea35e5fbe3e7002ab8b20df8a3ed34ad890d7f47d0eeeb0a4144e79

Observation 7f8c1543-a945-4910-af7e-409144b7f88f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.035535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.035535Z digest=sha256:200daccaafd19d56c2131e6816fd7cf8d60ff79095976cbb6ee1f2dde0e13c38

Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.040489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.040489Z digest=sha256:2ad2105cf09dfcc74b72094883744140f9442ecede506774cb5d94ae61aef181

Observation 9094410c-9abc-4c79-bf0a-6f01cc769b6e · outbound

This paper cites Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.395528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.045144Z digest=sha256:f9b23f1a638cd0d5c97731f66f38f678dcbea588010bbd83639e625dbfdc0a36

Observation b5f712da-2f25-4616-844c-628f256cccd3 · outbound

This paper cites Facial affective behavior analysis with instruction tuning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial affective behavior analysis with instruction tuning, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.384319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.049977Z digest=sha256:bd6f4d44c3054d1b0c1b3a19435e006c9e579769ca9f3df587a9005864ce2d37

Observation 88837485-f367-4190-a78b-1cc280f44245 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llama-vid: An image is worth 2 tokens in large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.373429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.054027Z digest=sha256:dedd9f78b49411342dbe2b77a97cd30f1aad0a705e828b803b24b8a1c587de9a

Observation 231af9a6-5224-4f1f-9aa0-310260037293 · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Photomaker: Customizing realistic human photos via stacked id embedding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.358412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.058317Z digest=sha256:235ca2ca038766a41206fcf8e45be4a471c5dca73aea8d9ab10edd0e91372f47

Observation fa8073b5-3bd1-43cd-963c-2284534b242b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.063000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.063000Z digest=sha256:8d1eede3a323e23e373ff8cfdd4894f4dfe38efc2df4d62279d20b7a338b2bed

Observation c2d2eb12-c3d7-4f98-85e0-64962d28268b · outbound

This paper cites Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.342396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.068366Z digest=sha256:a7bf394fab474e0c603144c92900b2ba282321eee1973780a1ed6300955b6992

Observation b1ae5ed9-7d2a-466e-baf4-923d799a0625 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Improved Baselines with Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.072952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.072952Z digest=sha256:2dee8309168ab3a050a3b1849ccd8fe0860c2a97581b537274614449aada8d36

Observation deaaf7f4-3e1b-4922-a469-7e7c51f095d5 · outbound

This paper cites Visual instruction tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.079056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.079056Z digest=sha256:21de55ea435c4d84c0eff6f4a4dd697aeed570c5be8aa57a947285e8dfbe092c

Observation 2a8b8f70-e38a-45c8-8359-777e250eb065 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.083552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.083552Z digest=sha256:96f68dbd11288b1b63ad84f0209a153a8019f4b382d21a919059f723d520681e

Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.087918Z digest=sha256:6b0464b4e11aab1dd722fe5f791c9b82d48f78e2fb10e9941a50bd0556fa3c4d

Observation 4d075e73-3a34-4aa7-874c-120ee1be9681 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.092524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.092524Z digest=sha256:c077b98887bb43563b29a6e5d336984ae436c0bb66fc0d7a6e790bcfe0ff0c75

Observation 803469f6-deef-4297-b87b-92833edf75ad · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.321731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.098492Z digest=sha256:9372778fdbc4165c166e54486ebead7737005dd9da8020699e397fa2a27cfaea

Observation e8f43fe3-bbf1-4cc1-bce9-5e410716a276 · outbound

This paper cites DamoFD: Digging into backbone de- sign on face detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness DamoFD: Digging into backbone de- sign on face detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.309781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.102402Z digest=sha256:19802118dc5734c1f670dfc90c6b591674247d59a0eb3248a4c1d85a28ef0741

Observation 1a21361f-515c-4376-a7f1-c07d45c73d82 · outbound

This paper cites Livingstone and Frank A.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Livingstone and Frank A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.297610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.109723Z digest=sha256:0b8de5b17b962a9989cc036c4af6b884a7fa54a89855255cd9d1b872d89aa267

Observation c9ccb828-d05c-4b47-ae5e-5b61b57bade5 · outbound

This paper cites Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.286051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.114235Z digest=sha256:8eb817f93c40d0ca898f68e5402853f6dcb23b27f180e15a85a1ef3329432186

Observation 2d697a4f-8b27-4320-b587-e8742ea614d5 · outbound

This paper cites The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.273230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.118502Z digest=sha256:13a4f1cb88ee2555f12408a60965b181c288f238a276cdd208e2ccbc43cd3407

Observation b5c470e0-abcf-40aa-ba94-8c820facbbd1 · outbound

This paper cites Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.257974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.123016Z digest=sha256:c5c8096d5c1c2aa80b3aa281268f304e0bbd9a767052ba73655057673ad1d110

Observation d0e643a4-7903-4f7c-91ee-b39f6e6d84e9 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.127444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.127444Z digest=sha256:86b102bc8c00c4a0dd54a67d0a05ec4efed26846421f890938113e74a908aaa4

Observation 8ffd7c41-e2d5-4641-a30c-47452385392f · outbound

This paper cites Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.246121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.131730Z digest=sha256:7d17fa6f673a5712b862134946b09597784a80bb3245b622a3630c54edab5f9d

Observation e82a982a-ec88-45ab-9a48-d0a347527cc1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.135546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.135546Z digest=sha256:229fa1dfcb1addf74db0652028b77bb7456cc9f5e5984c2b03859318f7eceea5

Observation 710122ae-8fba-4ba8-aa6f-3f5913cf3f7b · outbound

This paper cites The importance of emotional regulation in mental health.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The importance of emotional regulation in mental health

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.227238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.139183Z digest=sha256:2fcbb400945f22d2e570374d6b6ab24bacca99d5e81482bf849ef02590a7a6d7

Observation 719aa5fe-2e0a-4e13-9845-7ea75642b62f · outbound

This paper cites FaceXFormer: A Unified Transformer for Facial Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FaceXFormer: A Unified Transformer for Facial Analysis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.143445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.143445Z digest=sha256:76294835462c17f692a146a88769c558ea8a016aa2598eea620109b1ef1bcdfe

Observation 9f37c381-ba5c-4eb6-8586-9ccfaf984da2 · outbound

This paper cites Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.215683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.148000Z digest=sha256:25ed68c512930741653649f0ec31f190dd5017b276f372d6a48ef7353c3f1e96

Observation d662422b-4de8-4bca-bc29-6399216d3fe9 · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:55.202907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.152472Z digest=sha256:8a41ca54020bf0530456c82a4d63b1cb10d5594133f21f6f60e7392a2c7d560a

Observation f6f9678f-c6e2-4ba8-b2e7-9df8a433944e · outbound

This paper cites Gpt-4v(ision) system card.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4v(ision) system card

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.189876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.156778Z digest=sha256:b6042ff2561ec0796efe6a0609dadde0444f943ae1769b0116b925b513ce8f4b

Observation feebdaca-2d58-4d68-b35f-94a596c1ce9d · outbound

This paper cites Gpt-4 technical report, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4 technical report, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.178240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.160875Z digest=sha256:b4b1509574f05c3d747e58ab65783566a67bb87790766badec161a7a9d302932

Observation 68ab66d2-4688-4f8d-8a79-199d155881f5 · outbound

This paper cites Gpt-4o system card, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4o system card, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.166879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.167146Z digest=sha256:ca6eac4e3df2f76e212497920130ad695adef20c82e7d08401e068da18bfe4ba

Observation fed2f313-d31c-47f1-a744-9a6b3e0308f9 · outbound

This paper cites Digihuman: A con- versational digital human with facial expressions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Digihuman: A con- versational digital human with facial expressions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.152984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.171845Z digest=sha256:4063c46e8926f4d60966924cbdd21a5dfadd4e82fa2c1abc8118856048ca1360

Observation 1d57b4a1-9157-4f8f-a53c-d32938b2a18e · outbound

This paper cites Pantic, M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pantic, M

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.140334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.176751Z digest=sha256:3c99198df2b20d3b9aece7366b7a2b6775e1a5d61a07daed599bbd0156c69e03

Observation 5af1b47d-f09a-4c75-9ece-bce449cfd8f5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning transferable visual models from natural language supervi- sion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.181114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.181114Z digest=sha256:d91f954b2d4a7b6f93c693254c739aa103d51f2f74b7cc06c5c5e82d60c1c686

Observation 113794d7-32e1-42f4-adb7-8d05ccabfe0a · outbound

This paper cites Movie description.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Movie description

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.121966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.186009Z digest=sha256:390fe723344a340e8b83eb3a8afc80208280ea2781e6dcb563226567515e829e

Observation d423450c-b4bf-45a3-a3a2-63e59e6853fb · outbound

This paper cites Multi-view dynamic facial action unit detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multi-view dynamic facial action unit detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.111532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.190865Z digest=sha256:da93518202918528ae387ad5a27e8a60f519836c86e901bab1179dcd4ebd6882

Observation adc5adc1-c310-4d96-94c4-3e30f01709c8 · outbound

This paper cites Deep adaptive attention for joint facial action unit detection and face alignment.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep adaptive attention for joint facial action unit detection and face alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.098590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.195507Z digest=sha256:0dab5298ad5e08bf551f7a5d9b56b90d303d9b5aab47a4c91e0bfaba8eb8e131

Observation d57afeef-1816-4ffa-ab42-5dd4bc8d113b · outbound

This paper cites Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas).

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.084813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.200363Z digest=sha256:6164d284e49f211fb450d2fe10ec29612c37eaadec7a5ce631bdd5d8d448da7d

Observation e9a31929-d21b-4728-9114-4d844a9aa896 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gemini: A family of highly capable multi- modal models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.073989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.204861Z digest=sha256:e66f694ef4eda988901c3cd9abfdf6f79f5ae00c7b17678b36456d4a8bddda95

Observation 2954039a-fbb4-45d5-a87f-b04dd8aaa782 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2.5: A party of foundation models, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.209214Z digest=sha256:8b4e6450091edfc0c6e6bdf6ed476fdc7eda80227f5461a89d192939e5f39b4d

Observation e2ee8a76-9ba8-42df-b627-373ccf0dfbdd · outbound

This paper cites Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.053978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.213735Z digest=sha256:c8e746afbf62706724a0dcef67df589e4a40dd0a4c91b1fdd9e90bf3a900963f

Observation 96b16d77-d1c1-42b2-9937-6c6580196ac4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cider: Consensus-based image description evalua- tion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.042966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.218449Z digest=sha256:fb24cf0fe1f1e8399bad74bf4feb5abadce89beabeb50b50e7335331dee84418

Observation fe9dd5c5-1566-4bc3-808d-f1c7e92a8f29 · outbound

This paper cites A survey on the pipeline evolution of facial capture and tracking for digital humans.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness A survey on the pipeline evolution of facial capture and tracking for digital humans

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.028702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.222398Z digest=sha256:03782428fc6a55de1f01863210f982c7b4e9811b4af222e071cef02818ae210d

Observation f278228a-cd22-4c69-906e-39655cb90a70 · outbound

This paper cites Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.017717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.226871Z digest=sha256:7df01646c82f626619d132913ac7daa71de3edf1fd3c13753b02d4c5a2c9644f

Observation 0b0fa2dc-26a1-483e-b33e-485eb0196ed9 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.007121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.231149Z digest=sha256:828a8573358a445411f1faf6e02b42851ba7d4d991b1a2d1cd20a715db3e3720

Observation 329bd536-d2db-4793-a525-e6978bf8566a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.235456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.235456Z digest=sha256:4c67940532463d05754e737316a051398a5c703d84208052680866c88b4c4844

Observation f7cb10f3-2000-433b-95c5-bdfcf0a0d9a0 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.996676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.239911Z digest=sha256:da077ec16cd99d581cb3ac3d8ad7b3a1dc73b20b6815298a94cd88a5a844af88

Observation 8365d89e-3c8f-480b-913e-31a5ee19eee5 · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.985798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.244178Z digest=sha256:a3185c1532412b413083c07db75b774f2235aabc38cbee1a98b84a1427ea87e4

Observation 17a4cd68-123f-4ab6-a9b6-59c5b94b9864 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.248879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.248879Z digest=sha256:050efa2fdeaaee903a113d007090ab05cface9cc8864fa647f6e9d8b98f3f384

Observation 78a57c11-57c5-49b9-a415-014a2697bd64 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Msr-vtt: A large video description dataset for bridging video and language

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.253320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.253320Z digest=sha256:9a3acce27e68afc8fa69f4160872ccbfaabf5c3f11654ea28db95dad89cbd4ee

Observation b9baf577-4213-4296-898f-e4561352466e · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.842118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.258279Z digest=sha256:21e45e1b85efc7bf522c8b4d28d419e605d54dedbf548da31d2b5e035e1e0e34

Observation b86eb911-55bb-4a55-ad71-570e44686c28 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness xgen-mm (blip-3): A family of open large multimodal models, 2024

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.262867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.262867Z digest=sha256:22ce91000c7511c46c09975d8b400049609d2b24f13c9338532d11b6415440e2

Observation a80bd366-e3b8-4c07-bc84-872b9aa404ef · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.270705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.270705Z digest=sha256:bb3789e08d571ea1dcfe239869ba0b3fcd9d8875648ff70ca2ddf4ee58011faf

Observation 7120bbcc-1364-4dbf-9b61-41dc9cd72a46 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.814951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.274065Z digest=sha256:9ef1d53d74714afacd7d4bced08525ad82e57eca086b8b0212cbb543516fefe0

Observation 99cbcbe2-93a1-4092-83f4-6ed80486ba15 · outbound

This paper cites Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.805828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.277803Z digest=sha256:da48330ecd57b92f9d4fc2faa504029753204deab30fd4b04b80b77a3c1e767b

Observation 1d5ac5bb-0e4d-495e-9633-9bcc71042d74 · outbound

This paper cites Auformer: Vision transformers are parameter-efficient facial action unit detectors.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Auformer: Vision transformers are parameter-efficient facial action unit detectors

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.796041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.281283Z digest=sha256:482b2d1cb738f16e277ce3cb4cc93bc7d98a76fd4395f93fc55a460d3aef5731

Observation 262d4719-76df-40e9-88b2-d0e486e2a57e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.285100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.285100Z digest=sha256:9bfbb7cb637c7dd58f7f9b1786593e74d2d75960decb29f9dd1bc2f4df5da4d0

Observation 1ffd1b68-9ee3-4d03-8634-e20268023f95 · outbound

This paper cites Vision Transformer with Quadrangle Attention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vision Transformer with Quadrangle Attention

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.289229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.289229Z digest=sha256:b737b644cee9d9f88e8e83f40cf8e55796f8039b7fa4e22febe012e4d6433fa5

Observation 9812e6b7-9531-4f92-aa79-f3b95bdc06a8 · outbound

This paper cites Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.785326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.294736Z digest=sha256:97ff2ef08991df4c83fec0d2f8e6601092fcd13b4dbdf0aff4e12c3be63e2799

Observation ca82cb7d-5634-4986-921c-5a23b8072967 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava- next: A strong zero-shot video understanding model, 2024

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.774625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.298057Z digest=sha256:074b02b8dd069a7e74248c46288731960a9374cc42d3baa8e1d0853bec014548

Observation e265d731-5a5a-48a6-beac-ca464787be71 · outbound

This paper cites Facial expression recognition from near- infrared videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial expression recognition from near- infrared videos

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.764001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.301613Z digest=sha256:4a324544831a15c6728794668aad579b6ec0b8e92f3f240a3219d4ab5f83fd9c

Observation eace1683-3784-4d4d-a46a-932b0a4d921c · outbound

This paper cites LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.304945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.304945Z digest=sha256:305f49ce8ea4b61b5cdf7da256adc3a176c1102cc2c8847ac622d6121170946f

Observation 14e9c2b1-9476-4ab0-b54c-3afe638013e0 · outbound

This paper cites Deep region and multi-label learning for facial action unit detec- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep region and multi-label learning for facial action unit detec- tion

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.751577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.308759Z digest=sha256:efb09f0e762d7cb61e6a0ee0c0d3f3a98958d0af1f9b0c3ca6aebf1824a92e56

Observation 6d9ec7d2-adf8-4084-9f4a-740fce5b0c11 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MLVU: Benchmarking Multi-task Long Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.312278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.312278Z digest=sha256:6f3b5d93f334e417bacdc861092105f6b30395600dc8f686225ffc2991f69e0f

Observation 2276db77-201b-497d-a22f-600d3dd25b63 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Towards automatic learning of procedures from web instructional videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.739302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:33:54.316060Z digest=sha256:f0a43e1518f46d16933a86a09acce097610667c7d284446c4feed66886c3abdc

Observation 30e8f1e5-610c-4d29-a881-238bc6134e5d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.320561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.320561Z digest=sha256:105421fde55362f59bcc0bd44159199575dc630ec796a0082bb8a003725b50c7

Pith citing papers

Observation dfbf3f28-6334-4a74-834c-57a2e45d06e1 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.649306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:d14c1bc9a96b0fbe32e59f2ea0c52712177533dc422ea6264cc59c23b98612d2

Observation 9ba63ce5-3dbf-4d8c-90ca-ef657731fb73 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.796437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:1596cb3d2a80e4fd39ac2a857c95c07e6368520ac259edd72763f7ad9e3e3fb2

Observation 2f1f05b9-e8db-4b09-a634-0d7783dd5351 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.541729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.541729Z digest=sha256:51d803bdf367a0c4850e9c79e0ff12b7b92943a79b6c484fb238defb60800b77

Observation 260b9555-6759-4738-8c89-c4f26d70f407 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.336074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:fd980c0c1a3d7677611fae7de56df7adcd290d23e38191db111f5ad23ec0e8f3