Pith. sign in

Paper Citation Record · LEDGER

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

As of 12 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2501.07978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07978 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:54.320561Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:24:36.541729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:02:37.644844Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc93ae32-cc76-4acc-ba60-cbec4923ba3e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.884278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.884278Z digest=sha256:1b5e8ac96cfb80724b1f5f74b5f53c1a9d7d822faae74c061c2b983d5b7ded9e

Observation 4abc8b02-edbf-4fe1-a5c3-b75fdb296c4a · outbound

This paper cites Emotion recognition in speech using cross- modal transfer in the wild, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Emotion recognition in speech using cross- modal transfer in the wild, 2018

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.889385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.889385Z digest=sha256:c090d7fc7c90fce152e764fd32cfc0525e46feaeff2d776aab43cae159f63539

Observation cdd82447-b527-47b6-8bcc-f1a8e953315f · outbound

This paper cites Claude-3.5, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Claude-3.5, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.893802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.893802Z digest=sha256:134d5276786608357a258c1c50b214cfa39443d8e6a0760b4c491bc7e8c4a082

Observation d6f6185c-f02d-4cf3-a693-a66927114450 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.903232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.903232Z digest=sha256:e03dc7724176fec4d7932448f119fccb929f01f2a72da281785203d5623ba226

Observation 8cbe17d3-6f5a-421c-ba14-585c11660183 · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Collecting highly paral- lel data for paraphrase evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.907800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.907800Z digest=sha256:fcf925298d0a07f9a80a2b1d7b1347227659b5bc8d73ea466c71feb5bc83cdb0

Observation 14417cbe-13db-43dc-8f44-377172cdc49b · outbound

This paper cites FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.913411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.913411Z digest=sha256:6320af4e793046ee0d11e24aac812ab147f5beb939f59db521774a0fcdc9ed82

Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.918678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.918678Z digest=sha256:930bea96e5ab951568372154b32b5a05b5d2a3af5ccd6b61a35e5aaa3993aa7a

Observation a70c08be-1453-477b-a074-daa4cfb9a060 · outbound

This paper cites Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.923207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.923207Z digest=sha256:37409349e3315727b5d834d38ed960e98489dc6fbe07c29241fa4410773b98ff

Observation 058f14fb-70b5-4ece-a08c-26afc70ea1c0 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.927345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.927345Z digest=sha256:59e09778b724397b9c60ac9345c63f0e8dba3813f7cb828e3d34c57e5b210ab2

Observation 8d653a64-b7f1-4921-926f-8b741118b8ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.930933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.930933Z digest=sha256:bf0aa00f67ce04ffae48be612468264cf02e9dd537ecea7fa7bbc78292fd008b

Observation 184a51db-02cc-49ff-a15b-11671cca3d01 · outbound

This paper cites Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.934690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.934690Z digest=sha256:716f72209ca380d7cf490b2229df0ba8f8870d12463ff8ddbb405f681102c0c4

Observation 70be547e-fdad-44e1-8463-20e460e9a722 · outbound

This paper cites Diffusionrig: Learning personal- ized priors for facial appearance editing.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Diffusionrig: Learning personal- ized priors for facial appearance editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.937932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.937932Z digest=sha256:9a5ed9af30a9d8772f88053d0f646ad42e87cda38873b788cada59d9bbdc4c4b

Observation 868dbd0f-f384-43a1-8d93-af5d185986ac · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.942983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.942983Z digest=sha256:dc2b1432ff8f97c11b10e4d3b1002fc05a8cd83ce5590d74cb5dfac137e57c70

Observation 6236c81b-5727-4fb4-8b60-8bc7774588a9 · outbound

This paper cites Strongsort: Make deep- sort great again.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Strongsort: Make deep- sort great again

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.947008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.947008Z digest=sha256:8f1cd04428ebac1c27e7e35ce5d5408c644de868c2347d967adc0f3935346b62

Observation 17472c72-cec0-40a2-99b2-84c691b507e7 · outbound

This paper cites EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.950534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.950534Z digest=sha256:631ed04507a74834e1ca743ecf1ff67ae268aff68acd4732ca8b374bba4be2df

Observation 8e20eb0f-a12a-465a-9847-5d95dbc2c3b4 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.954454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.954454Z digest=sha256:34496ea5a0f8359463c45a320e2d3ee8bf6b61ffa7afd9dbb200eee4984486ba

Observation 2ee3927a-96fb-42f8-b2fd-123ecbc2904c · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.958465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.958465Z digest=sha256:8c76eddac5fb69d4e0225ab7ca46f8050afac127985ec3a8afd2b8b392a1e77d

Observation 289d0ca9-4987-45a4-9795-01f08b933ada · outbound

This paper cites Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.963217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.963217Z digest=sha256:52ee2e061d69e18f16cbceb5bbf814f6a7b27e94e77d95ca82b61039584b949e

Observation b69d2990-0a6e-4675-a16e-24c27a8da406 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LoRA: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.967225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.967225Z digest=sha256:1be8e71d64b577adb41999354e6583b3308341ed0214c07643f3c87efc0bfa92

Observation b7602365-cb26-4dfd-ba34-4fa8ff056ba9 · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multimodal Pretraining for Dense Video Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.972371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.972371Z digest=sha256:f4451e24f871dbff9bcd195c4fe8d558fbfc73a48da6f76d15df4ad61b869700

Observation 4d501ff5-cfb1-45a9-8ef6-2707817343ae · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.977080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.977080Z digest=sha256:4277101e3086a17dcc6dad8315596e6d8c13e55f739d5690bd42c2413511e230

Observation 33d0f7e5-2d02-4c8f-9663-7bed3fa6c57e · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.982395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.982395Z digest=sha256:e0f9648196c13425bcf7fb8d2b415b29449713f18d4c2d92122ac30230a685ae

Observation 0d82b7f6-3584-472c-9619-29f1ddd9bf3a · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.986831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.986831Z digest=sha256:7458e0e98803a3ccd0150222e801f1357f485f534b0ce6e9ba2514352602cf03

Observation ac506810-02a2-41f5-af59-04498bbaecb6 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.991893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.991893Z digest=sha256:c3bd165a526ec983989f10dd546157265cb89598ca73ffc632b0269efa49c90b

Observation b19a64f9-dbfe-445d-8e6b-174955e70bb1 · outbound

This paper cites Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.996900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.996900Z digest=sha256:b9512946674ea7ebb4f4722434347f29994d320d6e7bb2e4cac4728c12e90484

Observation 49941daa-11fa-4191-86d9-ffe12cf973be · outbound

This paper cites Afew-va database for valence and arousal estimation in-the-wild.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Afew-va database for valence and arousal estimation in-the-wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.000691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.000691Z digest=sha256:e88c3b9fdceec2275f583834b4d1c669956c31f7677341e558c28eeb07fc0947

Observation 61377989-453c-4780-bcc3-f74d63bdc9c6 · outbound

This paper cites Dense-captioning events in videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dense-captioning events in videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.005726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.005726Z digest=sha256:ff06bfde6de64df8a154ebc2f93d468a2ca70b6a2dbd3aae55816f64188a4559

Observation b64a0f55-59a7-4ea3-b690-8dfe4fe2cbb8 · outbound

This paper cites Context-aware emotion recognition net- works.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Context-aware emotion recognition net- works

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.430420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.016846Z digest=sha256:8ed76c829c1cf522969ea3b3178401ff30d4cf605ad1dc86a33df32ea2366411

Observation c2632afb-5b1a-4883-b728-12af4c6b4f4e · outbound

This paper cites Llava-next: What else influences visual instruction tun- ing beyond data?, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava-next: What else influences visual instruction tun- ing beyond data?, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.416512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.021962Z digest=sha256:02620f852a8985c68fd49eac8b57ea25103110b576477e864e9fbddfb857711c

Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.026558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.026558Z digest=sha256:0ceca0198dfda741f4cce57a20bc61493475590abeeee16b07aef489dd19705c

Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.031175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.031175Z digest=sha256:c9a8b7c90ea35e5fbe3e7002ab8b20df8a3ed34ad890d7f47d0eeeb0a4144e79

Observation 7f8c1543-a945-4910-af7e-409144b7f88f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.035535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.035535Z digest=sha256:200daccaafd19d56c2131e6816fd7cf8d60ff79095976cbb6ee1f2dde0e13c38

Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.040489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.040489Z digest=sha256:2ad2105cf09dfcc74b72094883744140f9442ecede506774cb5d94ae61aef181

Observation 9094410c-9abc-4c79-bf0a-6f01cc769b6e · outbound

This paper cites Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.395528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.045144Z digest=sha256:9e4dbd5efefac7bd59580ed6e00df54cd6de96b73787dc557a32a297e1834618

Observation b5f712da-2f25-4616-844c-628f256cccd3 · outbound

This paper cites Facial affective behavior analysis with instruction tuning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial affective behavior analysis with instruction tuning, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.384319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.049977Z digest=sha256:38de478e605f0e12b89da6bc645417f86ea5d05052b72bafe5c4cfcf6afa7d33

Observation 88837485-f367-4190-a78b-1cc280f44245 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llama-vid: An image is worth 2 tokens in large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.373429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.054027Z digest=sha256:45946318dc2c22e5be7ccb205b06921fb047327740829ba358c64cffb78b2508

Observation 231af9a6-5224-4f1f-9aa0-310260037293 · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Photomaker: Customizing realistic human photos via stacked id embedding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.358412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.058317Z digest=sha256:13c169b1dae470da63ae863a3972a3e5ba08bc7fb8d1b463c6634e95329db7e1

Observation fa8073b5-3bd1-43cd-963c-2284534b242b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.063000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.063000Z digest=sha256:8d1eede3a323e23e373ff8cfdd4894f4dfe38efc2df4d62279d20b7a338b2bed

Observation c2d2eb12-c3d7-4f98-85e0-64962d28268b · outbound

This paper cites Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.342396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.068366Z digest=sha256:ae207d96316a3a644646cbdc39273b4c756d74f8c80458523aaad3142763f092

Observation b1ae5ed9-7d2a-466e-baf4-923d799a0625 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Improved Baselines with Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.072952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.072952Z digest=sha256:2dee8309168ab3a050a3b1849ccd8fe0860c2a97581b537274614449aada8d36

Observation deaaf7f4-3e1b-4922-a469-7e7c51f095d5 · outbound

This paper cites Visual instruction tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.079056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.079056Z digest=sha256:21de55ea435c4d84c0eff6f4a4dd697aeed570c5be8aa57a947285e8dfbe092c

Observation 2a8b8f70-e38a-45c8-8359-777e250eb065 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.083552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.083552Z digest=sha256:96f68dbd11288b1b63ad84f0209a153a8019f4b382d21a919059f723d520681e

Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.087918Z digest=sha256:067d62aff9c5a86039b1d79695e270661cef269b1a17128cd8399bc21432136b

Observation 4d075e73-3a34-4aa7-874c-120ee1be9681 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.092524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.092524Z digest=sha256:cf8edf21ff14c9d75f3d513637f750c862ff24ab236e1075e05bd16681302146

Observation 803469f6-deef-4297-b87b-92833edf75ad · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.321731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.098492Z digest=sha256:8c3f68aea203363b86fc157091be9ae1be079ad2dadb45d54a41c03bfabe7a8c

Observation e8f43fe3-bbf1-4cc1-bce9-5e410716a276 · outbound

This paper cites DamoFD: Digging into backbone de- sign on face detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness DamoFD: Digging into backbone de- sign on face detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.309781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.102402Z digest=sha256:9269741e8819707c2ce5d36ba2d96f4377a570474eacdc8d1b953571e0912c81

Observation 1a21361f-515c-4376-a7f1-c07d45c73d82 · outbound

This paper cites Livingstone and Frank A.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Livingstone and Frank A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.297610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.109723Z digest=sha256:3ace43c1ec20feb8822a8f090ebc33330803400417bd0f14bedcd3fd90fca001

Observation c9ccb828-d05c-4b47-ae5e-5b61b57bade5 · outbound

This paper cites Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.286051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.114235Z digest=sha256:43d7cb7e48f8b42bf5e64053b2159a4f6955ab6176f5555600e41409e071ac32

Observation 2d697a4f-8b27-4320-b587-e8742ea614d5 · outbound

This paper cites The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.273230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.118502Z digest=sha256:311d32b1589f10930004435c3261f4ef4466bc0cee9c29d4e47796d752a81e84

Observation b5c470e0-abcf-40aa-ba94-8c820facbbd1 · outbound

This paper cites Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.257974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.123016Z digest=sha256:d9c28bebdd3a1fb02cc63b6671ca56167ca1fe2cbfcb3f425d8b665ce8cbb978

Observation d0e643a4-7903-4f7c-91ee-b39f6e6d84e9 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.127444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.127444Z digest=sha256:86b102bc8c00c4a0dd54a67d0a05ec4efed26846421f890938113e74a908aaa4

Observation 8ffd7c41-e2d5-4641-a30c-47452385392f · outbound

This paper cites Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.246121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.131730Z digest=sha256:1e8eb9060177c9a9d3dd2511d259929a8a8c59ae59b1f9f6f04c77a1535c9a19

Observation e82a982a-ec88-45ab-9a48-d0a347527cc1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.135546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.135546Z digest=sha256:229fa1dfcb1addf74db0652028b77bb7456cc9f5e5984c2b03859318f7eceea5

Observation 710122ae-8fba-4ba8-aa6f-3f5913cf3f7b · outbound

This paper cites The importance of emotional regulation in mental health.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The importance of emotional regulation in mental health

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.227238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.139183Z digest=sha256:3e5662b859c81b543bfe797942b1181f7bd62f9d45a55b7fa8b0c4325b5ac05a

Observation 719aa5fe-2e0a-4e13-9845-7ea75642b62f · outbound

This paper cites FaceXFormer: A Unified Transformer for Facial Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FaceXFormer: A Unified Transformer for Facial Analysis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.143445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.143445Z digest=sha256:4787cbefb4f4e77c4bc3638d741dad6ad2671d78af3e0a17ae62a6562d864e70

Observation 9f37c381-ba5c-4eb6-8586-9ccfaf984da2 · outbound

This paper cites Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.215683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.148000Z digest=sha256:ea2aa1f39825d853ebdd22aeb225127d9317f551a6858403a7be11805475da72

Observation d662422b-4de8-4bca-bc29-6399216d3fe9 · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:55.202907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.152472Z digest=sha256:1a958f6f6f85425863fee823df91f6643a465375e99a9440d0719e45d711b88e

Observation f6f9678f-c6e2-4ba8-b2e7-9df8a433944e · outbound

This paper cites Gpt-4v(ision) system card.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4v(ision) system card

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.189876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.156778Z digest=sha256:376ca3f12d91508ff5e546ada4df940ac5adcf75989b1be59ee285c492d72940

Observation feebdaca-2d58-4d68-b35f-94a596c1ce9d · outbound

This paper cites Gpt-4 technical report, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4 technical report, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.178240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.160875Z digest=sha256:71898fa58183a987f8d5f6a63be717120fa5ebb12f71811bbdcba56557a3d0ef

Observation 68ab66d2-4688-4f8d-8a79-199d155881f5 · outbound

This paper cites Gpt-4o system card, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4o system card, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.166879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.167146Z digest=sha256:c9f945eeeacb028634fe8fc933ff36abac79a3fa8a2703349f6c906d24136821

Observation fed2f313-d31c-47f1-a744-9a6b3e0308f9 · outbound

This paper cites Digihuman: A con- versational digital human with facial expressions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Digihuman: A con- versational digital human with facial expressions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.152984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.171845Z digest=sha256:905842a3fa0f8af7ee8814be310893d2e98813ea17ff7b38372fc6a4967bfa34

Observation 1d57b4a1-9157-4f8f-a53c-d32938b2a18e · outbound

This paper cites Pantic, M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pantic, M

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.140334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.176751Z digest=sha256:c4a864473fcf7765ce21cf3ad2f5606babfbd5de934b0a741b065df9a8f1f771

Observation 5af1b47d-f09a-4c75-9ece-bce449cfd8f5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning transferable visual models from natural language supervi- sion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.181114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.181114Z digest=sha256:d91f954b2d4a7b6f93c693254c739aa103d51f2f74b7cc06c5c5e82d60c1c686

Observation 113794d7-32e1-42f4-adb7-8d05ccabfe0a · outbound

This paper cites Movie description.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Movie description

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.121966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.186009Z digest=sha256:527067b764272bc30a12eee0f7a58882e537690aacba01b308d2f8982fd9ba1a

Observation d423450c-b4bf-45a3-a3a2-63e59e6853fb · outbound

This paper cites Multi-view dynamic facial action unit detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multi-view dynamic facial action unit detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.111532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.190865Z digest=sha256:d8f22c05652bdddb3a31bd39b47dfa5422a4b79ceff0a037a43d5c4bdf55c68b

Observation adc5adc1-c310-4d96-94c4-3e30f01709c8 · outbound

This paper cites Deep adaptive attention for joint facial action unit detection and face alignment.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep adaptive attention for joint facial action unit detection and face alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.098590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.195507Z digest=sha256:104464561168bc319baab6f2e09a7bc2b089f689c302a81f55a998750b328f3f

Observation d57afeef-1816-4ffa-ab42-5dd4bc8d113b · outbound

This paper cites Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas).

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.084813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.200363Z digest=sha256:3a1d48591cd410c4e68ad63fd2adae93f751e576a0e56fc9f42e221b4881f70d

Observation e9a31929-d21b-4728-9114-4d844a9aa896 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gemini: A family of highly capable multi- modal models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.073989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.204861Z digest=sha256:3785bcb24fd7bb7af71262556cc5af6ab9eafb4204eb63e21ae95c83f1fdd625

Observation 2954039a-fbb4-45d5-a87f-b04dd8aaa782 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2.5: A party of foundation models, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.209214Z digest=sha256:457017331cd189bb8c292678411e054ce5d367462e9a6940f97892d5f8dcf074

Observation e2ee8a76-9ba8-42df-b627-373ccf0dfbdd · outbound

This paper cites Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.053978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.213735Z digest=sha256:a94f7ff948356c799d65cb05a39b73ef1ffd3982beaad9b77978b834a14c6798

Observation 96b16d77-d1c1-42b2-9937-6c6580196ac4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cider: Consensus-based image description evalua- tion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.042966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.218449Z digest=sha256:91c3d3103358358aa8bc39adfd15f0b5dbb7a6e9b81d3983815bd60803b2328f

Observation fe9dd5c5-1566-4bc3-808d-f1c7e92a8f29 · outbound

This paper cites A survey on the pipeline evolution of facial capture and tracking for digital humans.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness A survey on the pipeline evolution of facial capture and tracking for digital humans

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.028702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.222398Z digest=sha256:6d49d9286a44cf6f7942642fe24107bbd8d99ef8158709aa4b068e4fd89b7a6a

Observation f278228a-cd22-4c69-906e-39655cb90a70 · outbound

This paper cites Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.017717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.226871Z digest=sha256:3e68ef04877c805c144301158b73ab8bbb230627ba0de2a8818ad8b4e93f34ff

Observation 0b0fa2dc-26a1-483e-b33e-485eb0196ed9 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.007121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.231149Z digest=sha256:5e60825200cfa89137f74202be82e5b02300001f98c5d5750c390c6d6ff2d719

Observation 329bd536-d2db-4793-a525-e6978bf8566a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.235456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.235456Z digest=sha256:4c67940532463d05754e737316a051398a5c703d84208052680866c88b4c4844

Observation f7cb10f3-2000-433b-95c5-bdfcf0a0d9a0 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.996676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.239911Z digest=sha256:79d5b945640fdf99f24bcff3691b9bfee56d5ac60f25efdb49ed4692ee496bc1

Observation 8365d89e-3c8f-480b-913e-31a5ee19eee5 · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.985798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.244178Z digest=sha256:20c1306dff2af1ef497569badf60d6dd859100b3cfef394795c082f0ba1f56bd

Observation 17a4cd68-123f-4ab6-a9b6-59c5b94b9864 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.248879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.248879Z digest=sha256:050efa2fdeaaee903a113d007090ab05cface9cc8864fa647f6e9d8b98f3f384

Observation 78a57c11-57c5-49b9-a415-014a2697bd64 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Msr-vtt: A large video description dataset for bridging video and language

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.253320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.253320Z digest=sha256:9a3acce27e68afc8fa69f4160872ccbfaabf5c3f11654ea28db95dad89cbd4ee

Observation b9baf577-4213-4296-898f-e4561352466e · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.842118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.258279Z digest=sha256:faa50cd59fa083eae8570db966d1f629b3b6193fd599472485f4bb17e2703b3d

Observation b86eb911-55bb-4a55-ad71-570e44686c28 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness xgen-mm (blip-3): A family of open large multimodal models, 2024

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.262867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.262867Z digest=sha256:22ce91000c7511c46c09975d8b400049609d2b24f13c9338532d11b6415440e2

Observation a80bd366-e3b8-4c07-bc84-872b9aa404ef · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.270705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.270705Z digest=sha256:bb3789e08d571ea1dcfe239869ba0b3fcd9d8875648ff70ca2ddf4ee58011faf

Observation 7120bbcc-1364-4dbf-9b61-41dc9cd72a46 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.814951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.274065Z digest=sha256:08787765cb3e5b3fad66615bf26180a57de8ad42785e9d61ec59dc67b4b240be

Observation 99cbcbe2-93a1-4092-83f4-6ed80486ba15 · outbound

This paper cites Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.805828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.277803Z digest=sha256:545bf4cbc1036bcd1f6755a1c1d8ae847337c9737aae666583bbd52941b8713a

Observation 1d5ac5bb-0e4d-495e-9633-9bcc71042d74 · outbound

This paper cites Auformer: Vision transformers are parameter-efficient facial action unit detectors.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Auformer: Vision transformers are parameter-efficient facial action unit detectors

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.796041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.281283Z digest=sha256:ecbdf137174ff26cc82a5be8b9acccbbcc0ba9b2b126a7433a8a950d4c80eb65

Observation 262d4719-76df-40e9-88b2-d0e486e2a57e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.285100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.285100Z digest=sha256:9bfbb7cb637c7dd58f7f9b1786593e74d2d75960decb29f9dd1bc2f4df5da4d0

Observation 1ffd1b68-9ee3-4d03-8634-e20268023f95 · outbound

This paper cites Vision Transformer with Quadrangle Attention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vision Transformer with Quadrangle Attention

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.289229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.289229Z digest=sha256:b737b644cee9d9f88e8e83f40cf8e55796f8039b7fa4e22febe012e4d6433fa5

Observation 9812e6b7-9531-4f92-aa79-f3b95bdc06a8 · outbound

This paper cites Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.785326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.294736Z digest=sha256:db3fbbba5d2e7beccd59a70e2432b3234e8bb8aabe2c86d3e2332738ab5e24b9

Observation ca82cb7d-5634-4986-921c-5a23b8072967 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava- next: A strong zero-shot video understanding model, 2024

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.774625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.298057Z digest=sha256:f98fcba35314359b2351ef0a6d50940d513c36e76e55d2c4c2b28bee073c693a

Observation e265d731-5a5a-48a6-beac-ca464787be71 · outbound

This paper cites Facial expression recognition from near- infrared videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial expression recognition from near- infrared videos

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.764001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.301613Z digest=sha256:03774965fee7459da1ac246533f903c0f7fea63ce507a416161819940e50d992

Observation eace1683-3784-4d4d-a46a-932b0a4d921c · outbound

This paper cites LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.304945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.304945Z digest=sha256:305f49ce8ea4b61b5cdf7da256adc3a176c1102cc2c8847ac622d6121170946f

Observation 14e9c2b1-9476-4ab0-b54c-3afe638013e0 · outbound

This paper cites Deep region and multi-label learning for facial action unit detec- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep region and multi-label learning for facial action unit detec- tion

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.751577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.308759Z digest=sha256:49b9d4e983c59bcf4b051287fdea52af98c6bab696801981f7c5fdc7f01c90ca

Observation 6d9ec7d2-adf8-4084-9f4a-740fce5b0c11 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MLVU: Benchmarking Multi-task Long Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.312278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.312278Z digest=sha256:6f3b5d93f334e417bacdc861092105f6b30395600dc8f686225ffc2991f69e0f

Observation 2276db77-201b-497d-a22f-600d3dd25b63 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Towards automatic learning of procedures from web instructional videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.739302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:33:54.316060Z digest=sha256:b3717e9b59b4a2c47f6719cef8a653db50ecdc189a68caefb5fa7843bb36ec36

Observation 30e8f1e5-610c-4d29-a881-238bc6134e5d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.320561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.320561Z digest=sha256:105421fde55362f59bcc0bd44159199575dc630ec796a0082bb8a003725b50c7

Pith citing papers

Observation dfbf3f28-6334-4a74-834c-57a2e45d06e1 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.649306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:2f9151d5eee187982519cb96576d9cb674564a42dbe7418006076008cd0bd81f

Observation 9ba63ce5-3dbf-4d8c-90ca-ef657731fb73 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.796437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:aa07caddf1b899d1e51b07df0781696e6497f733274cbbb2114225ca08cdfe32

Observation 2f1f05b9-e8db-4b09-a634-0d7783dd5351 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.541729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.541729Z digest=sha256:51d803bdf367a0c4850e9c79e0ff12b7b92943a79b6c484fb238defb60800b77

Observation 260b9555-6759-4738-8c89-c4f26d70f407 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.336074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:ec7a68f199927d9127dd478abbc6d9a37127b44c3cff40c3a24b53695ede58eb