Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:54.320561Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2501.07978.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:54.320561Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:24:36.541729Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T06:02:37.644844Z
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc93ae32-cc76-4acc-ba60-cbec4923ba3e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4abc8b02-edbf-4fe1-a5c3-b75fdb296c4a · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Emotion recognition in speech using cross- modal transfer in the wild, 2018
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd82447-b527-47b6-8bcc-f1a8e953315f · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Claude-3.5, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f6185c-f02d-4cf3-a693-a66927114450 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbe17d3-6f5a-421c-ba14-585c11660183 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Collecting highly paral- lel data for paraphrase evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14417cbe-13db-43dc-8f44-377172cdc49b · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70c08be-1453-477b-a074-daa4cfb9a060 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058f14fb-70b5-4ece-a08c-26afc70ea1c0 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d653a64-b7f1-4921-926f-8b741118b8ba · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184a51db-02cc-49ff-a15b-11671cca3d01 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70be547e-fdad-44e1-8463-20e460e9a722 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Diffusionrig: Learning personal- ized priors for facial appearance editing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 868dbd0f-f384-43a1-8d93-af5d185986ac · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6236c81b-5727-4fb4-8b60-8bc7774588a9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Strongsort: Make deep- sort great again
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17472c72-cec0-40a2-99b2-84c691b507e7 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e20eb0f-a12a-465a-9847-5d95dbc2c3b4 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee3927a-96fb-42f8-b2fd-123ecbc2904c · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 289d0ca9-4987-45a4-9795-01f08b933ada · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69d2990-0a6e-4675-a16e-24c27a8da406 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LoRA: Low-rank adaptation of large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7602365-cb26-4dfd-ba34-4fa8ff056ba9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multimodal Pretraining for Dense Video Captioning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d501ff5-cfb1-45a9-8ef6-2707817343ae · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video ReCap: Recursive Captioning of Hour-Long Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d0f7e5-2d02-4c8f-9663-7bed3fa6c57e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d82b7f6-3584-472c-9619-29f1ddd9bf3a · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac506810-02a2-41f5-af59-04498bbaecb6 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19a64f9-dbfe-445d-8e6b-174955e70bb1 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49941daa-11fa-4191-86d9-ffe12cf973be · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Afew-va database for valence and arousal estimation in-the-wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61377989-453c-4780-bcc3-f74d63bdc9c6 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dense-captioning events in videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64a0f55-59a7-4ea3-b690-8dfe4fe2cbb8 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Context-aware emotion recognition net- works
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c2632afb-5b1a-4883-b728-12af4c6b4f4e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava-next: What else influences visual instruction tun- ing beyond data?, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f8c1543-a945-4910-af7e-409144b7f88f · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9094410c-9abc-4c79-bf0a-6f01cc769b6e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5f712da-2f25-4616-844c-628f256cccd3 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial affective behavior analysis with instruction tuning, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 88837485-f367-4190-a78b-1cc280f44245 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llama-vid: An image is worth 2 tokens in large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 231af9a6-5224-4f1f-9aa0-310260037293 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Photomaker: Customizing realistic human photos via stacked id embedding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa8073b5-3bd1-43cd-963c-2284534b242b · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d2eb12-c3d7-4f98-85e0-64962d28268b · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b1ae5ed9-7d2a-466e-baf4-923d799a0625 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Improved Baselines with Visual Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deaaf7f4-3e1b-4922-a469-7e7c51f095d5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Visual instruction tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8b8f70-e38a-45c8-8359-777e250eb065 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d075e73-3a34-4aa7-874c-120ee1be9681 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803469f6-deef-4297-b87b-92833edf75ad · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e8f43fe3-bbf1-4cc1-bce9-5e410716a276 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness DamoFD: Digging into backbone de- sign on face detection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a21361f-515c-4376-a7f1-c07d45c73d82 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Livingstone and Frank A
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9ccb828-d05c-4b47-ae5e-5b61b57bade5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2d697a4f-8b27-4320-b587-e8742ea614d5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5c470e0-abcf-40aa-ba94-8c820facbbd1 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0e643a4-7903-4f7c-91ee-b39f6e6d84e9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffd7c41-e2d5-4641-a30c-47452385392f · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e82a982a-ec88-45ab-9a48-d0a347527cc1 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710122ae-8fba-4ba8-aa6f-3f5913cf3f7b · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The importance of emotional regulation in mental health
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 719aa5fe-2e0a-4e13-9845-7ea75642b62f · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FaceXFormer: A Unified Transformer for Facial Analysis
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f37c381-ba5c-4eb6-8586-9ccfaf984da2 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d662422b-4de8-4bca-bc29-6399216d3fe9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6f9678f-c6e2-4ba8-b2e7-9df8a433944e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4v(ision) system card
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation feebdaca-2d58-4d68-b35f-94a596c1ce9d · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4 technical report, 2023
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68ab66d2-4688-4f8d-8a79-199d155881f5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4o system card, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fed2f313-d31c-47f1-a744-9a6b3e0308f9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Digihuman: A con- versational digital human with facial expressions
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d57b4a1-9157-4f8f-a53c-d32938b2a18e · outbound
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5af1b47d-f09a-4c75-9ece-bce449cfd8f5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning transferable visual models from natural language supervi- sion
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113794d7-32e1-42f4-adb7-8d05ccabfe0a · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Movie description
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d423450c-b4bf-45a3-a3a2-63e59e6853fb · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multi-view dynamic facial action unit detection
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation adc5adc1-c310-4d96-94c4-3e30f01709c8 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep adaptive attention for joint facial action unit detection and face alignment
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d57afeef-1816-4ffa-ab42-5dd4bc8d113b · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas)
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9a31929-d21b-4728-9114-4d844a9aa896 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gemini: A family of highly capable multi- modal models, 2024
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2954039a-fbb4-45d5-a87f-b04dd8aaa782 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2.5: A party of foundation models, 2024
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2ee8a76-9ba8-42df-b627-373ccf0dfbdd · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96b16d77-d1c1-42b2-9937-6c6580196ac4 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cider: Consensus-based image description evalua- tion
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe9dd5c5-1566-4bc3-808d-f1c7e92a8f29 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness A survey on the pipeline evolution of facial capture and tracking for digital humans
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f278228a-cd22-4c69-906e-39655cb90a70 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b0fa2dc-26a1-483e-b33e-485eb0196ed9 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 329bd536-d2db-4793-a525-e6978bf8566a · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cb10f3-2000-433b-95c5-bdfcf0a0d9a0 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vatex: A large-scale, high- quality multilingual dataset for video-and-language research
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8365d89e-3c8f-480b-913e-31a5ee19eee5 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17a4cd68-123f-4ab6-a9b6-59c5b94b9864 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a57c11-57c5-49b9-a415-014a2697bd64 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Msr-vtt: A large video description dataset for bridging video and language
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9baf577-4213-4296-898f-e4561352466e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b86eb911-55bb-4a55-ad71-570e44686c28 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness xgen-mm (blip-3): A family of open large multimodal models, 2024
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80bd366-e3b8-4c07-bc84-872b9aa404ef · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7120bbcc-1364-4dbf-9b61-41dc9cd72a46 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 99cbcbe2-93a1-4092-83f4-6ed80486ba15 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d5ac5bb-0e4d-495e-9633-9bcc71042d74 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Auformer: Vision transformers are parameter-efficient facial action unit detectors
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 262d4719-76df-40e9-88b2-d0e486e2a57e · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffd1b68-9ee3-4d03-8634-e20268023f95 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vision Transformer with Quadrangle Attention
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9812e6b7-9531-4f92-aa79-f3b95bdc06a8 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca82cb7d-5634-4986-921c-5a23b8072967 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava- next: A strong zero-shot video understanding model, 2024
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e265d731-5a5a-48a6-beac-ca464787be71 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial expression recognition from near- infrared videos
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eace1683-3784-4d4d-a46a-932b0a4d921c · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e9c2b1-9476-4ab0-b54c-3afe638013e0 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep region and multi-label learning for facial action unit detec- tion
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d9ec7d2-adf8-4084-9f4a-740fce5b0c11 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MLVU: Benchmarking Multi-task Long Video Understanding
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2276db77-201b-497d-a22f-600d3dd25b63 · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Towards automatic learning of procedures from web instructional videos
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30e8f1e5-610c-4d29-a881-238bc6134e5d · outbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbf3f28-6334-4a74-834c-57a2e45d06e1 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9ba63ce5-3dbf-4d8c-90ca-ef657731fb73 · inbound
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f1f05b9-e8db-4b09-a634-0d7783dd5351 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260b9555-6759-4738-8c89-c4f26d70f407 · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.