Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:37.071305Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 6 inbound Pith citation observations for arXiv:2504.18152.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:37.071305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:14.442972Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T09:54:35.322765Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6253362-54a3-42af-bb60-8778b1f7f626 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d6365c-2465-476a-922b-1ea3cc9ce685 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Activitynet: A large-scale video benchmark for human activity understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05145958-747c-492b-9dd0-60300cec6725 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Temporalbench: Bench- marking fine-grained temporal understanding for multimodal video models, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c581161-9c3c-4e7f-9ecc-1e498632452c · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding A short note on the kinetics-700 human action dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b91f3263-6b80-45b1-8918-70a6136a25cf · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Honeybee: Locality-enhanced projector for multimodal llm
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 329b37c3-0136-49cf-a6d1-28dd703339eb · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c71f85-ad60-44fb-9b4c-f45d2e519b2a · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d09198-946a-4eac-ba30-656d6945dc2d · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7b1454-1c5f-4063-a6ca-d967f32c9de9 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding InstructBLIP: Towards general- purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a73caf0-d67f-4b98-8e3d-87317ae79772 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Slowfast networks for video recognition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 37765833-011e-4f85-9fad-f1d14ca330d5 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43f6846-b256-42ed-9430-f7dddfd610b4 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fd241236-c87c-4a43-962a-86e8d947695c · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video ReCap: Recursive Captioning of Hour-Long Videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d737399-fb69-4479-95d1-19112e32c04c · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d6ab48-bd36-4f64-b31e-afbb1ca58292 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Sapiens: Foundation for Human Vision Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc417616-a84c-452c-b250-42d4fea21a43 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ebaf115-fb52-4ca2-b111-9a730eace7e2 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca8bcb6-6517-4cc2-a3e2-54c5971eba90 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc613c64-309a-44d4-9dde-200c4c6776d6 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3dc5051b-9789-4ce7-884e-f4da1997bf7d · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3248a07f-f935-4cb6-904c-320ac50f0629 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Visual instruction tuning, 2023
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1867f8-7998-40ef-9541-f9787e7743f8 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ntu rgb+d 120: A large- scale benchmark for 3d human activity understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4c1d7de1-4828-45e5-b996-3f86c13ee30e · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Oryx mllm: On-demand spatial- temporal understanding at arbitrary resolution
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c8e4346-2e6c-4c6e-85c4-0cab8469c22b · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c93033d-365b-41b5-9fc1-f226698f9e75 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58167ab2-01b5-4f41-9e57-1e8d55d8feb7 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a81169-c87f-4371-b592-ce194c3de45a · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Videogpt+: Integrating image and video encoders for enhanced video understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2404dd98-3b13-478b-9b69-05872dacaccf · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a0e9dc-b128-4de9-8f8e-a536556a3619 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4 technical report, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbf9ad8-0ab1-4517-b250-c4ae5f32a2bc · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4v(ision) system card, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13d5981-c6f5-4696-8da6-97e15c214402 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4o system card, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174a6462-593c-41eb-b7df-fe10d46741c0 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Perception test: A diagnostic benchmark for multimodal video models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20ee508d-8d02-44b8-8a5d-05183e3c305e · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Learning transferable visual models from natural language supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a76a82-f8b0-4242-938c-0b91e6da7fe8 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa54454a-577b-4060-87f7-d59a63c8f4fd · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9731aca6-6b20-4f4a-a452-c233767bf89b · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Tomato: Assessing visual temporal reasoning capabilities in multi- modal foundation models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bfbb634d-5743-4eea-90d5-4d84d1699067 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Long-vita: Scaling large multi-modal models to 1 million tokens with leading short- context accuracy, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3cc4c6df-e62e-48f2-876c-d318a551c0cb · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2529593f-ba0c-4976-88a6-107cb97e3802 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gemini: A family of highly capable multi- modal models, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6d6dffb0-8429-472d-aefe-93b2da44e934 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Qwen2-vl
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85ba907a-3c2c-4fb1-80b2-f1914c26fd1c · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Qwen2.5: A party of foundation models, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46e5fd11-a4d3-489f-b500-6ef8e0a1259e · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ffb0ae59-5bc5-455c-b352-2f35074aed62 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1587230-678f-4da8-b828-1f90a9e1755a · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Pargo: Bridging vision-language with partial and global views, 2024
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c791026-0d08-469e-9d35-84e5ebc4babc · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mllm can see? dynamic correction decoding for hallucination mitiga- tion
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7de188c9-ffe2-41d6-8302-7162887e034e · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d3a4d7a4-0718-4bb5-81b4-306ad4f2d4e8 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab762dc-5a3f-452b-b42e-b436e4da57ae · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4bbe0040-c715-4c48-b78b-cb591f91e34f · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21d2ada-4f47-46f9-af86-9656484ec972 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding xgen-mm (blip-3): A family of open large multimodal models, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 09016623-c940-475b-b928-363be48ce85f · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a221c2-a909-483f-8456-7ad23aa17b7a · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Dense connector for mllms
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d590c2b-d881-4323-bd69-a367d09c6b26 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ureader: Univer- sal ocr-free visually-situated language understanding with multimodal large language model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d973a054-a419-4338-ac97-a8e041bfa837 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4160c195-12b8-4fa3-bbe4-45aa984a110f · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Sigmoid loss for language image pre-training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b668dc98-f0fd-4afe-ace5-1c2011107a11 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639d2f60-5fca-44d7-a07d-b3cbfc53c950 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0e204c-eed9-4d35-9456-4e775fb30aa4 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b862a0b-c9ee-4f52-9228-e2b09053df84 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10104e21-c442-4319-b540-30acae1bb0f8 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b08087d6-1adb-4d80-8c93-60d602ae4690 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228cece3-713c-4d90-aedd-c406d5cb8c42 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f058a47d-5ee2-425a-aa02-0059fffc0ae8 · outbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Apollo: An exploration of video understand- ing in large multimodal models, 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation da4ef8b4-87b9-44b4-bb5c-c0544257c32a · inbound
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · inbound
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50d51cd-1905-4ed7-b73f-622651c3aa42 · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 690280ac-a2d5-4e5b-b0dc-d60ab562d5e4 · inbound
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 251ad567-e0b9-480a-9e46-e318d7da0ca8 · inbound
HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b58fd2bb-2e51-4fd0-a17a-1a51846e9090 · inbound
HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.