Pith. sign in

Paper Citation Record · LEDGER

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

As of 22 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 6 inbound Pith citation observations for arXiv:2504.18152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18152 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:37.071305Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:14.442972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:35.322765Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6253362-54a3-42af-bb60-8778b1f7f626 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.425862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.425862Z digest=sha256:fb09516e60c11c146296160d454a5fe9a70d2eb76671a92b605da3ea72765420

Observation f1d6365c-2465-476a-922b-1ea3cc9ce685 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.432259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.432259Z digest=sha256:e4690d8d77ca118b9c87e03fa0915218af8709be65d50d79e01db56dcb1ed51b

Observation 05145958-747c-492b-9dd0-60300cec6725 · outbound

This paper cites Temporalbench: Bench- marking fine-grained temporal understanding for multimodal video models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Temporalbench: Bench- marking fine-grained temporal understanding for multimodal video models, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.439761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.439761Z digest=sha256:050fd8976314508232ec135b2d0ab1d957d4ff8e85122a63cd0bb782c0a29e1b

Observation 4c581161-9c3c-4e7f-9ecc-1e498632452c · outbound

This paper cites A short note on the kinetics-700 human action dataset.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding A short note on the kinetics-700 human action dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.398654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.446054Z digest=sha256:14a598bf731a6565a42ed95958770fa216e8a52d739b181af033871362c6f340

Observation b91f3263-6b80-45b1-8918-70a6136a25cf · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.380456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.451897Z digest=sha256:73605f94ddd4c1fcebfd5bbcdb8c16c21a16bfe6d6947f6889f84ce5102c62d9

Observation 329b37c3-0136-49cf-a6d1-28dd703339eb · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.458263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.458263Z digest=sha256:97779506bb70b78938625950708ebf33cbcec0cec39f8211927095ba52ffae67

Observation 40c71f85-ad60-44fb-9b4c-f45d2e519b2a · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.467816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.467816Z digest=sha256:437ad07d66ba734f98f3bc6fadfe46b7a5997c93d977adbf5ba5412f2bd6e58f

Observation f4d09198-946a-4eac-ba30-656d6945dc2d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.480460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.480460Z digest=sha256:069ce2847ec6b1f9729c33206c35852385c34b752b49ab7defba0eda0c168dfe

Observation 1a7b1454-1c5f-4063-a6ca-d967f32c9de9 · outbound

This paper cites InstructBLIP: Towards general- purpose vision-language models with instruction tuning.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding InstructBLIP: Towards general- purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.332644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.487976Z digest=sha256:b543af33ae7af826f52fd725ad7156af871256591cfd094e8ba4cbcb91d10236

Observation 3a73caf0-d67f-4b98-8e3d-87317ae79772 · outbound

This paper cites Slowfast networks for video recognition.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Slowfast networks for video recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.274061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.499430Z digest=sha256:ca88bb0918e5a3f450dd7e70a417eeda7012bd779275ffd7095028f099664ef4

Observation 37765833-011e-4f85-9fad-f1d14ca330d5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.507212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.507212Z digest=sha256:90a69f11e4f6aa315403953abfb93a70108ac651e81af0877cea27dc88a092d7

Observation a43f6846-b256-42ed-9430-f7dddfd610b4 · outbound

This paper cites Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.256530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.601046Z digest=sha256:c2202be70acf61fadbbf53c1cddc668fa69fc27bc2386c7dce7e6167199d1c5a

Observation fd241236-c87c-4a43-962a-86e8d947695c · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.655866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.655866Z digest=sha256:4254a2fd9dbabe445182796f7766b15a1213bdd53e8f573395eaa1f8a4df863d

Observation 0d737399-fb69-4479-95d1-19112e32c04c · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.743325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.743325Z digest=sha256:1b3e932339ea3a17337fd41b28cf9e3b8dfe8212eacb07cc299c248b043d753f

Observation 41d6ab48-bd36-4f64-b31e-afbb1ca58292 · outbound

This paper cites Sapiens: Foundation for Human Vision Models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Sapiens: Foundation for Human Vision Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.749272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.749272Z digest=sha256:e46c208d947cc158c43eeefe417583fe27995f740b71be133a7e2ab15ff5245d

Observation dc417616-a84c-452c-b250-42d4fea21a43 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.754788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.754788Z digest=sha256:dae19c144cbbf568fd57d60a7ac258b040a8c22ee5aa541085387e592c80f90e

Observation 8ebaf115-fb52-4ca2-b111-9a730eace7e2 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.760710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.760710Z digest=sha256:72aca622e762c1ad3faacbf48570e29f20a5a58a8629092584e94f54cd60f106

Observation aca8bcb6-6517-4cc2-a3e2-54c5971eba90 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.766075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.766075Z digest=sha256:869f07c73e4237a1fced7c79dc764132a72d7d736e592dd118394e0dd6f63878

Observation bc613c64-309a-44d4-9dde-200c4c6776d6 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.190267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.771185Z digest=sha256:449216c59222b39ac8230bf062cb17ce5b5a3b575904ae6aeb83a478b3aeba7e

Observation 3dc5051b-9789-4ce7-884e-f4da1997bf7d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.776594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.776594Z digest=sha256:de8b70371035bd73c78d1b6e4902bcbbd40e291051ebe01bece666f0f191cd0b

Observation 3248a07f-f935-4cb6-904c-320ac50f0629 · outbound

This paper cites Visual instruction tuning, 2023.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.782795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.782795Z digest=sha256:34e585c93923a8b5c212d597b2dbdf8930d5aa85330c7d545a73b9b8468a1fbb

Observation cc1867f8-7998-40ef-9541-f9787e7743f8 · outbound

This paper cites Ntu rgb+d 120: A large- scale benchmark for 3d human activity understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ntu rgb+d 120: A large- scale benchmark for 3d human activity understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.118678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.788075Z digest=sha256:98d813b00dc5f47e852f83612c09f64765baef825a9bda23d9bcf9249b5810df

Observation 4c1d7de1-4828-45e5-b996-3f86c13ee30e · outbound

This paper cites Oryx mllm: On-demand spatial- temporal understanding at arbitrary resolution.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Oryx mllm: On-demand spatial- temporal understanding at arbitrary resolution

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.082180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.793796Z digest=sha256:22e22644ebf9ae76ac8f0d35723465314c82b0e621d1948239afb58e168dabc4

Observation 2c8e4346-2e6c-4c6e-85c4-0cab8469c22b · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.799856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.799856Z digest=sha256:7fecb001c66f7d95f3578d455ba26d5f7fa588b8c92af452f758442a538efd93

Observation 1c93033d-365b-41b5-9fc1-f226698f9e75 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.805393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.805393Z digest=sha256:7ed6ae8025c007bbcd82793f8b5c3eb1d4b1b4e63c576a31a14b132697e3db69

Observation 58167ab2-01b5-4f41-9e57-1e8d55d8feb7 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.811052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.811052Z digest=sha256:aeb3d39bf4f36fecb35db86c6766d8ab50f33735578860307231108be6c4b5ab

Observation 87a81169-c87f-4371-b592-ce194c3de45a · outbound

This paper cites Videogpt+: Integrating image and video encoders for enhanced video understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Videogpt+: Integrating image and video encoders for enhanced video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:39.028094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.817414Z digest=sha256:7b4a58d1a0a72515ea074af7ee6248da8dcbb7cd969e8d4a888a8786812e8955

Observation 2404dd98-3b13-478b-9b69-05872dacaccf · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.824393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.824393Z digest=sha256:2e15abc3d073687d26c1ddf530fb88f2c359f75691c1b4d85404c2e68be41853

Observation 45a0e9dc-b128-4de9-8f8e-a536556a3619 · outbound

This paper cites Gpt-4 technical report, 2023.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4 technical report, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.830768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.830768Z digest=sha256:fcbf8248acf9ea12880846bd1357a463095f6527a7dea9a1f3743fd9a9a68cfc

Observation 1fbf9ad8-0ab1-4517-b250-c4ae5f32a2bc · outbound

This paper cites Gpt-4v(ision) system card, 2023.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4v(ision) system card, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.837286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.837286Z digest=sha256:68c66aacff60848c18d8c9f5f1847c686885a19d70451841a29fef274b04b58d

Observation d13d5981-c6f5-4696-8da6-97e15c214402 · outbound

This paper cites Gpt-4o system card, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gpt-4o system card, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.853994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.853994Z digest=sha256:0940c4b7f1a9c131cb2f1bdfc249743ed79f82bbe6e1fc4936493659e454abd3

Observation 174a6462-593c-41eb-b7df-fe10d46741c0 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Perception test: A diagnostic benchmark for multimodal video models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.931839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:35.907190Z digest=sha256:20b683d4a29dea6f42692ec85c6668c1ba1092430f68e06ed8f9c943fb693c1c

Observation 20ee508d-8d02-44b8-8a5d-05183e3c305e · outbound

This paper cites Learning transferable visual models from natural language supervision.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.060981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.060981Z digest=sha256:cbb1711e51fc80203de517cd0c8f18abc6cce476bf5cf54ff9dceb9ed266b74a

Observation c4a76a82-f8b0-4242-938c-0b91e6da7fe8 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.140464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.140464Z digest=sha256:48a4ef699672ee40cd33ca4f4d6dae5acacda0751d114b0116c61e595abf705b

Observation aa54454a-577b-4060-87f7-d59a63c8f4fd · outbound

This paper cites Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.878757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.147447Z digest=sha256:5b2e83dcf7238812bd7f8119b271d1512d188e3cb46e17c1f9666976afa355e6

Observation 9731aca6-6b20-4f4a-a452-c233767bf89b · outbound

This paper cites Tomato: Assessing visual temporal reasoning capabilities in multi- modal foundation models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Tomato: Assessing visual temporal reasoning capabilities in multi- modal foundation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.857122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.154383Z digest=sha256:93b12fb0f008a62f73f90aec148ead38abb24504026a08b20c8981d5779dd110

Observation bfbb634d-5743-4eea-90d5-4d84d1699067 · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short- context accuracy, 2025.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Long-vita: Scaling large multi-modal models to 1 million tokens with leading short- context accuracy, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.780426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.159465Z digest=sha256:d38b922ec85468c891442d81df2457c4b1626321e90de60aed4f7ce2068e2c56

Observation 3cc4c6df-e62e-48f2-876c-d318a551c0cb · outbound

This paper cites HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.166347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.166347Z digest=sha256:c48831e6ee6ae94fcdb29022558cf71e5e6002862578b56ebf4dbf4a4af5158f

Observation 2529593f-ba0c-4976-88a6-107cb97e3802 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Gemini: A family of highly capable multi- modal models, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.740076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.172023Z digest=sha256:ecbef9f53de997b01d6d4b3e703d0286c993fc3c7a1869954632e98a7a30d112

Observation 6d6dffb0-8429-472d-aefe-93b2da44e934 · outbound

This paper cites Qwen2-vl.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Qwen2-vl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.714576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.213541Z digest=sha256:5739a8185d1627864085cfa2768065e3e485c1a12aea25e5968deb5a1925f3a5

Observation 85ba907a-3c2c-4fb1-80b2-f1914c26fd1c · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Qwen2.5: A party of foundation models, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.549206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.364039Z digest=sha256:aff7185cda61dbfc46bb01345ed75a7d7006d6f484f3d2203a693374ad88653e

Observation 46e5fd11-a4d3-489f-b500-6ef8e0a1259e · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.503074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.370205Z digest=sha256:5d2e40adcf15d735571cd0874879cc3e2f1d1639798107650f5fa7aa82ea6483

Observation ffb0ae59-5bc5-455c-b352-2f35074aed62 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.375192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.375192Z digest=sha256:a71321b882b3dd26526143f6bff37edb3d4be077da6184e752cdcfe35d7a6735

Observation d1587230-678f-4da8-b828-1f90a9e1755a · outbound

This paper cites Pargo: Bridging vision-language with partial and global views, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Pargo: Bridging vision-language with partial and global views, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.364587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.380895Z digest=sha256:7a30c983c1f447fe6365e4b2101a3ec1eb26d4678e63aa4b67596776fd74232d

Observation 2c791026-0d08-469e-9d35-84e5ebc4babc · outbound

This paper cites Mllm can see? dynamic correction decoding for hallucination mitiga- tion.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Mllm can see? dynamic correction decoding for hallucination mitiga- tion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.261983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.386270Z digest=sha256:b1b95eecceb0b0aecf85720d24eb0b08e7365f7e8f9fb2dac48ea6d90680d6e8

Observation 7de188c9-ffe2-41d6-8302-7162887e034e · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.242333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.392500Z digest=sha256:ad71956fc97df924f362336733a118774bee3f776ca5a869b9fd44286418a10c

Observation d3a4d7a4-0718-4bb5-81b4-306ad4f2d4e8 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.399286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.399286Z digest=sha256:0e315d01f0a73c8914daafe3b3924b66b3c2a1edd25335b0909d6e3681cd0d02

Observation eab762dc-5a3f-452b-b42e-b436e4da57ae · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.223055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.406495Z digest=sha256:52d908d68bfed724fe167949ce4049ca541d84620d1f994f9cc04f5dd1dfb779

Observation 4bbe0040-c715-4c48-b78b-cb591f91e34f · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.479792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.479792Z digest=sha256:cc3ab80b133019c5f00dee728f5355780fcdfa0724cca52a932d7daba93a0f2b

Observation c21d2ada-4f47-46f9-af86-9656484ec972 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding xgen-mm (blip-3): A family of open large multimodal models, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.203806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.593934Z digest=sha256:6eb642cd6a208f98dc1079750bffc2167e07c8c7332ea83a3a02dd23203883e5

Observation 09016623-c940-475b-b928-363be48ce85f · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.710272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.710272Z digest=sha256:b83b20cb1d676c6bc13a170b3e59e728d8d53087eaa48a4f71e8459e189b192b

Observation 27a221c2-a909-483f-8456-7ad23aa17b7a · outbound

This paper cites Dense connector for mllms.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Dense connector for mllms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.182838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.718295Z digest=sha256:946f852dadca07e39af08e6d98bcc62b5a78b6e2d4e7eaa7f0296a721202db65

Observation 3d590c2b-d881-4323-bd69-a367d09c6b26 · outbound

This paper cites Ureader: Univer- sal ocr-free visually-situated language understanding with multimodal large language model.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Ureader: Univer- sal ocr-free visually-situated language understanding with multimodal large language model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:38.071376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.725847Z digest=sha256:7fd5095fe63c29b8d0a817f88a13d7e7dcdfb60e034792403c5ea5fdb38e5d5a

Observation d973a054-a419-4338-ac97-a8e041bfa837 · outbound

This paper cites mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.732058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.732058Z digest=sha256:b0c1c05db963ea5075d25885acbeeafb3c5ebb6d1580b4f7ec185e1f16f75fe0

Observation 4160c195-12b8-4fa3-bbe4-45aa984a110f · outbound

This paper cites Sigmoid loss for language image pre-training.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Sigmoid loss for language image pre-training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:37.971709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.737522Z digest=sha256:ada966251409c62d1ce85fb800ec3914b378ff8abd087b81066a1cbaf8198274

Observation b668dc98-f0fd-4afe-ace5-1c2011107a11 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.742235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.742235Z digest=sha256:ecfb3ee2c7d89762e1caa7006ef63b5b578dc409261b065f3a07120afe072f88

Observation 639d2f60-5fca-44d7-a07d-b3cbfc53c950 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.748616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.748616Z digest=sha256:f6b98e931b268993ca99b070017f1b962e75a42d9fdf00168ecc19ea1496c9da

Observation fc0e204c-eed9-4d35-9456-4e775fb30aa4 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.781880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.781880Z digest=sha256:a68480ce73e7c28eaa4c1e8384993380a8d94a1991d02946af7343b7343d2fa7

Observation 2b862a0b-c9ee-4f52-9228-e2b09053df84 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.872302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.872302Z digest=sha256:02f62b0571fed3152b0d6cedae5869a4d013b81927226b53f723adfa447d9ec5

Observation 10104e21-c442-4319-b540-30acae1bb0f8 · outbound

This paper cites Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:37.937729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:36.998431Z digest=sha256:c97adea45a8b8444631218cd36872634fe314a5fbc5c87f8eed341659c9a1ecc

Observation b08087d6-1adb-4d80-8c93-60d602ae4690 · outbound

This paper cites HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:37.058354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:37.058354Z digest=sha256:4c87a90396fca206c8415bf3f3cccbae5551da1833661f80be1b70b57213d4bb

Observation 228cece3-713c-4d90-aedd-c406d5cb8c42 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:37.064945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:37.064945Z digest=sha256:a19d4a5bb4788569711ecb400d51a4b7c002abff6afb6d89da40ac20ce6fca47

Observation f058a47d-5ee2-425a-aa02-0059fffc0ae8 · outbound

This paper cites Apollo: An exploration of video understand- ing in large multimodal models, 2024.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Apollo: An exploration of video understand- ing in large multimodal models, 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:37.731660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T10:28:37.071305Z digest=sha256:23b2b8642c0ed077bc4f6769b8958b3f0d331e85169c0976ef113e19a78621bb

Pith citing papers

Observation da4ef8b4-87b9-44b4-bb5c-c0544257c32a · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.442972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.442972Z digest=sha256:ce8103a076233ea7eecdc32fd6aee0c868ce2c9f4b28248fdcaffc1c061ee3af

Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.338147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.338147Z digest=sha256:0460ace84ee8ec44134fc782fc84df9d9e7136a73ca8b66a25e1de14291a92f4

Observation f50d51cd-1905-4ed7-b73f-622651c3aa42 · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.056011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:8b22868c89f54cfa94ae80141aeeb88aca5a4953c43b32bb41461fbb3579a336

Observation 690280ac-a2d5-4e5b-b0dc-d60ab562d5e4 · inbound

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification cites this paper.

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.320458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:26:21.016303Z digest=sha256:f66c6f1f1d7faf7412f25e89d4f9e8347d1a674976b983b3a0dc76c6858524c3

Observation 251ad567-e0b9-480a-9e46-e318d7da0ca8 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.056250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T04:53:11.830488Z digest=sha256:d280aadbd4abf2370bf7427ae88eb17596b2a916fba9146222368f27624b61f6

Observation b58fd2bb-2e51-4fd0-a17a-1a51846e9090 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.324230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T09:46:56.063333Z digest=sha256:fd62b07372f9b724b084c77c4a321eaafb129fbac2ed8e33eb1294c79a56d43d