Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T08:50:27.871886Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2604.23173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T08:50:27.871886Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5fd60d4-5da1-42ae-b3e7-634e9d122bfd · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition https : / / github
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e9b4570f-374a-49b7-92d7-b5057b6746cb · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c8cdd0e-be70-4ef3-a57f-28475efef9e0 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Frozen in Time: A Joint Video and Image Encoder for End- to-End Retrieval
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f26f2c02-c9c7-4403-8da3-743bb2554930 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition XMem++: Production-Level Video Segmentation from Few Annotated Frames
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6759dafd-8868-4720-a536-791728ff7481 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Corefer- ence Resolution through a Seq2Seq Transition-Based System
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01b4c86a-3986-4123-8c92-63e253b13e1c · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c71bc4cd-8fde-46c5-a55c-0526d8728d55 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Joint Multimedia Event Extraction from Video and Article
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7eebcc19-a9c6-4642-b3ee-d7820a122961 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 767e1352-4543-45af-b6e0-64216e8d3dde · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ShareGPT4Video: Improving Video Understand- ing and Generation with Better Captions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38d54ce3-335a-452d-9fb3-7ef616a232c9 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition V AST: A Vision-Audio- Subtitle-Text Omni-Modality Foundation Model and Dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f02f9336-0ece-4db3-967b-0c49fc8fd30b · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition XMem: Long- Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 119424ae-d507-4a02-a73c-9eb699106c1d · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Segment and Track Anything
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fae2782-2031-47a0-835c-4875430e5a47 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Zero-Shot Video Question Answering with Procedural Programs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebe15fdc-e2b7-4bc7-8045-685d8c28d27f · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Word-Level Coreference Resolu- tion
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 214f6afe-4e1a-42f3-a79b-c9e0099d7b46 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition DAPS: Deep Action Proposals for Action Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67dc4c8f-9d1e-4c53-8e4b-20ee3b7e09ca · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SlowFast Networks for Video Recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dee19dc4-5c2e-46f4-bae7-be3e3a394e14 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Action Transformer Network
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c27d6c7b-6794-4172-b272-4b384c50cce8 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Semi-Supervised Multimodal Coreference Resolution in Im- age Narrations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83027b6e-c2ac-4f63-929d-038cb2a6290d · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Who Are You Referring To? Coreference Resolution in Image Narrations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a26fbe97-bf40-4db1-a69d-2b6dddd5469d · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3fdf8a4-28ef-4508-b572-cfee712e33bd · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Fast Temporal Activity Proposals for Efficient De- tection of Human Actions in Untrimmed Videos
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a2ab035-2da3-4b5c-a2ea-b7e55174ded1 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VTIMELLM: Empower LLM to Grasp Video Mo- mentsVideo Action Transformer Network
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b806f9ed-f97a-4ab1-a9ea-5a112b30d0b0 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Action Genome: Actions as Compositions of Spatio- Temporal Scene Graphs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29f8e0d1-3f5a-4d90-9289-3cb86b4cab97 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6ef0d16-4bc6-46f6-b17e-91e79511920e · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Grounded Video Situation Recognition
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01e556e0-a124-4720-a4dc-5386ad82038b · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Kingma and Jimmy Ba
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation acc491be-d0f4-4bdb-ae1c-1d5988e66a5a · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.International Journal of Computer Vision (IJCV), 123(1):32–73
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d6ccf2e-2812-4b8d-8ae5-0f470269bbe9 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role Labeling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43975666-5855-4df1-a52a-50d1b6998428 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition TVQA: Localized, Compositional Video Question Answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b265e963-c916-49a7-bf23-ee5953bf9386 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition BLIP- 2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 928ae01a-504a-4a23-b085-ae0cc28929cb · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VILA: On Pre-training for Visual Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66527e88-d5a1-42a0-8c01-697c15cf6ac1 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67b043ed-f8c5-443c-b858-eb5a247b522e · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69887a59-b9d2-4306-98ad-b43f2acfe038 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MoMA: Multi-Object Multi-Actor Activity Pars- ing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c01aaa3-64fa-4af2-934a-7604b493ac9e · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2927b7c-2082-4279-859d-a7d2ccf2b0a4 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Major Entity Identification: A Generalizable Alternative to Coreference Resolution
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5fe1333-ec10-42ea-8935-3cdc449df7de · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92fa1bf3-ae89-4bbe-b20a-e9b5c90b7aaf · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b3e6cc8-b14e-4638-8b36-cfec3d5b77ee · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Which Coref- erence Evaluation Metric Do You Trust? A Proposal for a Link-Based Entity Aware Metric
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ed80c0f-0db6-4f7b-9034-77b733219f79 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VideoGLAMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36ca9cc9-1f11-4485-b37d-b63476214d6a · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8c1142b-96fd-47f9-b722-c0084bee742b · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2dc87950-735b-4c4f-bdaa-b080e31da4e5 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Identity- Aware Multi-Sentence Video Description
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b1db321-013e-4e36-9cdc-b0edfb449562 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bc7ec90-94c3-40f7-9da5-038858828236 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Qwen2.5-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution.Blog post: https://qwenlm.github.io/blog/qwen2.5-vl/
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2e9ac69-14a1-427a-931c-a120ccb61b2d · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MICAP: A Unified Model for Identity- Aware Movie Descriptions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8da2b91-4671-4350-9274-29708ceac499 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SAM 2: Segment Anything in Images and Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0f4996a-2bc4-4385-bc85-a3f65f7ea9c0 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8d5395d-b575-43d2-a51a-a43a152573ed · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Movie Description.International Journal of Computer Vision (IJCV), 123:94–120
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c19619df-c08c-447f-a8a3-ff06349cf4ee · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Object Grounding using Semantic Roles in Language Description
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2ca91a1-1a68-4777-8b50-07714d9f30cc · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Visual Semantic Role Labeling for Video Understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5846eb1d-9e95-412a-bfd5-a1708b1008ab · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VELOC- ITI: Can Video-Language Models Bind Semantic Concepts through Time? InConference on Computer Vision and Pat- tern Recognition (CVPR)
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ea43469-2cba-4504-9be5-d34441bf71e1 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Effi- cient Parameter-Free Clustering Using First Neighbor Rela- tions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e06e8275-7c63-4537-a501-26f0eb69d98c · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition End-to-End Generative Pretraining for Multimodal Video Captioning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 196b3e67-4de3-459c-adf2-537168ee7409 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Action Localization in Untrimmed Videos via Multi-Stage CNNs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d72c2cb-3914-404a-bdbb-35176f7b7336 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Actor-Centric Relation Network
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb637cd7-02cf-4ebb-ab84-343549e88ac6 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Unbiased Scene Graph Generation from Biased Training
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7671a9fc-1460-457d-815a-24df8c021001 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Long Term Spatio-Temporal Modeling for Action Detection.Com- puter Vision and Image Understanding (CVIU), 210
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 39124621-bd7d-405d-ac3f-12aa3f6ce112 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MovieQA: Un- derstanding Stories in Movies through Question-Answering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1e6903fd-efce-4b99-a116-d85e7b869e96 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition On Generalization in Coref- erence Resolution
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1cd80ef2-6eb8-4843-baea-d03e5eeb964f · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31aa72e1-dd44-4e54-9546-29c6c840f21d · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Lawrence Zitnick, and Devi Parikh
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe71fc0c-9156-40a5-b457-f4f455bc76f8 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Effec- tively Leveraging CLIP for Generating Situational Summaries of Images and Videos.IJCV
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30d3bcc8-0f56-4fc2-a711-ed08d1fcfa21 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MovieGraphs: Towards Understanding Human-Centric Situations from Videos
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1fe5fc1-dec3-463a-bbc3-86db0d6b851e · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Yoloe: Real-time seeing anything
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d63972a-4b40-421f-b0e7-c2224d9c66b8 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Segment Networks: Towards Good Practices for Deep Action Recog- nition
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a65125e1-11c6-4ed3-9600-e173705ebcfc · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9241731f-8b22-4df2-9138-2d8297dd70bf · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Demonstration Meets Typed Events: Type Specific Video Semantic Role Labeling via Multimodal Prompting and Retrieval
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 64dbb806-2a2a-4201-8e91-4382ba232fee · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Long-Term Feature Banks for Detailed Video Understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9ecbf17-778e-45ab-8246-33d12982c418 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7096acac-5063-44bf-9ecc-07f0b7ec918c · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition TubeDETR: Spatio-Temporal Video Grounding with Transformers
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bbd514c-21f6-41c5-b99f-4cea37362700 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cad60622-b58b-412d-a5c3-55d6378a9919 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Event Extraction via Tracking Vi- sual States of Arguments
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f126d71-9aa3-43ad-b798-6358d6c3dcea · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Track Anything: Segment Anything Meets Videos
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e40233af-9e7f-4d23-9105-e79b645cb135 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Panoptic Video Scene Graph Generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aaec3d26-26c8-4d62-a79b-42d8e42aa024 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16b9b879-28a0-4c3e-a26b-be991b26fbf4 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cb66a7f-62f9-4bd7-8f8c-178dc392acd6 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d83bff7-4fb2-40ed-9611-8d9cfe24d9be · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19e2f6b4-8eed-4640-bfb3-39f22eacf6fe · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5536d927-3cea-48a6-84b4-dc3db53c6d52 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video instruction tuning with synthetic data.TMLR
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b809598-985a-4b88-b931-2654ae26cb36 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role Labeling
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74a752a7-0e9b-4890-8f87-9a198f978264 · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Action Detection with 11 Structured Segment Networks
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3793d8cc-4406-40a6-adb1-e951abd084ea · outbound
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Streaming Dense Video Captioning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.